A court just ruled that training an AI on someone else's editorial work isn't fair use

Skye Jacobs

Posts: 2,239   +63
Staff
What just happened? The Third Circuit has ruled that Ross Intelligence infringed Thomson Reuters' copyrights by using Westlaw headnotes to train a legal research system designed to compete with Westlaw. The decision gives a federal appeals court's perspective on a question that has become central to the AI industry: when does training a system on copyrighted material qualify as protected fair use, and when does it amount to unauthorized copying for a competing product?

Ross developed an AI-based legal research platform that allowed lawyers to submit legal questions and receive relevant passages from court opinions. To train the system, the company worked with a third party, LegalEase Solutions, to create training materials that connected legal questions with passages from judicial opinions.

The material included Westlaw headnotes. Headnotes are brief summaries written by Westlaw editors that identify important legal issues in a court opinion. They are not part of the court's decision itself. Instead, they help lawyers quickly locate portions of an opinion that address a particular legal question.

Ross argued that the headnotes were too close to the underlying opinions to qualify for copyright protection. It also argued that its use of the headnotes was fair because it used them to train an AI system rather than republish them.

The Third Circuit rejected both arguments.

"Under Ross's framing, this case appears to concern the future of AI legal technology," US Circuit Judge Tamika Montgomery-Reeves wrote for the court. "But appearances can be deceiving. In truth, this is no more than an ordinary copyright case."

The court found that Westlaw's headnotes met the low originality threshold required for copyright protection. Editors had to decide which legal issue each headnote would cover, what facts and reasoning to include, and how to describe the point accurately and concisely.

"Each headnote had to convey enough factual background and reasoning to stand on its own without reference to the opinion or any other headnote, while accurately reflecting the corresponding opinion passage," Montgomery-Reeves wrote. "Because these decisions show the requisite minimal 'creative spark,' the headnotes are original."

That conclusion is important for AI companies working with material that sits on top of public records. Court decisions are public and cannot be copyrighted. But editorial content built around those decisions can be protected if it reflects independent judgment and expression.

The court also focused on the practical role Westlaw's material played in the Ross system. Ross did not use the headnotes merely to study legal writing or conduct internal research. It used them to help train a product that could answer legal research questions, which is a core Westlaw function.

The fact that Ross used the headnotes during training, rather than displaying them to users, did not settle the issue. The court said the training process was only an intermediate step in building a competing commercial service.

"Ross took an intermediate step of using the headnotes to train an AI program, which arguably presents a slight degree of difference in use," Montgomery-Reeves wrote. "But the undisputed evidence demonstrates that Ross used the headnotes to train an AI program for the benefit of its legal-research platform."

The court described that use as "minimally transformative at best." In other words, Ross applied a different technical method, but it was still pursuing a closely related commercial purpose. It used Westlaw's editorial work to improve a system aimed at the same legal research market.

The ruling also addressed market harm, a major part of the fair-use analysis. The panel found that Ross's use could affect the value of Westlaw's legal research services. It could also interfere with a potential market for licensing headnotes as AI training data.

"The undisputed evidence reflects that Thomson Reuters' materials possess a creative spark, and Ross aspired to be a direct competitor by using them for a highly similar purpose," Montgomery-Reeves wrote. "We therefore hold that Thomson Reuters' materials are copyrightable and that Ross' use was not fair."

The court's ruling does not create a broad rule for every AI training dispute. It was based on a specific record involving a direct competitor, copyrighted editorial material, and a legal research product that sought to perform much the same job as Westlaw.

Still, the case puts added weight on the source and function of training data. AI developers can often obtain public documents, but that does not necessarily give them the right to use a publisher's summaries, labels, rankings, or other editorial work built around those documents.

Permalink to story:

 
Good news. And not good as the same time.
The problem is this. What if we can force our AI companies
to pay for the work they "borrow"?
I mean China will follow, right?
No! It will keep stealing and thus making their AI
much cheaper, getting ahead, getting better
having no barriers of law or respect.

There is no way out of it. Either our AI companies do it, or theirs will.
And given how many dictators end up starting a bloody war, I would
fear China with leading AI going into a war for its dictator.
 
So basically, looking at other cases, training an AI system for a very specific purpose using specific sorts of documents may run afoul of copyright laws, especially if it directly competes with the source of the documents, but training a more general AI system with all sorts of random training data does not. (I know there were two lawsuits, one with Meta and one with Anthropic, that found their training itself did not violate copyright).
 
So basically, looking at other cases, training an AI system for a very specific purpose using specific sorts of documents may run afoul of copyright laws, especially if it directly competes with the source of the documents, but training a more general AI system with all sorts of random training data does not. (I know there were two lawsuits, one with Meta and one with Anthropic, that found their training itself did not violate copyright).


And it's so stupid. The AI summerize on top of google pretty much pushes visitors away, because it puts everything from one or multiple sources together. A win for Google, a not so win for the site owner, or who-ever created that content to begin with.

Like I block the complete META subnet from anything on websites of mines through Cloudflare, by just inserting their complete ASN. I'm active in a business in where Meta has no business in, but yet sends out hordes of crawlers, bots, to suck up data from websites of mines, without getting something in return.

Bot traffic is a serious problem. And the way AI gets it wrong is there too. If there's copyright on text, an image, an audio or video then there should be copyright for Ai as well. It's that simple.
 
Good news. And not good as the same time.
The problem is this. What if we can force our AI companies
to pay for the work they "borrow"?
I mean China will follow, right?
No! It will keep stealing and thus making their AI
much cheaper, getting ahead, getting better
having no barriers of law or respect.

There is no way out of it. Either our AI companies do it, or theirs will.
And given how many dictators end up starting a bloody war, I would
fear China with leading AI going into a war for its dictator.
I'm sure A.I companies compensating original authors won't make them go bankrupt, the billions of dollars the A.I business produce and will keep producing will be enough for them to sustain themselves and those getting compensated.
 
Good news. And not good as the same time.
The problem is this. What if we can force our AI companies
to pay for the work they "borrow"?
I mean China will follow, right?
No! It will keep stealing and thus making their AI
much cheaper, getting ahead, getting better
having no barriers of law or respect.

There is no way out of it. Either our AI companies do it, or theirs will.
And given how many dictators end up starting a bloody war, I would
fear China with leading AI going into a war for its dictator.

Yep, China WILL be the "leader" (owner) of the AI race. They are bankrolling
a lot of the anti AI-Data Center protests. They have MASSIVE (miles in acreage)
of solar panel & solar mirror farms along with huge water reserves out in the desert
areas of China, in hopes of getting the data center types to build there. And you can
100% bet, if/when they are built there, the CCP will have all the access to the data.
We've already seen some of their robots running at track meets, some boxing type
things. Yeah, they are a bit "sluggish", but just thing of their capabilities once the "AI"
things get programmed into them. LOL, and if it works out like it does in the movies,
they will turn on their makers.
 
Meanwhile the judge who already refused to break up Google even though it was found guilty of several antitrust violations has just ruled that AI Overviews are totally fine.

He justified the ruling with the nonsense explanation that sites can not expect to get traffic back from Google, because they didn't sign an agreement with Google about that.
 
I'm sure A.I companies compensating original authors won't make them go bankrupt, the billions of dollars the A.I business produce and will keep producing will be enough for them to sustain themselves and those getting compensated.
These so-called AI companies might well be producing millions in revenue, but how about profit? They will just continue to cook the books no doubt. Gotta get that juicy IPO. After all, (apparently)it's not real money is it?
 
Back