What happened
In Bartz v. Anthropic, before the U.S. District Court for the Northern District of California, the settlement announced in September 2025 became final on 20 July 2026 after a lengthy period of notice, objections and review. Final approval was granted by Judge Araceli Martínez-Olguín, who took over the file following the retirement of the judge who had presided over the case, William Alsup. According to public reporting, the settlement is now being implemented: claims are being processed and payments distributed to rights holders.
The development marks a turning point in the collision between generative AI and copyright. For the first time, an AI company accepted financial liability on this scale because of how its training data was acquired. The case has also become a reference framework for dozens of copyright suits against AI developers — but, as we will see, without creating binding precedent.
Background: books, shadow libraries and Claude
The suit began in 2024, when a group of authors brought a class action against Anthropic, the developer of the Claude models. They alleged that the company had copied millions of books without permission to train its models, downloading a large share of them from pirate "shadow libraries" such as Library Genesis and Pirate Library Mirror. It also emerged that the company had bought millions of print books and scanned their pages to build a digital library.
The court was therefore looking not at one act but at three: (1) scanning print books that had been purchased; (2) downloading books from pirate sources and keeping them in a permanent library; and (3) training models on text drawn from those sources. What decided the case was the court's insistence on analysing those three acts separately.
The key ruling: training and acquisition were separated
In June 2025, at the summary-judgment stage, Judge Alsup drew the legal map:
- Training is fair use. Training a language model on lawfully acquired books was, in the court's words, an "exceedingly transformative" use. The model does not use the works to republish them or to substitute for them, but to learn from them and generate new and different text. The court compared it to a person who reads widely and becomes a writer.
- Digitising purchased books is fair use. Scanning print books the company had legitimately bought, converting them into digital copies, was treated as a change of format and held to be fair use.
- Building a pirate library is not fair use. Downloading millions of books from pirate sites and storing them in a permanent central library was, however, a separate and independent infringement. That the copies were later used for training did not cure the unlawfulness of their acquisition. Acquiring a work that could have been bought through a pirate channel cannot be laundered by a legitimate downstream purpose.
That distinction split the case in two: the company won on training, but the piracy claim would go to a jury. When the court certified the case as a class action in July 2025, the risk multiplied, because the claims of every rights holder meeting the class criteria — not merely a handful of authors — were now consolidated in a single proceeding.
A timeline from complaint to final approval
- 2024: A group of authors sues Anthropic for copyright infringement in the Northern District of California.
- June 2025: Judge Alsup holds that training on lawfully acquired books, and scanning purchased books, is fair use; the pirate-library claim proceeds toward trial.
- July 2025: The case is certified as a class action on behalf of the rights holders of qualifying works.
- September 2025: The parties announce a $1.5 billion settlement, and the court grants preliminary approval.
- September 2025 – July 2026: The works list is finalised, rights holders are notified, and claims and objections are processed.
- 20 July 2026: Judge Martínez-Olguín grants final approval.
- September 2026: The settlement is in its implementation and distribution phase, with payments flowing through the claims process.
The calendar shows that almost a year can pass between announcing a class settlement and money reaching a rights holder. Much of that time went into determining which works were covered and how shares would be divided when a single work has more than one rights holder. Given the variety of publishing contracts, the author–publisher split becomes one of the thorniest technical details of settlements like this one.
Why settle? The arithmetic of statutory damages
Under U.S. copyright law, statutory damages — which can be claimed for registered works without proving actual loss — can reach $150,000 per work for wilful infringement. In a class action covering hundreds of thousands of works, that pointed, in theory, to exposure in the tens or even hundreds of billions of dollars. Facing a jury had become an existential threat for the company.
The parties announced a settlement in September 2025. Its publicly reported core terms are:
- Amount: $1.5 billion — the largest copyright payment to date in the AI context.
- Scope: roughly 500,000 works registered with the U.S. Copyright Office and meeting defined criteria.
- Per-work payment: about $3,000, shared among the rights holders in a work — such as author and publisher — under agreed allocation rules.
- Destruction: Anthropic destroys the copies it obtained from pirate sources.
- Limited release: the settlement covers past acquisition and use only; questions such as whether model outputs will infringe in future remain open.
What the settlement is not
The most common misreading of this outcome is that it sets binding "precedent." In fact, a settlement means the case was never taken to its end; the trial court's fair-use reasoning was never tested on appeal. In other words:
- Other courts may see it differently. In Kadrey, a case against Meta in the same district, another judge accepted that training was fair use on the record before him, but stressed that the risk of "market dilution" — AI output displacing the works it learned from — could prove decisive in future cases.
- Non-generative uses can come out the other way. In Thomson Reuters v. Ross, a Delaware ruling concerning a legal-research tool, using copyrighted content to build a competing product was held not to be fair use. The character of the use and its market effect are weighed afresh in every case.
- Courts outside the U.S. are not bound. In the United Kingdom, the High Court held in Getty Images v. Stability AI that model weights are not themselves an "infringing copy." In Germany, by contrast, the Munich Regional Court treated a model's "memorisation" of song lyrics as reproduction and ruled for the rights holder; we covered that line of reasoning in our report on the historic AI music copyright ruling in Germany.
- Major cases are still live. Suits such as The New York Times' case against OpenAI and Microsoft continue, with particular focus on outputs that reproduce copyrighted content.
That is why authors' groups greet the settlement with mixed feelings: a historic financial win, but one that leaves standing, at trial-court level, a "training is fair use" reading that many see as opening a door for the industry.
The real message for industry: the legal pedigree of data
The most practical lesson of the case is that AI compliance is shifting its focus from "what you do with the model" to "where and how you obtained the data." How a dataset was collected, under what licence it was acquired and how it is stored is now directly a billion-dollar risk item. We call this data provenance, and it covers questions such as:
- From which sources, on what date and on what legal basis was the training data acquired?
- Was pirated, leaked, or terms-of-service-violating scraped content filtered out?
- Can purchase, licence and destruction records be produced in an audit or in litigation?
- Were rights holders' machine-readable objections to text and data mining respected?
The requirement to destroy pirated copies underlines the point: unlawfully acquired data is a problem not only for damages but for its very right to continue existing. That reasoning may also lay the groundwork for future claims aimed at the model trained on unlawful data itself.
The European counterpart: transparency and the right to object
The European Union has chosen to manage this question through legislation rather than litigation. The 2019 Digital Single Market Directive grants a broad exception for text and data mining for scientific research, and a narrower one for other purposes that rights holders can override with a machine-readable opt-out. The EU AI Act, in turn, requires providers of general-purpose AI models to put in place a copyright-compliance policy that respects those opt-outs, and to publish a sufficiently detailed summary of the content used for training.
The contrast is striking. In the United States, companies train first and argue fair use in court later; in the EU, rights-holder objections and transparency operate up front. The Anthropic settlement confirms a point common to both worlds: how data was obtained sits at the centre of the legal analysis.
What it means for Turkey: training data without "fair use"
For Turkish lawyers, the most important lesson is that U.S.-style fair-use flexibility does not exist in Turkey. Law No. 5846 on Intellectual and Artistic Works provides a closed list of exceptions rather than a general, open-ended fair-use principle. Nor does Turkish law contain a dedicated text-and-data-mining exception of the kind found in the EU. Several consequences follow:
- Narrower room for manoeuvre. The transformative-use argument that won Anthropic the training issue may not find equivalent footing in Turkish law. Reproduction, whether temporary or permanent, generally requires the author's permission; training on copyrighted works without a licence or an express statutory basis carries significant risk.
- The acquisition distinction still matters. Copying from pirate sources can be treated as an infringement of the reproduction right whatever the purpose — and in Turkey criminal sanctions may accompany civil liability.
- The personal-data dimension. Books, articles and web content often contain personal data, so training data must be assessed under data-protection law as well as copyright; we examine this in training data: the KVKK and GDPR questions to settle first.
- Ownership of outputs is a separate question. Lawful training does not settle who owns what the model produces; we explore that in who owns the copyright of AI-generated images?
- Cross-border data flows. Training models developed in Turkey on datasets abroad raises a multi-layered compliance question spanning copyright and data transfers, discussed in training AI across borders.
A practical roadmap for companies
- Build a data inventory. Document the training-data sources of the models you develop or fine-tune. If you rely on third-party models, obtain written assurances from the provider about data provenance.
- Rewrite procurement contracts. Indemnification clauses protecting you against copyright claims, together with representations about data sources, have become critical in AI supply contracts.
- Move toward licensed data. Licensing deals with publishers and content owners have spread rapidly since this case. They may look costly, but they are a measured investment set against litigation risk.
- Screen for pirated content. Detect and remove content from known pirate sources in your own datasets, and keep a record of having done so.
- Honour opt-outs. If you scrape the web, document that you respect rights holders' technical objection signals; for any company entering the EU market this is now a de facto requirement.
- Manage the output side. Filters that stop a model reproducing copyrighted content verbatim, together with sound user terms, are an important defensive layer as litigation shifts toward outputs.
What authors and publishers in Turkey can do now
The settlement applies to works registered in the United States and to one company's conduct, but its logic gives rights holders everywhere practical tools:
- Reserve your rights expressly. Add clear text-and-data-mining reservations to publishing contracts, website terms and machine-readable signals. Even though Turkish law has no TDM exception, an express reservation strengthens your position in foreign markets — above all in the EU, where opt-outs have legal effect.
- Gather evidence carefully. Where a model reproduces long, verbatim passages of a work, record the prompts, dates and outputs in a way that can be verified later. Evidence of memorisation was central to the German lyrics case.
- Know which law applies. Under Turkey's Law No. 5718 on Private International Law and Procedure, claims arising from intellectual-property rights are governed by the law of the country where protection is sought. A Turkish author may therefore rely on different national laws against the same model depending on where the infringement occurs.
- Act collectively. Collecting societies and publishers' associations are better placed than individual authors to monitor use, negotiate licences and bring claims. The Anthropic class action showed how much leverage collective action creates.
- Treat licensing as an opportunity. Curated, high-quality Turkish-language content is scarce and valuable for model developers. Rights holders who organise their catalogues and metadata can negotiate from a position of strength rather than waiting for litigation.
What to watch next
Three themes will be decisive. First, the outcome of the major output-focused cases: that training may be fair use does not mean a model is free to reproduce copyrighted content. Second, the maturing of the licensing market: by supplying a de facto per-work value, this settlement will directly shape licensing negotiations with publishers. Third, legislative moves: the enforcement of the EU's transparency obligations and other countries' choices on text-and-data-mining exceptions will define the data strategies of global companies.
Frequently asked questions
Does this mean AI training is now entirely lawful?
No. The trial court held that training on lawfully acquired books was fair use on the facts before it, but that reasoning was never tested on appeal, other courts may reach different conclusions, and it binds no country outside the United States.
Could a non-U.S. author benefit from this settlement?
Only if their work fell within the class definition, which turned on criteria such as the work appearing in the pirated datasets and being registered with the U.S. Copyright Office. Nationality alone was not decisive; the settlement's own terms and its administrator determine eligibility and deadlines.
Does the settlement bind other AI companies?
No. A settlement binds only its parties. But its per-work figure will serve as a de facto reference point in similar cases and in licensing negotiations.
Could an author in Turkey bring a similar claim?
An author who alleges that their works were reproduced and adapted without permission can seek remedies under the Law on Intellectual and Artistic Works. The absence of a general fair-use defence in Turkish law may make the defendant's position harder than in the United States, though the less developed class-action mechanism produces a different picture in practice.
Is a company that merely uses an off-the-shelf model at risk?
Even without training anything itself, a company that integrates a model into its product can become part of the liability debate, particularly if outputs reproduce copyrighted content. That is why assurances in provider contracts matter.
Is $3,000 per work a lot or a little?
It is low against the theoretical ceiling of statutory damages and high against the price of a single book licence. Its real significance is that it provides the first large-scale market data point on what a work is worth for AI training.
Does destroying the pirated copies affect the trained model?
The publicly reported terms concern destroying files obtained from pirate sources. That does not mean deleting models already trained, but the question of a "model trained on unlawful data" remains open for future litigation.
How is data provenance documented in practice?
By keeping a file of data sources, acquisition dates, licence or purchase records, filtering and destruction logs, and records of compliance with opt-out signals. That file is a company's strongest defence in an audit and in any dispute.
Expert Opinion
This section reflects my personal assessment as the founder of this site and an AI ethics & compliance counsel.
In my view, the real importance of the Anthropic settlement is that it moves the AI-and-copyright debate from an abstract question — "is training allowed?" — to a concrete, auditable one: "how did you obtain the data?" The court's separation of training from acquisition is sound legal technique, because the legitimacy of the end does not cleanse the unlawfulness of the means. A model built on a pirate library stands on legally compromised foundations, however innovative it may be.
At the same time, reading the settlement as a "price tag" for the industry is dangerous. The per-work figure makes the cost of piracy measurable, but it also risks turning copyright infringement into a calculable operating expense. The purpose of law is not to price infringement but to prevent it. A sum that large companies can pay does not strike the same balance for small developers or independent creators.
My advice to companies in Turkey is clear: do not build strategy on U.S. fair-use debates. Our system is stricter and its exceptions narrower. Documenting the provenance of training data, moving toward licensed data, and rewriting supplier contracts around copyright risk are the most valuable steps available today. And to the creative sector I would add this: the settlement shows that works retain economic value in the age of AI, and that this value can be enforced in law. In the years ahead, AI compliance will be judged less by a model's performance than by the legal history of its data.
This article is for information only and does not constitute legal advice. The facts are based on public court decisions and press reporting on the case (e.g. TechCrunch, 20 July 2026); the analysis and assessments are the author's own.