top of page

Can Ai Learn From Copyrighted Books? The Copyright Questions Raised By Bartz V.

2 hours ago
7 min read

Introduction : Artificial intelligence (AI) is increasingly reliant on vast amounts of human generated content. Large language models (LLMs) are trained on data sets of books, articles and other written materials. This allows the AI to learn patterns of information and create responses, but it raises important copyright issues. When the AI learns from a copyrighted work, is that unlawful copying? This question was considered by the United States District Court for the Northern District of California in Bartz v. Anthropic. Three authors claimed that Anthropic had illegally copied their books while developing its AI systems. Anthropic had obtained books through multiple channels, including purchasing and digitising books as well as downloading them from online pirate libraries.


On 23 June 2025, the Court ruled that Anthropic's use of copyrighted books for training its LLMs were fair use. However, it differentiated between training and maintaining a permanent library of books obtained from pirate sources, which were not protected by fair use. The dispute eventually led to a $1.5 billion settlement concerning approximately 482,460 books obtained from pirate sources. The settlement did not create a general licensing requirement for future AI training or invalidate the Court's finding regarding training on lawfully acquired books. The significance of Bartz goes beyond the United States. In India, the Delhi High Court has considered a similar question in ANI Media Pvt. Ltd. v. Open AI Op Co LLC on the storage and use of copyrighted material for training Chat GPT.


How Does AI Learn from Copyrighted Works


A large language model requires extensive training on vast quantities of text. For use in a computer, that material typically has to be digitized first, potentially involving copying and processing copyrighted works. This raises copyright issues since copyright holders have an exclusive right to copy their works.


It might be used for a different purpose than the one for which it was created, a book is created to be read, while an AI company might copy the same work as part of a training set for a language model. This was the critical difference, according to Bartz.


The Bartz vs Anthropic Dispute


The Authors Andrea Bartz, Charles Graeber and Kirk Wallace Johnson Accused Anthropic of stealing their Copyrighted Books in the Development of its AI Systems Essay


Anthropic buys books in two different ways.


  • It buys printed books and then scans them and puts them in a searchable central library, from which it takes extracts for the purposes of training its AI.

  • It also buys millions of books from online pirate libraries such as LibGen and stores them in its central library.


The case therefore raises two different but related issues.


  • Can copyrighted books be copied and used for AI training?

  • Can an AI company buy and keep copyrighted books from unauthorized sources for AI use?


The Court treats the two issues differently.


Fair Use and AI Training


The main US legal issue was one of fair use under Section 107 of the US Copyright Act. The Court was considering whether the use of books for AI training was sufficiently different in purpose from their original purpose.


The Court found the use of the books for training highly transformative as the books were not used for their ordinary purpose of reading but as training material for creating large language models.


The Court nevertheless drew a distinction between the lawful purchase and digitisation of books and the acquisition of pirated copies. The judgment therefore did not establish that AI companies have an unfettered right to copy any copyright material found on the internet.


Lawfully Acquired Material: Where Anthropic purchased books and digitised them for training the Court found the use to be fair. The transformation of the books into training data served a purpose different from ordinary consumption.

Pirated Material: The position was different where Anthropic obtained books from pirate sources and retained them in its central library. The Court refused to extend fair-use protection to this conduct. A transformative purpose did not automatically justify unlawfully obtaining copyrighted works.


This is a significant distinction in that it shows that the purpose of using copyrighted material cannot necessarily cure the illegality of how that material was obtained.


The India Position : Can Section 52 Apply to AI Training ?


The Indian Position: Can The Indian position is different because the Copyright Act, 1957 works with particular fair-dealing exceptions, not like the United States with their broad use doctrine.


Section 52(1)(a) provides for fair dealings with a work for specific purposes like private or personal use, research, criticism or review, and reporting of current events amongst others. Therefore, a US finding of transformative nature of AI training does not automatically apply to an Indian copyright dispute. The Indian courts have to look at it with the parameters of Section 52.


ANI Media Pvt Ltd vs Open AI Opco Llc : The Indian Connection


On 24 July 2026, the Delhi High Court was asked to consider whether Open AI's storage and use of ANI's copyrighted news material for training Chat GPT constitutes copyright infringement. ANI claimed infringement and Open AI relied on Section 52(1)(a) and argued that the storage and use of ANI's works for training their language models would fall within the research-related exception. The Court was asked to consider if the storage of ANI's works for training of large language models could amount to “private or personal use, including research” under Section 52(1)(a) and whether AI training could fall within the provision. The Court noted that AI training could contribute to scientific knowledge, technological development and AI research and on a prima facie basis found that the requirement of purpose and fairness were met.


It therefore found that the storage of ANI's works for training purposes did not, on the material before it, amount to copyright infringement. However, the decision relates to interim relief and does not mean that every use of copyrighted material by AI systems is protected. The Court also separately considered questions regarding AI outputs and any possible memorisation or substantial reproduction of ANI's works and demonstrated that AI training and AI-generated outputs raise separate copyright questions.


Bartz and ANI : Two different Approaches to a similar problem


The significance of Bartz to India is not that it introduces US style fair use principles into Indian copyright law, but rather that it illustrates that different stages in the life cycle of AI may raise distinct copyright issues.


In Bartz the US court was considering whether scanning books for the purposes of training an AI system was a transformative use. By contrast, in ANI v. Open AI the Delhi High Court was considering whether the storage of copyrighted works for the purpose of training an LLM would be covered by the research exemption in Section 52(1)(a).


Though different in many respects, both cases raise the same fundamental question for copyright law: should the law be concerned with the purpose for which copyrighted material is copied, or is the mere fact of copying sufficient to engage in copyright infringement?


The Indian decision demonstrates that purpose and fairness are relevant factors in deciding whether a use is included in the exception in Section 52, although there are still some limitations placed on the section.


The Question of Pirated Training Data


The piracy issue in Bartz is an additional comparison point. An AI company making use of material lawfully available to them for a purpose within Section 52 is a different legal question to an AI company knowingly obtaining pirated copies and retaining them in a private database.


The Bartz settlement reinforces this practical importance as in the case, the $1.5 billion settlement concerned the pirated books and did not establish whether future AI training using lawfully acquired material would automatically require a licence.


For India, this opens an issue of their own, can an AI company invoke Section 52 if the underlying training material itself was obtained through an infringing source? Bartz does not have the answer as Indian courts would have to independently apply Indian copyright law.


Unresolved Questions


Despite these developments, several questions remain open.


  • First, it remains to be seen whether the reasoning regarding AI training would apply equally to different categories of copyrighted works, including fictional books, academic works and news articles.

  • Second, market harm is crucial; if AI-generated content begins to replace the works that human authors create, then courts examining potential copyright infringement will need to consider how the economic interests of copyright owners should be evaluated.

  • Third, the question of the treatment of pirated training datasets is important. Bartz shows what the possible consequences of retaining such a repository of unlawfully acquired works could be.

  • Finally, AI training and AI outputs must be distinguished; even if training is protected in a given case, a particular AI output may independently raise copyright concerns.


Conclusion


Bartz v. Anthropic delivers an instructive copyright message on the dangers of AI training that is not solely defined by the mere establishment of the fact that copyrighted works have been copied. The purpose of copying and mode in which it was obtained can drastically change the legal analysis The Court found Anthropic's use of books for the purposes of AI training to be a fair use but rejected the extension of the same logic to its acquisition and retention of pirated books. The matter is more differentiated for India as the Copyright Act provides specific fair dealing exceptions. The Delhi High Court's decision in ANI v. OpenAI puts forth that, in appropriate circumstances, Section 52(1)(a) can account for the storing of copyright works for the purpose of AI research and training.


Ultimately, the question is not about whether AI can learn from copyrighted works. Rather, it is if copyright law can be differentiated between legitimate technological research, the necessary copying of works to train AI and the unlawful appropriation of copyrighted works.


Author: Rida Fathima in case of any queries please contact/write back to us via email to content@khuranaandkhurana.com or at  Khurana & Khurana, Advocates and IP Attorney


References


  1. Bartz et al. v. Anthropic PBC, No. C 24-05417 WHA, U.S. District Court, Northern District of California, Order on Fair Use, 23 June 2025.

  2. Dave Hansen, “The Bartz v. Anthropic Settlement: Understanding America's Largest Copyright Settlement,” Kluwer Copyright Blog, 10 November 2025.

  3. Hadar Y. Jabotinsky & Michal Lavi, “AI, Fair Use, and Legal Risk: Governance Implications of Bartz v. Anthropic,” Oxford Business Law Blog, 4 November 2025.

  4. The Copyright Act, 1957, Section 52, Government of India.

  5. ANI Media Pvt. Ltd. v. OpenAI OpCo LLC, CS(COMM) 1028/2024, Delhi High Court, Judgment dated 24 July 2026.

  6. Bartz v. Anthropic Settlement, Yale University Press.

  7. Authors Guild v. Google, Inc., 804 F.3d 202 (2d Cir. 2015).

Comments


bottom of page