{"id":259170,"date":"2026-08-23T15:00:00","date_gmt":"2026-08-23T15:00:00","guid":{"rendered":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/08\/23\/is-it-legal-to-train-ai-models-on-copyrighted-books-its-complicated-techcrunch\/"},"modified":"2026-08-23T15:00:00","modified_gmt":"2026-08-23T15:00:00","slug":"is-it-legal-to-train-ai-models-on-copyrighted-books-its-complicated-techcrunch","status":"publish","type":"post","link":"https:\/\/entertainment.runfyers.com\/index.php\/2026\/08\/23\/is-it-legal-to-train-ai-models-on-copyrighted-books-its-complicated-techcrunch\/","title":{"rendered":"Is it legal to train AI models on copyrighted books? It\u2019s complicated | TechCrunch"},"content":{"rendered":"<p> <br \/>\n<\/p>\n<div>\n<p id=\"speakable-summary\" class=\"wp-block-paragraph\">You probably know by now that the AI models powering ChatGPT, Gemini, Claude, and other chatbots are trained on seemingly infinite databases of published works, containing hundreds of millions of books, online articles, academic papers, and basically anything you can find on the internet. Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods. That seems illegal, right? <\/p>\n<p class=\"wp-block-paragraph\">The reality isn\u2019t that simple.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cI think one of the issues with this entire area of law and this entire area of technology is there\u2019s a lot going on,\u201d Cathy Gellis, an attorney with expertise in intellectual property, copyright, and technology, told TechCrunch. \u201cIt\u2019s very complex and there are a lot of raw feelings about what is happening, both for and against.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Last year, in one of the first rulings of its kind, Judge William Alsup ordered Anthropic to pay a mammoth <a href=\"https:\/\/techcrunch.com\/2025\/09\/05\/screw-the-money-anthropics-1-5b-copyright-settlement-sucks-for-writers\/\" target=\"_blank\" rel=\"noopener\">$1.5 billion copyright settlement<\/a> to a group of writers whose works were used to train the company\u2019s AI models. At face value, this seemed like a moral victory favoring authors, but Judge Alsup actually ruled that Anthropic\u2019s AI training was lawful. What Alsup penalized Anthropic for was pirating these books from illegal online shadow libraries.<\/p>\n<p class=\"wp-block-paragraph\">\u201cLike any reader aspiring to be a writer, Anthropic\u2019s LLMs trained upon works not to race ahead and replicate or supplant them \u2014 but to turn a hard corner and create something different,\u201d the judge wrote, comparing the way an LLM ingests trillions of words to a writer\u2019s study of literature.<\/p>\n<p class=\"wp-block-paragraph\">Gellis thinks the ruling is more advantageous for AI companies. What\u2019s a $1.5 billion fine to a company projecting about <a rel=\"nofollow noopener\" href=\"https:\/\/www.reuters.com\/business\/anthropic-ipo-valuation-hinges-190-200-billion-2028-revenue-forecast-sources-say-2026-08-15\/\" target=\"_blank\">$200 billion<\/a> in annual revenue by 2028?<\/p>\n<p class=\"wp-block-paragraph\">\u201cI think it is generally good news for AI training that he looked at what was going on and really sort of thought it analogous to reading a copyrighted work as opposed to copying a copyrighted work,\u201d Gellis said. \u201cCopyright law hinges on copying, but it doesn\u2019t hinge on using the work or experiencing the work, consuming the work, reading the work.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Copyright law <a rel=\"nofollow noopener\" href=\"https:\/\/en.wikipedia.org\/wiki\/Copyright_Act_of_1976\" target=\"_blank\">hasn\u2019t been updated<\/a> since 1976, which means that judges have to figure out how to interpret guidelines from 50 years ago when confronting legal questions that have the potential to shape the future of the AI industry.<\/p>\n<p class=\"wp-block-paragraph\">\u201cEverybody is very worried right now because the law is all over the place, and it\u2019s because of this question,\u201d Jason Henderson, Senior Attorney and Founder of the IP &amp; Media Practice at JWL International, told TechCrunch. \u201cThey know that the AI model has been trained on so much stuff, and the law has not really caught up to that question.\u201d<\/p>\n<p class=\"wp-block-paragraph\">These questions often hinge on fair use law \u2014 namely, whether use of a copyrighted work is \u201ctransformative\u201d enough to be considered legally permissible.<\/p>\n<p class=\"wp-block-paragraph\">Fair use is a carve out of copyright law that allows for the use of copyrighted materials without explicit permission, protecting the ability to comment and iterate on copyrighted works through criticism, parody, education, and other means. Judges consider specific factors when deciding if something is fair use, including the purpose and nature of the work, the amount used, and its impact on the market.<\/p>\n<p class=\"wp-block-paragraph\">\u201cCopyright is always about protecting and growing the market,\u201d Henderson noted. \u201cThe courts are kind of all over the place in their reasoning [in AI cases]. What\u2019s tending to win is if what you\u2019re doing is you\u2019re training on somebody\u2019s property because your purpose is to directly compete, then the courts will frown on it\u2026 If what you\u2019re doing is not going to compete, then the courts are tending to find ways that it will be okay.\u201d\u00a0<\/p>\n<p class=\"wp-block-paragraph\">Henderson is referencing a case in which the media and technology company Thomson Reuters sued the research firm Ross Intelligence for copying its content in order to build a competing, AI-based legal platform.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">\u201cRoss\u2019s use is not transformative because it does not have a \u2018further purpose or different character\u2019 than Thomson Reuters\u2019s,\u201d Judge Stephanos Bibas <a rel=\"nofollow noopener\" href=\"https:\/\/fingfx.thomsonreuters.com\/gfx\/legaldocs\/xmvjbznbkvr\/THOMSON%20REUTERS%20ROSS%20LAWSUIT%20fair%20use.pdf\" target=\"_blank\">wrote<\/a> last year.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In that case, Judge Bibas decided that it was not fair use to train on Reuters\u2019 content to make a new platform that would directly compete with it. While authors could potentially argue that chatbots are competing with them by using their works to generate new, synthetic books, that argument has not yet prevailed in court.<\/p>\n<p class=\"wp-block-paragraph\">When it comes to the relationship between AI and copyright, Gellis finds it helpful to narrow down what we\u2019re actually talking about \u2013 the way we think about copyright in terms of AI training is quite different from how we think about copyrighting AI-generated content.\u00a0<\/p>\n<p class=\"wp-block-paragraph\">In one case, <a rel=\"nofollow noopener\" href=\"https:\/\/law.justia.com\/cases\/federal\/appellate-courts\/cadc\/23-5233\/23-5233-2025-03-18.html\" target=\"_blank\">Thaler v. Perlmutter,<\/a> the court ruled that if a work is 100% AI-generated, it\u2019s not copyrightable, which opens a whole new can of worms \u2013 how can we definitively prove whether or not a work was generated using AI, and if so, how do we know what percentage of it was created or assisted with AI?<\/p>\n<p class=\"wp-block-paragraph\">\u201cIf you write your novel in [Microsoft] Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel,\u201d Gellis said. \u201c[AI] is forcing us to look at a whole bunch of decisions that we kind of ignored for a while.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Most AI companies are still lodged in pending litigation over these issues, which means that we won\u2019t have a definitive solution to these problems any time soon. <\/p>\n<p class=\"wp-block-paragraph\">\u201cWhat you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things, and it\u2019ll take later states of litigation to figure out which one will prevail,\u201d Gellis said. \u201cBut in the meantime, all these decisions are shaping everything that\u2019s happening. It would be kind of foolish for the AI companies to ignore them.\u201d<\/p>\n<\/div>\n<p><em>When you purchase through links in our articles, <a href=\"https:\/\/techcrunch.com\/techcrunch-affiliate-monetization-standards\/\" target=\"_blank\" rel=\"noopener\">we may earn a small commission<\/a>. This doesn\u2019t affect our editorial independence.<\/em><\/p>\n<p><br \/>\n<br \/><a href=\"https:\/\/techcrunch.com\/2026\/08\/23\/is-it-legal-to-train-ai-models-on-copyrighted-books-its-complicated\/\" target=\"_blank\" rel=\"noopener\">Source link <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>You probably know by now that the AI models powering ChatGPT, Gemini, Claude, and other chatbots are trained on seemingly infinite databases of published works, containing hundreds of millions of books, online articles, academic papers, and basically anything you can find on the internet. Most published authors have, without their knowledge or consent, contributed to [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":259171,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[14],"tags":[],"class_list":{"0":"post-259170","1":"post","2":"type-post","3":"status-publish","4":"format-standard","5":"has-post-thumbnail","7":"category-tech"},"_links":{"self":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/259170","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/comments?post=259170"}],"version-history":[{"count":0,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/posts\/259170\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media\/259171"}],"wp:attachment":[{"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/media?parent=259170"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/categories?post=259170"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/entertainment.runfyers.com\/index.php\/wp-json\/wp\/v2\/tags?post=259170"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}