Absorbing a writer's style by reading their books and training a language model on hundreds of thousands of those same books look similar on the surface — both draw something from someone else's text to produce something new. Yet one faces no legal constraints, while the other is being litigated simultaneously in multiple U.S. federal courts.

Since 2023, hundreds of American authors — including George R. R. Martin, Jodi Picoult, and John Grisham — have sued OpenAI, Meta, Anthropic, and others. Among these cases is a class action coordinated by the Authors Guild. The claim is consistent across suits: these companies used the authors' books to train AI without permission, and the resulting models now compete in the very markets those authors depend on. None of the cases has reached a final verdict yet.

Copyright Protects Expression, Not Ideas

These lawsuits aren't resolving quickly, and the reason lies in what copyright actually protects.

Copyright doesn't protect ideas. Anyone can write about "a lonely detective solving a case in an old mansion." What's protected is the specific expression — the particular sentences and descriptions — that brings that premise to life. The idea-expression distinction is one of the oldest principles in copyright law.

That principle cuts both ways in the AI training debate. AI companies argue that their models don't reproduce a given author's sentences but instead extract statistical patterns from billions of pieces of text — the same process, they say, by which a human absorbs style and structure through reading. Authors counter that extracting those patterns is itself an unauthorized appropriation of their work's value. Neither argument settles cleanly under existing case law.

Scale and Reproducibility Are Where the Two Acts Diverge

Put the two activities side by side, and the differences start to show.

Human reading vs. AI data trainingWhen a Person Learns from BooksMemory is imperfect and gets reshapedA few thousand books in a lifetimeTime and interpretation sit in betweenWhen AI Trains on BooksCan reproduce source text in some casesProcesses millions of documents at onceTraining data feeds directly into output

Start with reproducibility. If a person reads a book and then writes down its sentences verbatim, that's deliberate copying. AI models aren't designed to deliberately memorize their training data, but under certain conditions they can generate output that closely resembles it. Evidence filed in The New York Times's lawsuit against OpenAI includes instances where prompting the model in specific ways produced text nearly identical to the original articles.

Scale differs too. Even a voracious reader gets through, at most, a few thousand books in a lifetime. Books3, the dataset reportedly used to train Meta's LLaMA models, contained roughly 196,000 books — a substantial portion of which were confirmed to have been scraped from pirate book-sharing sites. The impact of one person's reading and the impact of a platform processing hundreds of thousands of books on the market for those authors' work are simply not calculated on the same scale.

Market substitution is the fourth factor in the U.S. fair use test. A person who reads a novel and is influenced by it to write a new one is unlikely to displace the market for the original. But if an AI trains on the style and narrative techniques of hundreds of thousands of authors and then mass-produces similar genre writing, the argument that this erodes a meaningful slice of those authors' market is one courts could credit. U.S. courts simply haven't settled that question yet.

Some AI Companies Are Signing Licensing Deals Ahead of Any Ruling

Some companies aren't waiting for a verdict. OpenAI has signed content licensing deals with News Corp (parent company of The Wall Street Journal), the Associated Press, and the Financial Times. The AP deal has been reported to be worth several million dollars annually, though most of these agreements keep their exact terms confidential.

Whether these deals stem from a sense of legal obligation or are simply business partnerships done at scale is open to interpretation either way. But the fact that a company begins paying for content while litigation over that very question is still pending reads, at minimum, as a signal that the company isn't fully confident in its legal position.

This isn't a distant concern for Korean authors and content directors. Korean publications with English translations or international distribution may well be included in these datasets. Text already public on the open web — blogs, newsletters, long-form columns — has often ended up in the training data of numerous models without separate consent, via open web datasets like Common Crawl.

What Options Authors Have Now, and Their Limits

Until the courts rule, the options genuinely available to authors are limited.

For web content, you can block major AI crawlers by adding rules to your robots.txt file. Identifiers for major crawlers — OpenAI's GPTBot, Common Crawl's CCBot, and others — are publicly documented, and following the technical instructions takes under 30 minutes to set up. The catch: this only prevents future scraping. It doesn't retroactively apply to data already collected.

Published books are a different story. There's still no official channel for excluding your book from a given model's training data. Some AI companies say they accept removal requests, but there's no external way to verify whether those requests are actually honored.

The more realistic path may be collective action — authors' organizations pooling their bargaining power to negotiate licensing terms. In the music industry, collective negotiations with streaming platforms took years to reshape how rights holders are protected, and the outcome remains imperfect. Still, it's a path more likely to produce results than any individual negotiating alone against a major corporation.

Knowing what copyright actually protects is what lets you judge what you can reasonably demand and what you can't. Ideas aren't protected; expression is. How AI processes that expression is a question the courts are still working out right now. Some companies have already run the numbers — and moved — before that standard was set.