OpenAI has been accused by many parties of training its AI on copyrighted content sans permission. Now a new paper by an AI watchdog organization makes the serious accusation that the company increasingly relied on nonpublic books it didnโt license to train more sophisticated AI models.
AI models are essentially complex prediction engines. Trained on a lot of data โ books, movies, TV shows, and so on โ they learn patterns and novel ways to extrapolate from a simple prompt. When a model โwritesโ an essay on a Greek tragedy or โdrawsโ Ghibli-style images, itโs simply pulling from its vast knowledge to approximate. It isnโt arriving at anything new.
While a number of AI labs, including OpenAI, have begun embracing AI-generated data to train AI as they exhaust real-world sources (mainly the public web), few have eschewed real-world data entirely. Thatโs likely because training on purely synthetic data comes with risks, like worsening a modelโs performance.
The new paper, out of the AI Disclosures Project, a nonprofit co-founded in 2024 by media mogul Tim OโReilly and economist Ilan Strauss, draws the conclusion that OpenAI likely trained its GPT-4o model on paywalled books from OโReilly Media. (OโReilly is the CEO of OโReilly Media.)
In ChatGPT, GPT-4o is the default model. OโReilly doesnโt have a licensing agreement with OpenAI, the paper says.
โGPT-4o, OpenAIโs more recent and capable model, demonstrates strong recognition of paywalled OโReilly book contentย โฆ compared to OpenAIโs earlier model GPT-3.5 Turbo,โ wrote the co-authors of the paper. โIn contrast, GPT-3.5 Turbo shows greater relative recognition of publicly accessible OโReilly book samples.โ
The paper used a method called DE-COP, first introduced in an academic paper in 2024, designed to detect copyrighted content in language modelsโ training data. Also known as a โmembership inference attack,โ the method tests whether a model can reliably distinguish human-authored texts from paraphrased, AI-generated versions of the same text. If it can, it suggests that the model might have prior knowledge of the text from its training data.
The co-authors of the paper โ OโReilly, Strauss, and AI researcher Sruly Rosenblat โ say that they probed GPT-4o, GPT-3.5 Turbo, and other OpenAI modelsโ knowledge of OโReilly Media books published before and after their training cutoff dates. They used 13,962 paragraph excerpts from 34 OโReilly books to estimate the probability that a particular excerpt had been included in a modelโs training dataset.
According to the results of the paper, GPT-4o โrecognizedโ far more paywalled OโReilly book content than OpenAIโs older models, including GPT-3.5 Turbo. Thatโs even after accounting for potential confounding factors, the authors said, like improvements in newer modelsโ ability to figure out whether text was human-authored.
โGPT-4o [likely] recognizes, and so has prior knowledge of, many non-public OโReilly books published prior to its training cutoff date,โ wrote the co-authors.
It isnโt a smoking gun, the co-authors are careful to note. They acknowledge that their experimental method isnโt foolproof and that OpenAI mightโve collected the paywalled book excerpts from users copying and pasting it into ChatGPT.
Muddying the waters further, the co-authors didnโt evaluate OpenAIโs most recent collection of models, which includes GPT-4.5 and โreasoningโ models such as o3-mini and o1. Itโs possible that these models werenโt trained on paywalled OโReilly book data or were trained on a lesser amount than GPT-4o.
That being said, itโs no secret that OpenAI, which has advocated for looser restrictions around developing models using copyrighted data, has been seeking higher-quality training data for some time. The company has gone so far as to hire journalists to help fine-tune its modelsโ outputs. Thatโs a trend across the broader industry: AI companies recruiting experts in domains like science and physics to effectively have these experts feed their knowledge into AI systems.
It should be noted that OpenAI pays for at least some of its training data. The company has licensing deals in place with news publishers, social networks, stock media libraries, and others. OpenAI also offers opt-out mechanisms โ albeit imperfect ones โ that allow copyright owners to flag content theyโd prefer the company not use for training purposes.
Still, as OpenAI battles several suits over its training data practices and treatment of copyright law in U.S. courts, the OโReilly paper isnโt the most flattering look.
OpenAI didnโt respond to a request for comment.


