Explainer

NYT v OpenAI: The Case That Could Decide Who Owns AI's Training Data

A newspaper sued the most valuable AI company in the world over the oldest question there is: if you built your fortune out of my work, do you owe me anything? However it ends, the answer will set the terms for everyone whose words, images or code went into the machine.

The fight over NYT vs OpenAI is the closest thing this decade has to a constitutional question about machines: when a company reads everything you have ever written and builds a product that answers the questions your writing used to answer, what exactly did it take from you? The New York Times sued OpenAI and Microsoft in December 2023, in the Southern District of New York, alleging that millions of its articles were copied at scale to train large language models without a licence, and that those models can spit back its journalism in near-verbatim form. OpenAI says this is fair use. As of this writing, no one has won.

That last sentence matters more than the headlines suggest. Nearly three years in, the case is at summary judgment before Judge Sidney Stein, who denied most of OpenAI’s motion to dismiss back in March 2025 and let the core copyright claims proceed while trimming some of the DMCA claims. Related suits from the New York Daily News and the Center for Investigative Reporting have been folded in alongside author claims. In early September 2026, as reported, both sides filed duelling summary judgment briefs, and the US Department of Justice filed a brief broadly supporting the position that training models on copyrighted text is generally fair use. A ruling is expected in the coming months. Anyone telling you they know how it ends is selling something.

What the Times is actually alleging

Strip away the press-release language and the complaint makes two distinct claims that are worth separating, because they fail or succeed for different reasons.

The first is about inputs. The Times says its archive — decades of reporting, much of it expensive, some of it dangerous to produce — was ingested wholesale into training corpora without permission or payment, and that its material was disproportionately valuable to the models precisely because it is clean, edited, factual prose. This is the same argument at the heart of nearly every case about the data AI is trained on: the corpus was not a natural resource lying around. Somebody made it.

The second is about outputs. The Times put exhibits in front of the court showing models reproducing long passages of its articles close to verbatim, and producing fabricated content falsely attributed to the paper. This is the regurgitation evidence, and it is the part people are too quick to wave away. If a system can be induced to emit a paywalled investigation nearly word for word, the claim that the model merely learned statistical patterns gets harder to hold with a straight face. Memorisation is not an abstraction; it is a copy sitting inside the weights, waiting for the right prompt.

The commercial theory ties the two together. The Times argues that a chatbot which summarises, paraphrases and occasionally reproduces its journalism is not a scholar quoting a source — it is a substitute product, competing with the paper for the same reader, built out of the paper’s own work.

Why OpenAI’s defence is not frivolous

It is tempting, if you write for a living, to treat the fair-use argument as a fig leaf. It isn’t, and pretending otherwise is how you lose an argument you should win.

OpenAI’s position rests on a few points that are genuinely strong. Facts are not copyrightable — the Times owns its expression, not the events it reports. Training is, on its face, a transformative use: the output of the process is a set of weights, not a library of articles. And OpenAI can point to real precedent. Two California federal decisions in 2025, Bartz v. Anthropic and Kadrey v. Meta, found that training on copyrighted books was transformative enough to qualify as fair use on the facts presented. The DOJ brief, as reported, goes further still, framing training as a public good with implications for scientific progress and competitiveness.

On regurgitation, OpenAI has argued the Times engineered its exhibits with adversarial prompting — essentially, that you have to work hard to make the model misbehave, and that memorised emission is a rare bug rather than a product feature. There is something to this. A system that can be coaxed into reciting an article after sustained effort is not the same as a system that hands it over on request.

The honest version of this dispute is not good versus evil. It is two serious readings of a law written for photocopiers, applied to a machine that learns.

The Times’ answer to the precedent problem is narrow and, to my eye, well aimed: it does not contest what those courts found, it says those plaintiffs never proved market substitution. Novelists could not show that a chatbot replaced their novels. A newspaper can at least try to show that an answer engine replaces the click, the subscription and the ad that funded the reporting. Fair use has always turned heavily on the fourth factor — the effect on the market for the original — and that is the ground the Times has chosen to fight on. It is the right ground.

Why a judgment matters more than a cheque

Here is the part that most coverage gets backwards. The biggest number in this field so far did not come from a court deciding anything. The Anthropic author settlement — roughly $1.5 billion, finally approved in 2026, covering around half a million pirated books at about $3,000 each — was widely reported as a landmark. It was, in one narrow sense: it is the largest known copyright recovery on record, and it put real money in the hands of real authors.

But a settlement decides nothing. It is a price, paid to make a question go away. Its structure is telling: the class released claims about acquisition and copying — the inputs side — and left the harder question of outputs largely untouched. No rule was written. No future defendant is bound. What the industry learned from it is not “don’t do this” but “this costs about this much,” and for a company raising at hundreds of billions, that is a line item, not a deterrent.

A litigated judgment is a different animal. It produces a rule that binds people who were never in the room. It tells every model builder, and every writer, photographer, musician and programmer whose work sits in a scrape, what the default is. And crucially, it is a rule arrived at in public, on a record, with evidence tested by an adversary — not a number negotiated privately between two parties who both have reasons to keep the reasoning secret.

That distinction is the whole reason to care whether this case is tried rather than bought. Settlement is how a well-capitalised industry converts a question of principle into an operating expense. It is the most reliable mechanism by which a rule that might have constrained the powerful becomes, instead, a fee the powerful can afford and nobody else can charge.

The discovery fight nobody expected

One sub-plot deserves attention because it shows how strange the terrain is. In the consolidated litigation, the plaintiffs sought a sample of ChatGPT conversation logs to test how often models actually reproduce their work. OpenAI initially proposed a 20-million-log sample, then later tried to narrow production to conversations surfaced by keyword search. Magistrate Judge Ona Wang rejected that narrowing in late 2025, and Judge Stein affirmed the order compelling the full sample.

OpenAI framed its resistance as a privacy defence for its users, which is not a cynical framing — those logs contain other people’s confessions, medical questions and business plans. But notice the shape of the problem. To find out whether a machine is copying journalism, a court has to reach into a private record of what millions of people said to it. The evidence of infringement and the intimate data of the public now live in the same box. That is not an accident of this case; it is what happens when a single product becomes both the library and the reading room.

What this means if your work is in the pile

If you have ever published an essay, posted code, uploaded a photograph or written documentation, you are almost certainly in a training set. You will never be notified, you cannot audit it, and in most cases you cannot get out. So what does the outcome actually change for you?

  • If training is held to be fair use across the board, the practical answer is that your published work is free input for anyone with enough GPUs. Licensing deals will still happen — large publishers have leverage and lawyers — but they become commercial courtesies, not obligations. Individual creators get nothing, because nobody negotiates with one person.
  • If the Times wins on market substitution, the rule is narrower than the triumphalism would suggest. It would likely mean training is broadly permissible but becomes infringing when the output competes directly with the source. That is a rule big rights-holders can enforce and a freelance illustrator cannot.
  • If the case settles, we learn the price of The New York Times and nothing else.

Notice that two of those three outcomes concentrate power. This is the pattern worth naming, and it is exactly how technology gets captured: a genuinely open question gets resolved in a forum where only the largest participants can afford to show up, and the settlement — whether it is a cheque or a judgment — hardens into an arrangement between incumbents. The individual writer is present in this case only as a statistic in someone else’s corpus.

The question was never whether machines may read. It is who collects the rent on what the reading built.

The question underneath

Copyright is a poor instrument for what is actually being contested here, which is why the arguments feel slightly off-key on both sides. Copyright asks whether a copy was made. The real dispute is about value: an enormous amount of it has been created by assembling human work into a machine, and almost none of it is flowing back to the humans whose work made the assembly possible.

The same confusion runs through the mirror-image argument about who owns what AI makes. On the input side, the industry’s position is that using protected work is transformative enough to be free. On the output side, it would very much like the results to be ownable. You cannot comfortably hold both — that training is too abstract to be copying, and that generation is concrete enough to be property — without an account of why the line falls exactly where it is most profitable.

I do not think there is a clean answer, and I am suspicious of anyone who offers one. A rule that makes all training infringing would hand the future to whoever already owns the largest archives, which is not a victory for writers. A rule that makes all training free hands it to whoever already owns the largest compute, which is not a victory either. The interesting design space is in between — compulsory licensing, collective bargaining for creators, attribution that carries payment, transparency requirements about what was ingested — and almost none of that is available to a district judge deciding a fair-use motion. Courts answer the question in front of them. This one is much bigger than the question.

Which is why the judgment still matters, even though it cannot solve the problem. Whatever Judge Stein rules, as reported the losing side will appeal, and the rule that eventually settles will be made by people who had the resources to stay in the fight for a decade. The rest of us do not get a vote. We just supplied the training data.

Kenney Jacob is the author of Captured, a history of who takes, who pays, and who fights back.

Frequently asked questions

What is the NYT vs OpenAI case about?

The New York Times sued OpenAI (and Microsoft) alleging that its journalism was copied at scale to train large language models without licence or payment, and that the models can reproduce or closely paraphrase its articles — competing with the paper using its own work. OpenAI has argued its training constitutes fair use. The litigation has been lengthy and its status keeps moving, so check current reporting for where it stands.

Why does the case matter beyond the New York Times?

Because it goes at the foundational question for the whole industry: whether training a commercial model on copyrighted material without permission is fair use. A ruling either way would reshape what AI companies must licence and pay for, and would set the baseline for every writer, artist, musician and publisher whose work sits in a training set.

How is this different from the Anthropic author settlement?

A settlement resolves one dispute without establishing law — money changes hands and the underlying legal question stays open. A litigated judgment would produce a precedent courts and companies must follow. That is why the industry watches contested cases more closely than large settlements, however big the cheque.

← All articles