Explainer

Your Words Trained the Machine. Where's Your Cut?

Every AI model is built on a mountain of human work — our posts, reviews, photos, code, conversations — gathered mostly without asking and never paid for. The quiet expropriation behind the magic.

Somewhere in the guts of every large AI model sits a version of you. Not your name or your face, but the shape of how you write, the review you left on a restaurant, the answer you posted on a forum at 2am to help a stranger fix their code, the photo you uploaded of your dog. This is AI training data: the collective output of billions of ordinary people, scraped mostly without consent and almost always without payment, then fed into machines that now generate text, images and code for profit. The models did not learn to write from nowhere. They learned from us. And that raises a question the industry would rather you didn’t dwell on — your words trained the machine, so where is your cut?

I want to be careful here, because this subject attracts a lot of heat and not much light. This isn’t a conspiracy. Nobody broke into your house. What happened is more mundane and, in a way, more unsettling: a shared commons of human expression, built over three decades of people posting online for free, was reclassified as raw material — and then as a private asset worth hundreds of billions. This piece is about how that happened, why it’s genuinely hard to fix, and what a fairer arrangement might look like.

What training data actually is

An AI language model is, at heart, a machine for predicting patterns. To make one that writes like a human, you feed it an enormous quantity of human writing and let it absorb the statistical regularities — which words tend to follow which, how an argument is structured, what a recipe looks like versus a legal contract. To make one that paints, you feed it images with their descriptions. The capability of the model is downstream of the data: better and broader data, better model. This is why the race for AI has quietly become a race for data.

And where did that data come from? Overwhelmingly, from the open internet. The most capable systems were trained on colossal scrapes of the web — encyclopaedias, news archives, forums, blog posts, product reviews, question-and-answer sites, code repositories, image libraries, digitised books. Broad web scraping at this scale hoovers up whatever is reachable. Some of it was published deliberately by professionals. Most of it was the incidental exhaust of normal life online: the comment you left, the Wikipedia edit, the Stack Overflow answer, the recipe blog, the fan wiki. Billions of small, uncoordinated acts of human expression, gathered into a corpus no licensing department could ever have negotiated for in advance.

That last point matters, because it’s the fact the whole industry is built on. The scale is what makes the data valuable and what makes paying for it look impossible.

The sleight of hand

Here is the move I want you to see clearly. When you posted a review or answered a question, you were contributing to a commons — a shared pool of human knowledge that anyone could read and benefit from. That was the implicit deal of the open web, and it was a good deal. The value was diffuse and public.

Then that same pool got fed into a handful of models owned by a handful of companies, and the value stopped being diffuse. It concentrated. The commons everyone built together became the private foundation of a commercial product that a few firms now charge rent on. Nothing was formally stolen — much of it was technically public — but the character of the thing changed completely. The people who supplied the raw material were not partners in the enterprise. They were, at best, an unpaid input, and mostly they didn’t know they were an input at all.

Shared contribution went in; private asset came out. The people who supplied the raw material were not partners in the enterprise — they were an unpaid input, and mostly they didn’t know they were an input at all.

If that pattern feels familiar, it should. It is the same move I keep returning to whenever I look at how technology gets captured: a resource that belongs to everyone gets fenced off and turned into something a few people own. It’s the logic of surveillance capitalism, where your behaviour became a raw material for prediction; and it runs straight into the harder question of who owns AI in the first place. The raw material is always us. The ownership is always somewhere else.

The honest complexity

Now, if this were simply theft, the answer would be easy: pay people back. It isn’t simple. There are at least three genuine complications, and any fair solution has to survive them.

First, the individual contribution is tiny and the micro-payment is impractical. A single comment or product review contributed an almost immeasurably small amount to a model trained on trillions of words. If you tried to trace value back to each person and cut them a cheque, it would be worth a fraction of a penny and the accounting would cost more than the payment. This is the industry’s favourite defence, and it isn’t wrong on the arithmetic. The trouble is that it slides from “we can’t pay each of you individually” to “therefore we owe nothing to any of you collectively,” which does not follow at all.

Second, not all data is the same. A professional novelist’s book, a journalist’s investigation or a photographer’s portfolio is deliberate, skilled, often commercially licensed work — and copying it to train a commercial model is the subject of a serious and unresolved legal fight. A casual tweet is different in kind. Lumping the two together muddies the argument. The novelist has a claim rooted in copyright law; the person who left a helpful forum answer has something more like a moral claim than a legal one. Both matter, but they are not the same claim, and I dig into the legal side — including the messy question of who owns what AI makes from that scraped work — in a companion piece.

Third, learning from prior work is not inherently wrong. Every human writer learned to write by reading other writers without paying them. The AI companies lean hard on this analogy — the model, they say, learns patterns the way a person does. There’s something to it. The counter-argument is about scale and effect: a person who reads a thousand novels and writes their own is not the same as a machine that ingests millions and can then flood the market with substitutes at zero marginal cost, competing with the very people it learned from. Analogy is not identity. But the honest position is that “it’s just learning” is neither obviously true nor obviously false — it’s the thing being argued about.

Hold those three together and you get the real situation: a genuine wrong (uncompensated value taken at scale) tangled up with genuine difficulty (you can’t easily price or refund it). That tension is why the responses on offer are all partial.

The range of proposed fixes — and their limits

A licensing-and-lawsuit landscape is now emerging, and alongside it a set of proposed remedies. None is complete.

  • Licensing deals. The most active response: AI firms signing agreements to pay large platforms, news organisations and archives for access to their data. This is real money changing hands, and it’s progress of a sort. But notice who gets paid — the platform that hosts your content, not you who made it. When a forum or a photo library licenses its corpus, the cheque goes to the company that holds the data, and the actual authors typically see none of it. Licensing can settle the fight between corporations while leaving the individual where they started.
  • Collective data funds. The idea that AI firms pay into a shared pool — a levy on models trained on public data — which is then distributed to creators, or to the public, or reinvested in things the commons needs. This has the right shape, because it matches a collective harm with a collective remedy instead of chasing impossible micro-payments. The hard parts are governance and honesty: who administers the fund, who counts as a contributor, and whether it becomes a genuine redistribution or just a small fee that buys permanent permission cheaply.
  • Opt-outs. Mechanisms letting you tell crawlers not to train on your site or your work. Useful, and better than nothing, but structurally weak. Opt-out puts the entire burden on the individual, assumes you even know your data is being taken, and typically arrives after the models are already trained. A right you have to discover and manually enforce, one site at a time, is not much of a right.
  • Data dignity. The most ambitious idea — associated with thinkers like Jaron Lanier — that data should be treated as labour, with people recognised as the producers of the value AI extracts and entitled to an ongoing stake in it. As a principle it reframes the debate correctly: this is uncompensated labour, not free-floating material. As a mechanism it’s still largely aspirational, running into the same measurement problem as everything else.

The pattern across all four is telling. The remedies that are easy to implement (opt-outs, platform licensing) do the least for ordinary people, and the ones that would do the most (collective funds, data dignity) are the hardest to build. That is not an accident. It’s what capture looks like when it’s working: the cheapest concessions get made, the structural ones get called impractical.

The enclosure lens

Step back far enough and the newest technology of our age is running an old script. Centuries ago, land that villages had farmed and grazed in common was fenced off and turned into private property, and the people who had depended on it were left with nothing but their labour to sell. That was the enclosure movement, and it is the cleanest analogy we have for what is happening to human expression now.

The open web was a commons. Not a perfectly fair one, but a genuinely shared resource that people contributed to without expecting to be paid, because the benefit flowed back to everyone. AI training is the fence going up around that commons. The shared pool of human writing and image-making is being enclosed, converted into a private input, and monetised by whoever owns the model. The commoners — us — still get to use the land, in that we can use the AI tools. We just don’t own any part of what was built from our work.

The commons everyone built is being fenced off, converted into a private input, and monetised by whoever owns the model. We rent access to capabilities distilled from our own contributions.

Enclosure is not a story of villains twirling moustaches. It’s about who has the power to redraw the boundaries of ownership, and in whose favour. The uncomfortable truth is that the boundary has already moved. The models are trained. The question is no longer whether to prevent the enclosure, but whether the people whose expression was enclosed get any lasting stake in what it produces.

What fairer would look like

I don’t think the answer is to unwind AI, and I don’t think individual micro-payments are coming. But “fair” is not the same as “impossible.”

It would mean treating training data as a collective asset with a collective return — a levy or a fund, transparently governed, that channels value back to the public that supplied it, as direct distributions, public investment, or shared ownership of the resulting infrastructure. It would mean licensing deals that pass real value through to the actual authors, not just the platforms sitting on their content. It would mean opt-in and provenance built in by default, so consent is the starting condition rather than a right you have to hunt down. And it would mean naming what the data is: not free raw material, but the accumulated, uncompensated labour of billions of people, which entitles them to more than a thank-you.

The through-line I keep coming back to is simple. Somebody takes, somebody pays, and eventually somebody fights back. Right now the taking is nearly complete and the paying has barely begun. The models learned to speak in our voice, built on work we gave away before anyone told us it was worth a fortune. Asking where our cut is isn’t entitlement. It’s the first honest question of the AI economy — and the answer we settle on will tell us who this technology is really for.

Kenney Jacob is the author of Captured, a history of who takes, who pays, and who fights back.

Frequently asked questions

What is AI training data?

The vast body of text, images, code, audio and video used to teach an AI model. Much of it was scraped from the open web and platforms — the collective output of billions of people — usually without their explicit consent or any payment.

Was my data used to train AI?

If you've posted online — reviews, comments, photos, code, forum answers — it very likely was, directly or indirectly. Most large models were trained on broad web scrapes that swept up ordinary people's contributions alongside published work.

Should people be paid for AI training data?

It's a live and contested question. Some argue for licensing, royalties or collective data funds; others say individual micro-payments are impractical and the fix must be collective. Either way, the current default — take everything, pay no one — concentrates the value with whoever owns the model.

← All articles