News-jack
Who Owns What AI Makes From Your Work? Inside the Copyright Fight
The machine learned to write, draw and compose by reading almost everything we ever made — most of it without asking. The lawsuits now deciding who owns the output are really deciding an older question: who gets the gains.
Every time an AI copyright lawsuit lands in the news, the same uncomfortable question sits underneath it: who owns what an AI makes from your work? The large language and image models that now write, draw, code and compose were trained on the writing, art, code and music of millions of people — most of whom were never asked, never credited, and never paid. That single fact is the root of the biggest fight in creative industries right now. It splits into two separate legal questions that often get tangled together, and it points at something older and larger than the law: a familiar pattern of who takes, who pays, and who eventually fights back.
Let me separate the two questions first, because almost every confused argument online comes from mixing them up. Question one is about the input: was it lawful to train a model on copyrighted work scraped from the open web without permission? Question two is about the output: when the model produces a poem, an image or a function, can that be copyrighted at all — and if so, by whom? The courts are answering these very differently, and neither answer is settled.
The input fight: training on work nobody licensed
Start with the training data, because that is where the money and the anger are. To build a model that writes like a human, you feed it an enormous corpus of human writing. To build one that paints, you feed it human paintings. The most capable systems were trained on colossal scrapes of the internet — books, news archives, forums, code repositories, stock-photo libraries, artist portfolios — gathered at a scale no licensing department could ever have negotiated in advance. The companies building these models largely treated the open web as free raw material.
Creators and rights-holders disagree, and a wave of copyright lawsuits has followed. Authors, visual artists, musicians, news organisations, stock-image companies and software developers have all sued major AI firms, arguing that copying their work to train a commercial model is straightforward infringement. The AI companies mostly lean on a fair use defence (or its equivalents in other legal systems): they argue that training is transformative — that a model learns statistical patterns rather than storing and reselling the originals, much as a person learns to write by reading. Rights-holders counter that the output competes directly with the very people whose work was taken, and that “the machine only learned from it” is a convenient way to describe industrial-scale copying.
Where does this stand in 2025? Honestly, unsettled. Courts in different countries are still deciding, and early rulings have pointed in more than one direction — some sympathetic to the transformative-use argument for the act of training itself, others far more hostile to how the data was acquired, especially where pirated copies were involved. Some disputes have been settled quietly, and a few reported settlements have involved large sums, but there is no single tidy precedent that tells you “training is legal” or “training is theft.” Anyone claiming certainty is selling something. What we can say is that the direction of travel is toward licensing: more AI firms are now signing paid deals with publishers and archives, which is itself a quiet admission that the free-scraping era was legally shaky.
The most capable systems were built on a scrape of human creativity that no licensing department could ever have negotiated in advance — and that scale is exactly what is now being argued over.
It’s worth being fair to the other side here, because the creative industries do not have a monopoly on the public interest. Requiring a signed licence for every sentence and sketch a model ever saw would, in practice, hand the entire future of AI to whichever handful of companies could afford to buy the whole internet — which happens to be the same handful that can already afford it. A strict permission regime could entrench the incumbents rather than protect the small creator. There is a genuine tension between rewarding individual authorship and keeping powerful tools from becoming the private property of a few. Both of those are real values, and the honest position is that they are in conflict.
The output fight: can a machine’s work be owned?
Now flip to the other end. Suppose the model exists and you prompt it. Who owns the result?
Here the law has been clearer, at least in outline, and the 2025-era direction in several jurisdictions has held reasonably firm: copyright protects human authorship. A work generated purely by a machine, with no meaningful human creative control over the specific expression, has generally been treated as not copyrightable — it falls into a kind of public domain by default. Registration offices and courts have repeatedly rejected attempts to claim copyright in “art made autonomously by an AI,” and guidance has leaned toward asking how much a human actually shaped the final expression rather than just pressing go.
The interesting cases are in the middle. If you write a detailed prompt, generate a hundred variations, then select, edit, arrange and rework the output into a finished piece, how much of you is in there? The emerging answer — still hedged, still evolving — is that the human-authored parts (your selection, arrangement, edits) may be protectable, while the raw machine-generated material underneath may not be. That is a messy, case-by-case line, and it will be litigated for years. But the underlying principle is old and, I think, right: copyright is a reward for human creative labour, not for owning a machine that produces output at scale.
Notice the strange asymmetry this creates. AI firms want maximum freedom on the input — the right to ingest everyone’s work without a licence — while their users often want maximum ownership on the output. You can’t cleanly have both. If the output of a trained model is “just maths, not really anyone’s expression” when that helps a company avoid paying for training data, it is hard to then insist that the same output is a fully authored, ownable creative work when a customer wants to sell it. The law is slowly, awkwardly, refusing to grant both wishes at once.
Who should actually get paid?
Strip away the legal vocabulary and the real question is about compensation and consent. Millions of people did unpaid labour — every blog post, every uploaded photograph, every answer written to help a stranger on a forum, every open-source commit — and that labour became the training substrate for tools now worth staggering amounts of money. The value those people created did not disappear. It moved. The question of who owns AI is, underneath, the question of who owns the accumulated output of everyone who ever put something online.
There are a few honest answers to “who should get paid,” and they are not mutually exclusive:
- The direct creators whose work is identifiable in the training set — authors, artists, musicians, photographers, developers — have the strongest moral claim, and increasingly a legal one.
- The platforms that hosted the data often claim a cut, sometimes via terms of service that quietly granted them the right to license your uploads for AI training. Whether that “consent” was ever meaningful is a fair thing to be angry about.
- The public at large has a claim too, because a great deal of training data is collectively produced knowledge — encyclopaedias, public discussion, shared reference works — that no single person owns.
What almost nobody defends, once you say it plainly, is the current default: that the value flows overwhelmingly to the model owner, and to almost no one else. That is not a law of nature. It is a choice about where the money stops.
The deeper pattern: enclosure in a new costume
This is where I want to zoom out, because I don’t think this is fundamentally a story about copyright doctrine. It is the latest instance of a move I keep seeing repeated across the history of technology: the same move, a new machine, every time. Something that was common — held loosely, shared, produced by many — gets fenced off, concentrated, and then rented back to the very people who made it.
We have watched this happen before. Common land was enclosed and the people who worked it became tenants on ground they used to share. Craft knowledge was pulled into factories and workers were paid a wage to operate machines that embodied skills they once owned. This is precisely how technology gets captured: a genuine collective advance — and AI is a genuine advance — arrives wrapped in an ownership structure that quietly routes the gains to whoever controls the machine. The productivity is real. The question is always the same: who gets to keep it.
Your writing, your images, your code were the common land. The model is the fence. And the subscription is the rent you now pay to work ground you once shared.
With AI the enclosure is unusually elegant, because the raw material is us. Your writing, your images, your answers, your code — the accumulated creative output of an entire species — becomes the enclosed commons. The model is the fence around it. And the monthly subscription is the rent you pay to draw water from a well your own labour helped fill. When the tool that learned from a million photographers is then sold back to photographers as the thing that undercuts their rates, the loop is complete. This is the mechanism behind what some now call technofeudalism — a world where a few platform lords own the essential infrastructure and everyone else pays to access what was built, in large part, from their own contributions.
I want to be careful not to make this sound like a conspiracy. Mostly it isn’t. It is what happens by default when a powerful general-purpose technology meets an ownership structure that has no obligation to share the upside. Nobody has to be a villain for the value to end up concentrated. That is exactly why it keeps happening — and why it takes deliberate effort, and usually organised pressure, to bend the outcome toward something fairer.
What fair compensation and consent could look like
So what would a better version look like? Not a ban on AI — that ship has sailed, and honestly the tools are too useful to wish away. The goal is to change who the gains flow to. A few directions seem both plausible and fair:
- Opt-in and real consent for training data. The default should flip from “everything is scrapable unless you sue” to “your work is used only if you allow it.” Machine-readable ways to say yes or no to training, respected in practice, would restore a shred of the consent that was skipped the first time around.
- Collective licensing and royalties. Music worked out, imperfectly, how to pay millions of rights-holders when their work is played, through collecting societies and blanket licences. A similar model — pooled licensing that pays into the pot when work is used to train, or is clearly reflected in output — could route money back to creators at scale without requiring a lawsuit per sentence.
- Provenance and attribution. Even where cash is hard, credit is not. Systems that can point back to influential sources make attribution — and eventually payment — technically possible rather than hand-waved away.
- Shared upside, not just individual payouts. Because so much training data is collectively produced, part of the answer may be collective too: public funds, data trusts, or dividends that treat the training commons as something the public part-owns, not merely a free input for private firms.
None of these is a finished solution, and each has hard edges. But they share a principle worth holding onto: the people whose labour and creativity make a technology possible should not be the ones it is quietly used against. Consent before taking. A share of what your work helps create. Credit where it is due. These are not radical demands — they are the ordinary terms we expect in almost every other exchange.
The copyright fight will grind on in courtrooms for years, and the precedents matter. But the case files are really a proxy for the older argument underneath them. Every powerful new machine forces the same negotiation about who benefits, and the outcome has never been decided by the technology alone — it has been decided by whether the people who fuel it insist on their share. AI has simply made the stakes, and the raw material, unmistakably personal. It is your work in the machine. The only real question is whether you get any say in what it does next.
Frequently asked questions
Is it legal to train AI on copyrighted work?
That is exactly what the courts are still deciding. Some training uses are being argued as fair use or fair dealing; other cases have ended in large settlements. The law is unsettled and varies by country, so treat any confident answer with caution.
Can you copyright something made by AI?
In several jurisdictions, purely machine-generated work with no meaningful human authorship cannot be copyrighted, while work where a human shapes the output substantially may qualify. The line is being drawn case by case.
Do creators get paid when AI trains on their work?
Mostly not, so far. Some settlements and licensing deals have begun to route money to rights-holders, but the vast majority of the writing, art and music used to train today's models was taken without payment — which is the heart of the fight.