India explainer
Your Data Is the New Oil — and India Is the New Oilfield
‘Data is the new oil’ is a cliché — but clichés can be true. If data is the resource, India is one of the richest fields on earth, and the question is the old one: who owns what gets extracted?
“Data is the new oil.” You have heard it in boardrooms, at tech conferences, in the mouths of ministers announcing the next great leap. It is a tidy line, and like most tidy lines it hides more than it reveals. What it points at is real: there is a resource being pulled out of the ground, refined, and sold at enormous markup, and whoever controls that pipeline controls a great deal of the future. But the metaphor also breaks in ways that matter — and once you see where, you start to see something I would call data colonialism: raw human behaviour extracted from places like India, refined into valuable products somewhere else, and quietly returned as something we now have to rent back. This is the same move I keep seeing across the history of technology — a new machine, the same old pattern of who takes and who pays.
Where the oil metaphor holds
Let me give the metaphor its due, because it is not wrong to start there. Oil in the ground is nearly worthless to the person standing on top of it. It becomes valuable only when someone with capital, refineries, and distribution turns crude into petrol, plastics, and jet fuel. The wealth accrues not to the land but to whoever owns the refining and the pipes.
Data works the same way at that first step. Your search queries, your location traces, your voice notes, your late-night scrolling, the exact half-second you hesitate before tapping “buy” — each fragment is close to worthless on its own. Aggregate a billion such fragments, run them through the expensive machinery of storage, modelling, and machine learning, and you get something worth staggering sums: the ability to predict and shape what people do next. The value is not in the raw stuff; it is in the refining, and the refining is owned by very few.
So far the analogy earns its keep. It explains why the companies sitting on the largest behavioural datasets are among the most valuable enterprises ever built, and why they guard that data so fiercely. This is one instance of a broader pattern I have written about separately — how technology gets captured — the same story wearing a data-shaped mask.
Where it breaks — and why the break matters
Here is the crack. Oil is rival and finite: if I burn a barrel, you cannot burn the same barrel, and that scarcity is the whole economic logic. Data is the opposite — non-rival and effectively infinite. Your behavioural trail can be copied a thousand times at no cost, sold to a dozen buyers at once, fed into one model today and another tomorrow, and none of it is used up. The “barrel” never empties.
That difference changes the moral arithmetic completely. When an oil company pays you for the crude under your land, the transaction at least resembles a fair trade: they take the thing, you no longer have it, you get compensated. When a platform takes your data, you still “have” it — your life is unchanged in the moment — so it feels free, like nothing left your pocket. That is the trick. You did give something away — not the copy, but the leverage. You handed over raw material that will be refined into a system built to predict and nudge you, and you were paid nothing because the taking felt like nothing.
The oil metaphor lets companies claim they are merely “harvesting” a resource, when what they are really doing is enclosing a commons that we all generate and none of us are paid for.
The metaphor is useful for explaining the refining, and dangerous for explaining the taking. It naturalises extraction — as if the resource were just lying there — when in fact the resource is us, continuously, flowing out cheaply only because we have not built the institutions to price it, tax it, or refuse it.
Why India is the new oilfield
Now place that pattern on a map and ask where the richest field is. Increasingly, the answer is India — and it is worth being precise about why.
Consider the raw geology of it: a population well over a billion, hundreds of millions of whom came online not gradually but in a rush, over a handful of years, on the back of some of the cheapest mobile data in the world. That is a first-generation internet user base arriving at extraordinary speed — people whose entire digital life, from first search to first payment, is recorded from the start.
Layer on top of that India’s digital public infrastructure — the identity layer, the payments rails, the account and data-sharing frameworks built at genuinely impressive scale. I have written elsewhere about the double edge of Aadhaar and UPI, and the tension is exactly this: the same rails that let a street vendor accept a payment and a migrant worker prove who they are also generate an extraordinarily fine-grained, real-time record of how a billion people transact and move. Built well, that infrastructure is a public good. Left ungoverned, it is one of the most legible behavioural fields on earth.
Diversity multiplies the value. Dozens of languages, enormous variation in income, geography, and behaviour, all captured in one interoperable system. For anyone training the next generation of models, that is not just a big dataset but a uniquely varied one — and variety is exactly what these systems are hungry for. India is attractive to the extractors not despite being complex but because it is complex, online, and fast.
How the extraction actually works
Strip away the branding and the mechanism is old and simple. Raw behaviour is taken cheaply here — usually in exchange for a free app, a convenient service, a few rupees of cashback. It is shipped, physically or legally, to infrastructure owned elsewhere, where it is refined at great expense into high-value products: recommendation engines, credit-scoring models, ad-targeting systems, and increasingly the large AI models everyone now wants to build a business on top of. Then those refined products are sold back into the same market the raw material came from, at prices set by the refiner.
Notice the shape. Value flows out as cheap crude and flows back as expensive finished goods. The margin — the gap between what the raw behaviour cost to acquire and what the refined product sells for — stays at the centre, with whoever owns the refinery and the model weights. The person in Kochi or Kanpur whose clicks trained the system does not get a share of that margin. They get to be a customer.
This is why the question of who owns AI is not separate from the data question — it is the same debate one step downstream. The models are the refineries, and everyone else lines up to rent access to their own aggregated selves.
The colonial echo
I do not use the word “colonialism” loosely, or as mere rhetoric. I mean that the structure rhymes.
The old colonial economy ran on a specific arrangement: extract raw materials from the periphery — cotton, indigo, spices, ore — ship them to the centre, manufacture finished goods, and sell those goods back at a markup, while discouraging the periphery from building its own industry. India knows this arithmetic intimately. Raw cotton left; finished cloth returned; the loom that could have kept the value at home was kept idle. The colony supplied the input and bought the output, and was kept from becoming much more than a supplier and a market.
Raw material out, finished goods in, and the machinery of value kept firmly at the centre. Change “cotton” to “behaviour” and “loom” to “model,” and the sentence still reads true.
The data economy can run on the same arrangement if we let it. Raw behaviour is the new commodity crop; the model is the new mill. The risk is that India again supplies the input and buys the output — trains the systems with its people’s lives and then rents those systems back — while the machinery that turns the one into the other, and the wealth it throws off, stays offshore. It is not a perfect one-to-one: the coercion of empire is not the coercion of a terms-of-service checkbox. But the flow of value, and the concentration of control, follow a path we should recognise, because we have walked it before and paid dearly to get off it.
The enclosure lens
There is an even older frame that clarifies all of this: enclosure. Before the factory, before the colony, there was the commons — land that belonged to everyone and no one, that ordinary people used to graze animals, gather wood, and survive. Enclosure was the process of fencing that commons, declaring it private property, and turning the people who had always used it into tenants or trespassers. The grass did not change. The fence changed everything.
Data is a commons of exactly this kind. It is generated collectively — none of your behavioural data means much without everyone else’s to compare it against — and yet it is being fenced off piece by piece, aggregated into private enclosures, and sold back to the very people who generate it. You are being turned into a tenant of your own behaviour. This is the through-line I keep returning to across the whole history of technology: the same move, a new machine, every time. Someone finds a commons, draws a fence around it, and what was shared becomes something you have to pay to enter. The land, the loom, the airwaves, and now the data trail — different machines, identical move. It is also the mechanism underneath what some now call technofeudalism: when the fences are digital and the landlords own the platforms, rent replaces ownership as the basic economic relationship.
What keeping the value at home could look like
If that is the diagnosis, what is the honest prescription? Techno-optimism (build more infrastructure and prosperity follows) ignores who controls the infrastructure. Pure refusal (opt out, delete the apps) is not available to the hundreds of millions for whom these systems are now how you get paid, prove who you are, and reach the state. The more serious answer is to change who owns the refinery and who captures the margin. A few directions seem worth taking seriously, each with real caveats.
Data rights that have teeth. If behavioural data is a resource, the people who generate it should have enforceable rights over it — to know what is collected, to move it, to withhold it, to have it deleted. India has taken formal steps toward a data-protection regime, and that direction matters. But rights on paper are not rights in practice; the caveat is enforcement. A right you cannot exercise without a lawyer, or one riddled with exemptions, mostly relocates power rather than redistributing it.
Governance of the public infrastructure. India’s digital public infrastructure is a genuine achievement, and the strongest card in this hand — precisely because it is, in principle, public. Rails built and governed in the public interest can be a real alternative to private enclosure: shared plumbing that many players build on, rather than one owner’s toll road. The caveat is that “public” is a claim to be verified, not assumed. Done right, it means transparent governance, strong limits on how the data can be repurposed, and independent oversight with the power to say no. Done wrong, it simply builds a more efficient pipeline to the same offshore refineries.
Domestic refining, not just domestic drilling. Keeping value at home ultimately means owning the mills, not just guarding the crude — the ability to build and run the models here, on our own compute, in our own institutions. That is expensive and slow, and there is a real risk of merely swapping foreign landlords for domestic ones. But it is the part of the chain where cheap behaviour becomes expensive capability, and a country that supplies the input while renting all the output has not escaped the old arrangement — only modernised it.
Collective, not just individual, bargaining. Because data is a commons, the individual opt-out is a weak instrument; your lone refusal barely dents a billion-record dataset. The more promising ideas treat data as collectively held — data trusts, cooperatives, public mandates that bargain on behalf of many at once — so the people who generate the resource can negotiate over the margin rather than clicking “accept” one at a time. This is early and unproven, but it matches the shape of the problem, which the individual consent model never has.
I will not oversell any of this. Every one of these directions can be captured in turn — that, after all, is the pattern I keep describing. The point is not that we have found the exit. It is that the field is not neutral geology waiting to be drilled; it is a commons being enclosed, and who ends up owning the fences is still, for now, an open question.
“Data is the new oil” was always half a truth. The other half is the one the slogan is designed to keep you from noticing: the oilfield is you, the barrel never empties, and what decides whether the value stays here or flows out is who we let build the fences — and who holds the keys.
Frequently asked questions
What is data colonialism?
The idea that today's extraction of personal data echoes historical colonialism: raw material — here, human behaviour — is taken cheaply from many people, processed elsewhere, and sold back as high-value products, with the profit and control staying at the centre.
Why is India central to the global data economy?
Scale. With over a billion people rapidly coming online through cheap data and digital public infrastructure, India generates one of the largest streams of behavioural data anywhere — extremely valuable for training and targeting AI systems.
Is ‘data is the new oil’ actually true?
As a metaphor, largely yes: data is a raw resource that becomes valuable once refined, and control of it confers power. But unlike oil, data is not used up and can be copied endlessly — which makes who owns and governs it even more important.