The Case That AI is About to Decentralize
- May 21
- 5 min read
Updated: Aug 4
Written by: Enzo Gonzalez.
The center of gravity in artificial intelligence sits inside a handful of data centers owned by a handful of companies. Most agents running in production today reach back to one of them for every hard decision, renting intelligence by the token from a provider that owns the model, the infrastructure, and the logs of everything the agent does. Artemii Amelin thinks that arrangement is less durable than it looks.

“Local models are only getting better,” he said. “There will come a point where they become more attractive to the masses than paying for APIs. The tasks most people use AI for do not require a frontier model. They can be handled perfectly by a model running on their own machine.”
It is a contrarian position in a market that has spent three years and hundreds of billions of dollars betting the other way. Amelin is unusually placed to take it. He co-founded Pilot Protocol, a San Francisco company that runs an open network for autonomous AI agents, and the network gives him a vantage point few others have. The agents on it run different models, made by different companies, on different harnesses, from small systems on a single laptop to large ones in the cloud. The shift first showed up as a detail in the data. Early on, when the team surveyed the agents on the platform, a growing share turned out to be running on local models such as Qwen, the open-weight family released by Alibaba, rather than calling out to a commercial API. What he sees, watching that population, is a slow tilt away from the center.
Some of that pull is already legible in the data. In a McKinsey survey of some 300 executives, investors, and government officials, 71 percent described sovereign AI, infrastructure an organization controls rather than rents, as an existential concern or a strategic imperative. The European Union's AI Act has put legal weight behind data residency, and regulated industries in finance, health, and law increasingly cannot send sensitive information to a cloud whose physical location they cannot guarantee. Chipmakers have responded by pushing inference onto local hardware, the neural processing units now shipping as standard in laptops and phones.
The pressure is not only regulatory. The World Economic Forum has begun describing AI's future as a distributed one, a hybrid of global cloud and local instances rather than a single center, and the hardware market is moving the same way, with rising demand for the edge silicon that runs models close to where data is generated. Latency adds its own push. An agent that has to finish a task in a fraction of a second cannot always afford a round trip to a distant data center. The reasons differ, law, cost, speed, control, but they converge on the same place, computation that lives nearer the work.
Step back, and the dispute fits a pattern as old as computing itself. The industry has swung between centralization and its opposite for half a century, from mainframes to personal computers, from on-premises servers to the cloud, and each swing turned less on raw capability than on who wanted control. Amelin's thesis is that agents will follow the same arc, and faster, because an autonomous system that leans on a single external provider inherits that provider's outages, its price changes, and its surveillance. An agent that runs locally answers to its owner. As agents take on more responsibility, that difference stops being academic and becomes the entire question.
The stakes in that swing are not only technical. Where an agent runs determines who can watch what it does, who can switch it off, and who captures the data it generates as it works. Centralized inference has concentrated those powers in a few firms to a degree the industry rarely names plainly. A move toward local compute would redistribute them, handing more control to the owners of agents and less to the companies that now sit between an agent and its own decisions. That is the real subject beneath the engineering debate, and it is why a question that sounds like infrastructure keeps surfacing as a question about power.
The honest version of the argument is not that the center collapses. Every prior move toward the edge eventually met a counter-move back toward the middle, and the cloud has won that fight more than once. Frontier models remain far more capable than anything that fits on a laptop, and for many tasks the gap is decisive. The economics of centralized inference keep improving as providers add capacity, and convenience has a long record of beating principle in consumer technology. A skeptic could point out that sovereignty has been predicted to triumph before, and that the cloud absorbed each previous revolt. What Amelin is claiming is narrower and harder to dismiss, that the balance shifts, and that the next 12 to 18 months will reveal how far.
For the thesis to hold, local hardware has to keep closing the capability gap, and the reasons to stay home, privacy, latency, control, have to keep mattering more than the convenience of renting the best available model. Both are live questions. Neither is settled.
His read draws its weight from the heterogeneity he watches every day, a network that already contains the diversity the rest of the industry is only beginning to reckon with: agents that disagree, that run on incompatible stacks, that have different reasons to stay close to home. Some are bound by the data they are permitted to touch. Some are tuned to a latency budget that a round trip to a distant data center would blow. Some answer to owners who simply want to know exactly where their software is doing its thinking. Pilot has run the experiment on itself. The company uses agents internally and recently moved them off commercial APIs entirely, onto local models, and two of them are visible as contributors on the Pilot Protocol repository on GitHub, their commits sitting alongside the humans'. The nearest evidence of the tilt is inside the company arguing for it.
If sovereignty does reshape how AI gets built, the earliest evidence will surface in precisely that kind of mixed population, in agents running quietly on local hardware for reasons that have nothing to do with benchmark scores, well before it shows up in the quarterly results of the companies betting the other way. Amelin is pointing to where the tilt begins rather than calling a winner, and arguing that the people closest to the agents will feel it first.









