Ricard Boada & Sofia Gumuzio
Volta co-founders
9 min
Intelligence has outgrown the ground it runs on.
For most of computing history, progress was gated by the processor – faster chips, denser transistors, better silicon. That constraint has moved. The frontier of AI is no longer set by ideas or algorithms. It is set by megawatts, by land, by fibre, by the physical machinery of computation. Intelligence has become an infrastructure problem – and infrastructure is the one thing the industry keeps treating as an afterthought.
Consider how a frontier model actually gets trained today. Power is procured from one party, land from another, cooling from a third, networking from a fourth and compute rented in slices from whoever has capacity that quarter. Each layer was designed in isolation, by people who never had to make it work as one system. The result is a chain of seams, and every seam is a place where cost accretes, where reliability frays and where a training run can die at hour 400 because two vendors disagree about whose problem it is.
We started Volta because we no longer believe that seam is a detail. It is the defining constraint of the age.
Every general-purpose technology eventually becomes a utility. Electricity did. You do not build your own power station, negotiate separately for turbines and transmission and metering, and hope they interoperate. You plug in. The complexity is absorbed by whoever owns the whole chain, from the generator to the socket.
Compute has not made that leap. It is still sold in parts, to be assembled by the buyer, at exactly the moment the stakes have never been higher. Volta exists to finish the transition – to deliver the Utility of Compute: power, capital, software and compute, owned and operated as a single vertically integrated stack. Not a marketplace reselling someone else's capacity. The whole of it, from the electron entering the substation to the FLOP leaving the cluster.
The physical expression of that idea is what we call an AI factory. Not a data centre with a fashionable name – a purpose-built machine in which the grid connection, the power hall, the cooling, the networking fabric and the racks are engineered together, toward a single number: useful compute, delivered without interruption.
For most of the last decade, the scarce thing in AI was the GPU. If you could get allocation, you could build. That era is closing. The chips are still hard to get, but they are no longer the binding constraint. The binding constraint is now everything around them: firm power, energised land, a grid connection, enough cooling to keep a rack from throttling, and the ability to hold tens of thousands of processors in lockstep for months at a time. The bottleneck has moved from the silicon to the substation.
Start with power, because everything else waits on it. In the United States, the generation and storage sitting in interconnection queues – waiting for permission to connect to the grid – passed 2,290 gigawatts at the end of 2024, according to Lawrence Berkeley National Laboratory. That is roughly twice the entire installed US power fleet. The typical project that reached commercial operation that year had spent about five years in the queue, up from under two in 2008. Historically, only around 14 per cent of queued capacity is ever built.
Read that again, because it is most of the story in one number. Connecting new power to the grid takes about five years, and the large majority of projects never make it. JLL puts grid connections in primary markets at more than four years on average. Gartner expects that by 2027, 40 per cent of AI data centres will be constrained not by chips or capital but by electricity they cannot get delivered.
Meanwhile the demand curve is close to vertical. US data-centre electricity use was about 176 terawatt-hours in 2023, roughly 4 per cent of the country's power. The Department of Energy projects it reaching 325 to 580 terawatt-hours by 2028 – as much as 12 per cent of national demand. The one thing in the queue that is rising rather than falling is natural gas, up more than 70 per cent year on year in 2024. The market has already worked out that renewables plus a five-year wait does not power a training cluster.
So the industry is doing the only thing it can: building its own power, on site, behind the meter. This is no longer exotic. xAI stood up a 100,000-GPU cluster in Memphis in around four months by bypassing the grid and running on truck-mounted gas turbines – more than 500 megawatts of them. Meta is powering a Texas site with hundreds of mobile turbines, and a Louisiana campus that scales to five gigawatts. The Stargate campus in Abilene runs on aeroderivative turbines rated at 1.2 gigawatts. Oracle has signed for up to 2.85 gigawatts of fuel cells at a single New Mexico facility. One tracker counts at least 46 data centres already running on roughly 56 gigawatts of behind-the-meter power.
But on-site generation is not a shortcut so much as a different queue. Gas turbines from the three manufacturers who build them at scale are now booked into 2028 and 2029. Large transformers carry four-year lead times, and prices have roughly tripled. The genuinely scarce assets in this build are not "data centres" – they are energised land, a gas interconnect, turbines, transformers, switchgear, water rights and air permits, plus the operational grit to assemble them on schedule. A gigawatt on a press release is not a gigawatt delivering electrons, and the gap between the two is measured in years. Every quarter of that gap is expensive: a gigawatt of AI cloud is worth something like 10 to 12 billion dollars of annual revenue, so bringing 400 megawatts online six months early is worth billions. In this build, speed is the moat – and speed is decided at the substation.
Assume you have solved power. The next wall is heat, and it arrives sooner than most people expect. A conventional data-centre rack drew 5 to 15 kilowatts. An Nvidia GB200 NVL72 rack draws 120 to 130. Air simply cannot move that much heat: its practical ceiling sits around 40 kilowatts per rack, above which the airflow required becomes, literally, hurricane-force in the cold aisle. Water carries roughly 3,300 times more heat per unit of volume than air. Past about 40 kilowatts you are on direct-to-chip liquid cooling whether you planned for it or not; past 100 you are looking at immersion.
This is not a swap you retrofit into a leased hall. A liquid-cooled 130-kilowatt rack needs its cooling designed with the silicon, plumbed through the building, matched to the local climate and water position. Get it wrong and the chips throttle – an H100 starved of cooling drops its clock within seconds – and because a training job is synchronised across thousands of GPUs, one throttled chip drags the whole run down with it. Goldman Sachs expects roughly three-quarters of AI servers to be liquid-cooled by the end of 2026. The building and the chip are now one machine, and have to be designed as one.
Here is the part that reframes everything, and it comes from Meta's own engineers. Training Llama 3's 405-billion-parameter model took a cluster of 16,384 H100 GPUs 54 days. Over that run it suffered 419 unexpected interruptions – about one every three hours – and more than half traced back to GPUs or their onboard memory. A single failed NVLink stalls the entire job, because the collective operation blocks until every participant reports in. The synchronised power draw of that many chips swung by tens of megawatts, stressing the grid feeding them.
And yet Meta held effective training time above 90 per cent. Only three of those 419 failures required serious manual intervention. That is the number that matters. At frontier scale, hardware failure is not an edge case to be avoided; it is the steady state to be engineered around – with checkpointing, hot spares, fast detection and automated recovery. Reliability here is not a GPU you rent. It is a systems property spanning power quality, the network fabric, the cooling loop and the orchestration layer, and no single one of those can deliver it alone.
Now put it together. The default way to get AI compute is to rent GPUs from a cloud, which leases power from a utility, over a grid connection nobody in the chain controls, in a hall someone else cooled to a spec chosen before these chips existed. Four or five owners, four or five incentives, and no single party with a view of – or responsibility for – the whole system.
That model is fine for what it was built for: bursty, forgiving, multi-tenant cloud. It is the wrong shape for a single synchronised machine drawing 50 megawatts that has to run flat out for months, where a cooling decision and a power-quality event and a network fault are the same problem wearing different hats. Every seam between owners is a place where cost accretes, where reliability frays and where the answer to "whose fault is this" is a meeting rather than a fix. You cannot rent your way to a system property.
Volta's answer is to remove the seams by owning them. We secure the power and the land first – deploying institutional capital at the source, where the megawatts and the interconnect actually are, rather than bolting energy onto a site chosen for other reasons. We design the hall around liquid-cooled, high-density racks instead of forcing next-generation silicon into last-generation buildings. We own the network fabric. And we operate power, cooling and compute as one system, with a single party accountable from the electron entering the substation to the FLOP leaving the cluster.
This is, in the older sense of the word, a utility: the hard, capital-intensive chain assembled once, by whoever is willing to own all of it, so the customer can simply draw on it. It is financed like infrastructure because that is what it is – multi-billion-pound, multi-year assets, not a software rollout. None of it is glamorous. Interconnect queues, transformer lead times, water permits and checkpoint recovery do not demo well. They are also the entire game.
We think the winners of this buildout will not be whoever accumulates the most GPUs. They will be whoever can turn a megawatt into useful compute – reliably, and first. That is the problem we have chosen, and the one we will keep writing about here.
If you are training at the frontier, scaling inference or allocating capital to what is becoming the defining asset class of the decade, we would like to talk. And if you would rather build the foundation than rent one, we are hiring.