How much AI can you get from a kilowatt?

For years, the efficiency conversation in data centres has revolved around a familiar set of numbers. How much power does the IT equipment consume? How much additional energy goes into cooling and other infrastructure? How efficiently does the facility use the electricity it receives?

Power Usage Effectiveness became the shorthand for that conversation. Now AI is beginning to make that metric feel incomplete.

The question is no longer only how efficiently a data centre consumes electricity. It is also what that electricity produces. As GPU clusters become denser and AI workloads consume enormous amounts of power, the more consequential measure may be how much useful computation can be generated from every kilowatt.

Sunil Gupta, Co-founder, Managing Director and CEO, Yotta Data Services, believes this is where the industry’s efficiency debate needs to move. “The measure should be how much compute I can generate per kilowatt, how much compute I can generate per dollar or per rupee,” Gupta says.

For AI infrastructure, that could eventually translate into a metric such as tokens generated per kilowatt. The objective is not simply to consume less electricity, but to extract substantially more useful AI output from every unit of power consumed.

That seemingly simple shift has implications for everything underneath the AI workload, from the GPU and cooling system to the power infrastructure and even the design of the data centre campus.

PUE was never designed for AI

Traditional data centre efficiency was largely a facility-level equation. A lower PUE meant less energy was being spent on cooling and other overheads relative to the power reaching IT equipment.

AI introduces another layer. The computing equipment itself has become vastly more power-intensive, while the output expected from that power has also increased. A data centre could have an efficient cooling system and still deliver relatively poor compute efficiency if the underlying hardware is inefficient.

Gupta argues that efficiency therefore needs to be viewed across the entire stack. “My B300 is much more powerful from inferencing tokens point of view,” he says, pointing to the industry’s push towards greater output from successive generations of GPUs.

The resulting metric is less about the facility in isolation and more about the relationship between electricity and useful AI work. “If you are able to generate much, much more tokens for the least amount of kilowatt, it means you are more efficient,” Gupta adds. That changes the conversation from infrastructure efficiency to computational efficiency.

The megawatt is getting denser

The shift becomes more apparent when looking at the physical infrastructure supporting AI. For much of the data centre industry’s history, power density was a constraint, but not the defining constraint. AI is changing that. Gupta says that a building designed for 50 MW can increasingly be expected to accommodate 100 MW or more as workloads become denser.

The implication is that the value of a campus is no longer determined simply by how much land it controls. The ability to bring large amounts of power into that land, distribute it safely and remove the resulting heat is becoming equally important.

That also changes the economics. Gupta estimates that building a conventional data centre in India costs around $6 million per MW, including land, building and mechanical, electrical and plumbing infrastructure. A facility deployed as GPU capacity can require roughly $40 million per MW once the compute layer is included.

The difference is substantial. GPU infrastructure can therefore require roughly six to seven times the capital intensity of a conventional colocation facility. But the more important point is that the expensive component is also the component that cannot simply be purchased speculatively.

The race starts before the GPU arrives

An operator cannot economically hold billions of dollars of GPUs waiting for customers. Yet customers ordering large AI clusters cannot afford to wait years for the underlying infrastructure to be created.

This creates an unusual planning equation. Land, buildings, substations and power infrastructure take time to develop, but represent a relatively small proportion of the total investment in a GPU deployment. GPUs are far more expensive, but can be purchased against committed demand. “The things which are not very costly in the overall scheme of things, but which takes a lot of time, better do them upfront and in advance,” Gupta points out.

That is why Yotta has focused on having land, core-and-shell facilities and power infrastructure ready ahead of demand. The urgency is driven by the deployment window. “If I cannot deliver to a global customer GPU in five months, I’m not getting the business,” Gupta says. “They will have 10 other options.”

For an AI infrastructure provider, therefore, efficiency is also a function of time. A megawatt sitting behind an incomplete building has no economic value to a customer waiting for compute.

Cooling every watt of compute

Higher compute density creates another challenge. Every additional watt consumed by a GPU ultimately becomes heat that has to be removed. This is where liquid cooling enters the equation, but Gupta argues that the industry should not view it as a simple replacement for air cooling.

In a large GPU deployment, liquid cooling can handle the majority of the heat generated by GPUs and critical components, while networking equipment, storage, CPUs and other components continue to require air cooling.

“Your design now is a hybrid cooling part where 80% of the workload will be cooled with liquid and 20% of the workload will still be cooled by air,” Gupta says. That makes AI-ready infrastructure an engineering challenge extending well beyond the IT rack.

Higher rack densities can require additional transformers, generators and chillers. Existing facilities may not have sufficient space around the building or on the roof to accommodate that additional infrastructure.

“AI-ready is obviously a good jargon to speak, but essentially, you’re talking about a hyper-density data centre which can take liquid cooling to the racks and the density becomes 100 kilowatt-plus per rack,” Gupta adds. The efficiency of the GPU therefore cannot be separated from the efficiency of the infrastructure supporting it.

The power behind the kilowatt

The next question is where all that electricity will come from. For Gupta, sustainability cannot be reduced to simply buying renewable power. Solar and wind introduce intermittency, creating a requirement for storage or other mechanisms to ensure that power remains available when renewable generation falls.

Battery energy storage systems could play a role by storing electricity when renewable sources are producing and releasing it when they are not.

But Gupta also sees a longer-term role for small modular reactors. “The first solution I’m looking forward to is SMRs,” he says. The attraction is not simply nuclear power itself. It is the possibility of generating electricity behind the metre, close to the data centre, thereby reducing dependence on distribution infrastructure.

“You can be behind the metre, right next to your data centre building. You can generate your own power, plug into your building, and you are not then impacting the distribution grid at all,” Gupta states.

For now, commercial viability remains the central barrier. SMRs are therefore a future possibility rather than an immediate answer to the AI industry’s power requirements. But the direction of travel is clear. As AI campuses become larger and more power-intensive, the data centre operator will increasingly have to think like an energy operator as well.

From megawatts to useful intelligence

This is ultimately what makes the question of AI per kilowatt more important than it first appears. The industry is entering an era in which more power does not automatically mean more competitive infrastructure. The same megawatt can produce very different amounts of useful computation depending on the efficiency of the GPU, the cooling architecture, the power infrastructure and the workload being executed.

That is why Gupta believes the next generation of efficiency metrics must look beyond the facility itself. “If I’m efficient across the board, my GPU efficiency can be lower, my cost of building is lower, my cost of chip is lower, my cost of power is lower, everything is lower, then my end-to-end price to the end customer will be much lower,” he explains. The AI infrastructure race is therefore becoming less about securing the largest possible amount of electricity and more about what can be done with it.

PUE will continue to matter, renewable energy will continue to matter, cooling efficiency will continue to matter, but for AI, the ultimate measure may increasingly be what comes out at the other end of the power metre.

How many tokens? How much compute? How much useful intelligence?

Per kilowatt.

AIData centerGPUYotta
Comments (0)
Add Comment