Tech Features

When power becomes the bottleneck, efficiency becomes capacity

Published

on

Kayvan Karim, Programme Director of MSc Software Engineering, School of Mathematical and Computer Sciences, Heriot-Watt University Dubai 

For the past few years, the AI infrastructure race has largely been measured in scale: more GPUs, larger data centres and greater power capacity. That expansion is continuing, but the economics are beginning to change. Competitive advantage may increasingly depend not only on securing additional power, but on extracting more useful AI work from the power already available. Part of the reason is a change like AI demand. We are moving from relatively simple prompt-and-response systems towards agents that can reason across multiple steps, call tools, inspect results, revise plans and continue working autonomously. Anthropic’s latest Economic Index notes that Claude usage is increasingly shifting towards long-running agentic tasks and that more computationally intensive conversations tend to be associated with higher-value outputs.

This transition could have significant implications for infrastructure demand. A chatbot might generate one answer to one prompt. An agent performing a software-development, research or business task may invoke a model dozens of times, maintain a large context, call external tools and generate many intermediate reasoning steps before producing an outcome. Demand may therefore grow in two ways: more people using AI and more inference, model calls, and tokens processed for each task. Major AI laboratories are already working on this efficiency problem. OpenAI says one of its primary inference objectives is to serve more tokens from the same hardware, using techniques including scheduling, caching, kernel optimisation and improved model implementation. It also describes GPT-5.6 as being trained to accomplish more work per token. Google DeepMind is pursuing a similar direction: its Gemini 3.6 Flash was designed for scaled agentic workloads and uses fewer output tokens than its predecessor on several evaluations. In contrast, its recent agentic video system reduced token consumption by up to 88% for that workload.

Other approaches address efficiency at the model architecture level. DeepSeek-V3, for example, uses a Mixture-of-Experts design with 671 billion total parameters but activates 37 billion per token, so only part of the network is used for each computation. Meta has similarly worked on inference efficiency through grouped-query attention and more efficient tokenisation; the Llama 3 tokeniser was reported to require up to 15% fewer tokens than Llama 2 for equivalent text. Taken together, these approaches show that model capability is increasingly being developed alongside the cost of delivering it.

Model-level efficiency, however, is unlikely to remove the infrastructure constraint on its own. Global data-centre electricity consumption was approximately 415 TWh in 2024, according to the International Energy Agency, and its base case projects this to reach around 945 TWh by 2030. AI is expected to drive most of that growth. Efficiency is therefore improving while aggregate demand continues to rise. One reason is the Jevons, or rebound, effect: efficiency improvements reduce the resources required for each unit of work, but lower costs can also encourage greater overall use. If agents become much cheaper to operate, organisations may respond by deploying more of them, running them for longer, or applying them to tasks that were previously uneconomic. Efficiency can reduce the compute required for an individual task while still increasing total demand.

That increased demand meets infrastructure that cannot expand as quickly. Models and software can improve quickly, but grids, substations, transformers and power-generation infrastructure usually have much longer development cycles. The IEA notes that while a data centre can sometimes be developed within two or three years, the broader energy infrastructure required to support it often involves longer planning and construction periods. Where grid capacity is constrained, each available megawatt becomes a more valuable production resource. The amount of power available remains important, but so does the amount of useful computation that can be produced within that power envelope.

That changes how we should understand capacity. Improvements in accelerator performance, model architecture, workload scheduling, caching, utilisation and inference software can increase computational output without increasing a site’s electrical connection. OpenAI’s recently reported Jalapeño inference hardware illustrates the direction of travel: the company says the chip can deliver more AI work per unit of power while increasing throughput and reducing latency. Efficiency can therefore act as a form of virtual capacity. If two operators each control 100 MW, but one can consistently deliver substantially more useful AI work within that power envelope, their nominal capacity may be identical while their productive capacity is not.

The same constraint applies to physical space and cooling. AI systems are concentrating more computational power into individual racks, increasing both power density and heat output. Packing more accelerators into the same building only creates useful capacity if the electrical and thermal infrastructure can support them. This is one reason liquid cooling is moving from a specialist technology towards a more central part of AI data-centre design. Microsoft, for example, has introduced a closed-loop chip-level cooling architecture that it says eliminates evaporative water consumption for cooling and could avoid more than 125 million litres of water annually per data centre. The example also shows why power, cooling, water use and rack density cannot be treated independently.

The same shift creates a measurement problem. Power Usage Effectiveness, or PUE, has been valuable for showing how much facility energy is required beyond the electricity IT equipment consumes. It does not, however, measure whether that IT equipment is producing useful work efficiently. Uptime Institute’s 2025 survey placed average PUE at around 1.54 and noted that the headline industry figure had changed little for six years. Uptime has consequently argued for productivity measures that relate computational work to energy consumption. For AI inference, tokens per kilowatt-hour might offer one operational measure. Still, even that is incomplete: an efficient model that solves a task in 1,000 tokens may be more valuable than one generating 10,000. A more useful long-term measure may therefore be useful AI work per unit of energy, water and infrastructure.

Capacity will remain essential. The AI industry will continue to build larger data centres, secure new power supplies and deploy large quantities of computing hardware. As agentic systems create more persistent inference demand and physical resources become harder to expand, however, efficiency may increasingly determine the productive value of that capacity. Operators that can support more useful computation within the same power, cooling, water, and space constraints can accommodate more workloads without waiting for equivalent growth in physical infrastructure.

For AI infrastructure, installed megawatts will remain a headline measure. The more consequential measure may increasingly be how much useful AI work those megawatts can support.

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending

Exit mobile version