The old cloud data centre was a real estate model wrapped around uptime engineering. It sold virtual machines, storage, software services and geographic redundancy. The AI data centre is different. It is closer to a power plant, a semiconductor packing problem and a financial instrument, all compressed into a building full of GPU racks.
That shift is now visible in the numbers. The International Energy Agency estimates that data centre electricity consumption rose by more than 15% between 2024 and 2025, taking the sector to almost 500 terawatt-hours a year, or a little over 1.5% of global electricity demand. Its base case still moves from roughly 415 TWh in 2024 to about 950 TWh in 2030.
The growth is not evenly distributed. It clusters around AI infrastructure: accelerated servers, high-bandwidth memory, dense networking, liquid cooling, substation capacity, grid studies and a construction schedule that often moves faster than the local utility planning cycle.
- AI data centre economics are shifting from square feet to secured megawatts, rack density and time-to-power.
- Blackwell-era racks have pushed mainstream AI infrastructure into liquid-cooled, high-voltage, high-density design.
- Connection queues contain real demand and speculative demand, which makes grid planning harder.
- Hyperscale capex is no longer just a software-sector line item. It is competing with energy-sector investment.
What is an AI data centre?
An AI data centre is a facility designed to train, fine-tune or run large AI models at scale. The phrase gets used loosely, but the real distinction is technical: AI workloads rely on tightly linked accelerators, usually GPUs or other specialised chips, that must communicate with extremely low latency and very high bandwidth.
That makes the facility less modular than the older cloud model. In a standard hyperscale data centre, many workloads can be distributed across thousands of general-purpose servers. In a high-end AI cluster, thousands of accelerators need to behave like one machine. The building becomes part of the computer.
NVIDIA's GB200 NVL72 is a useful marker. The system links 36 Grace CPUs and 72 Blackwell GPUs in a liquid-cooled rack-scale design, with a 72-GPU NVLink domain and 130 terabytes per second of low-latency GPU communication. The point is not the branding. The point is the architecture. Compute, memory, networking, power conversion and cooling are collapsing into one physical envelope.
For operators, the operating question changes. It is no longer "How many servers fit in this hall?" It is "How many usable kilowatts can be delivered to each rack, cooled continuously, backed up, connected to fibre and kept stable when the workload swings?"
Training, inference and the shape of demand
Training runs create sustained bursts of compute demand. Inference, especially for consumer products and enterprise AI applications, spreads demand across the day. The industry is now adding reasoning models, multimodal generation and agentic AI systems that can call tools, write code, search databases and run multi-step tasks without a user staring at each intermediate step.
That matters for infrastructure. Short text requests can be cheap in energy terms. Rich document analysis, image generation, video generation and autonomous agents can run far hotter. The facility has to be sized for peak concurrency, not the average elegance of a model benchmark.
Why rack density is rewriting the building
In older enterprise environments, a rack drawing 5 kW to 10 kW was ordinary. Some cloud halls moved higher, especially for dense compute. AI racks have broken that mental model.
The IEA's April 2026 update traces the jump in NVIDIA rack power from roughly 13 kW for Ampere-era systems in 2020 to about 130 kW for current Blackwell rack designs. Announced Rubin architecture moves toward 600 kW per rack, with a path toward 1 MW. A rack-sized object begins to behave like heavy electrical equipment.
| Design area | Traditional cloud data centre | AI data centre pressure point |
|---|---|---|
| Compute layout | General-purpose servers spread across resilient pools. | GPU clusters arranged to minimise latency and maximise accelerator utilisation. |
| Rack power | Often planned around single-digit to low double-digit kilowatts per rack. | High-density racks moving through 100 kW and beyond, with future designs far higher. |
| Cooling | Air cooling, hot/cold aisles and evaporative systems. | Direct-to-chip liquid cooling, rear-door heat exchangers, immersion trials and warmer closed loops. |
| Electrical design | Conventional AC distribution with established UPS and backup patterns. | Higher voltage, more DC distribution, power electronics, solid-state transformer research and local batteries. |
| Development constraint | Land, fibre, tax treatment and market proximity. | Grid connection, transformer lead times, secured energy, water politics and capex financing. |
High density does save space. It also moves the stress elsewhere. Floors carry heavier racks. Busways and cabling have to tolerate higher loads. Mechanical rooms grow. Fire suppression, leak detection, vibration, maintenance paths and commissioning procedures become more exacting. A modern AI hall looks clean. Underneath, it is a dense electrical and thermal system with little tolerance for improvisation.
Liquid cooling moves into the core design
Air still matters. It will remain in data centres for years. But air alone is a poor match for the hottest AI racks, where the heat is concentrated in accelerators, memory and networking hardware packed into a small footprint.
Direct-to-chip liquid cooling routes coolant through cold plates attached to the hottest components. Rear-door heat exchangers remove heat as it leaves the rack. Immersion cooling places server hardware in dielectric fluid, though operational maturity, serviceability and supply-chain comfort vary by deployment. The direction is plain enough: liquid is becoming a default design assumption for the densest AI zones.

Cooling is also where the public-resource debate gets messy. Water-based evaporative systems can reduce electrical demand for cooling, which helps power usage effectiveness. They can also create local tension in water-stressed regions. Dry coolers and closed-loop liquid systems can reduce onsite water consumption, but the full footprint still includes electricity generation, semiconductor manufacturing and construction.
PUE remains useful, but it is not enough. A facility can improve PUE while shifting stress to water, backup fuel, grid reinforcement or upstream power generation. Water usage effectiveness, carbon intensity by hour, rack utilisation and curtailment behaviour are becoming part of the same diligence conversation.
The grid queue is the market signal
It is tempting to treat power as just another data centre input. That is no longer sufficient. In AI infrastructure, power is the product bottleneck.
The simple claim that the world lacks enough renewable generation misses the harder problem. Solar, wind, nuclear, geothermal and gas supply all matter, but data centres need firm, deliverable power at a specific node, on a specific timeline, with transformers, substations, switchgear and studies completed. Grid connection can take five to ten years in some jurisdictions. Transformer lead times can run two to three years. Gas turbines can stretch toward five.
Connection queues are where the tension shows up first. They include serious projects, land-bank options, speculative applications and duplicate requests from developers trying to secure a place before competitors do. Utilities then face a planning problem: build infrastructure too slowly and real projects leave; build it around inflated queues and ordinary ratepayers may carry stranded costs.
Texas offers the cleanest public signal. According to IEA analysis of ERCOT data, the large-load connection queue, about three-quarters data centres, rose from roughly 63 GW in December 2024 to 130 GW four months later, then to more than 230 GW by January 2026. The state's all-time peak demand is 85 GW.
Phantom demand is a planning problem, not just a hype problem
"Phantom demand" does not mean all proposed data centres are fake. It means the queue is larger than the set of projects likely to reach construction. Developers file early because connection rights have option value. Utilities cannot treat every filing as guaranteed load, but they also cannot wait for perfect certainty.
This is where data centre policy is becoming electricity policy. Minimum demand charges, exit fees, upfront connection payments, non-firm connection offers and large-load rate classes are no longer niche regulatory instruments. They decide who pays for substations and who absorbs the risk when a 500 MW project changes schedule.

The new energy sourcing stack
Hyperscalers spent the last decade becoming major buyers of renewable power purchase agreements. That model still matters. It does not fully solve the 24-hour physics of an AI cluster.
The IEA notes that operators are now using a wider stack: corporate PPAs, hourly matching, battery energy storage systems, onsite generation, nuclear and geothermal deals, and in some markets natural gas capacity that can be built faster than large transmission projects. The result is a less tidy story than the annual renewable-procurement headline suggests.
Annual matching can make a company's books look cleaner than the local grid serving the data centre at 9 p.m. in a constrained region. Hourly matching is stricter. It asks whether low-emissions generation is available when the load actually runs. That tends to push buyers toward storage, firm clean power or demand flexibility, none of which is free.
Some operators will pursue onsite power to avoid waiting years for a grid upgrade. That can mean gas turbines, fuel cells, dedicated renewable generation, batteries or hybrid structures. The trade-off is local. Faster power can reduce queue exposure. It can also invite scrutiny over emissions, air permits, fuel logistics and whether private generation weakens or supports the public grid.
Capital and sovereign AI
The AI infrastructure cycle is now large enough to change capital markets. The IEA estimates cumulative data centre investment of US$3.9 trillion between 2026 and 2030. It also says leading hyperscalers and neo-clouds have announced a 75% increase in 2026 capital expenditure, reaching US$715 billion.
This is why AI data centre analysis cannot stop at model quality or chip benchmarks. A GPU cluster that cannot secure power is stranded inventory. A site that cannot get transformers is a press release. A company that has to fund land, grid interconnection, servers, cooling plants and multi-year energy contracts is making an industrial balance-sheet decision.
The Stargate Project made that explicit in January 2025, when OpenAI and SoftBank announced plans to invest US$500 billion over four years in US AI infrastructure, starting with US$100 billion. The announcement named Texas, power, land, construction and equipment as part of the same buildout. Software companies do not usually publish requests for power and land. AI companies now do.
Governments have noticed. AI data centres are increasingly treated as sovereign AI infrastructure, tied to national security, industrial policy, research access, tax incentives and digital competitiveness. The IEA cites policy support across the United States, China, the European Union and Asia; it also notes that Singapore has moved from a pause on new projects toward a managed call for additional capacity.
There is a political bargain underneath that shift. Governments want local AI capability, skilled jobs, cloud resilience and investment. Communities see water stress, power bills, backup generators, traffic, noise and a small permanent workforce after construction ends. The winners will not be the regions with the loudest AI slogans. They will be the ones that can price those trade-offs before the queue fills.
What to watch next
The market will keep talking about GPUs. Watch secured megawatts. Watch transformer slots, interconnection deposits, water permits, backup-fuel strategy, rack densities, model utilisation rates and the cost of debt. Those figures will reveal more than most product demos.
Facility design will keep fragmenting. Some sites will run hybrid halls, with conventional cloud servers on air and AI clusters on liquid. Some will push toward high-voltage DC distribution. Some will use warmer cooling loops to reduce chiller dependence. Some will discover that the local grid, not the server vendor, sets the delivery date.
The cold data point is still Texas: by January 2026, ERCOT's large-load connection queue had climbed above 230 GW, while the state's all-time peak demand was 85 GW.
Planning AI-ready data centre capacity in Singapore?
Browse Singapore data-centre operators, colocation providers and infrastructure specialists who can help with power, high-density racks, liquid cooling and AI-ready facilities.
Frequently asked questions
What makes an AI data centre different from a normal data centre?
A traditional cloud data centre spreads many independent workloads across thousands of general-purpose servers. An AI data centre is built around tightly coupled GPU clusters that must act like one machine, with very high bandwidth and very low latency between accelerators. That forces much higher rack power density, liquid cooling and a building designed around delivered megawatts rather than floor space.
How much power does an AI rack use?
Far more than a conventional rack. Older enterprise racks often drew 5–10 kW. The IEA traces NVIDIA rack power from about 13 kW for 2020 Ampere-era systems to roughly 130 kW for current Blackwell designs, with announced Rubin-generation racks moving toward 600 kW and a path toward 1 MW. At those levels a rack behaves like heavy electrical equipment.
Why do AI data centres need liquid cooling?
Because heat is concentrated in a small footprint of accelerators, memory and networking. Air cooling struggles past roughly 30–50 kW per rack, so the densest AI zones use direct-to-chip cold plates, rear-door heat exchangers or immersion. Liquid is becoming a default design assumption rather than an exotic add-on.
Why is grid connection the real bottleneck for AI data centres?
GPUs and racks can be bought far faster than firm, deliverable power can be connected. Grid connection can take five to ten years in some markets, transformer lead times two to three years, and large gas turbines up to five. Connection queues also mix real and speculative projects, making it hard for utilities to plan — which is why power, not chips, increasingly sets the delivery date.
What is the status of data centres in Singapore?
Singapore paused approvals for new data centres around 2019–2022 over land, power and water constraints, then moved to a managed approach. Under its Green Data Centre Roadmap it has reopened capacity through controlled allocations that prioritise energy efficiency and access to green energy — a shift from outright pause to a managed call for additional capacity.
