// cloud & ai infrastructure · advanced

Compute as the New Currency: How AI Compute Demand Is Reshaping Hardware Markets and Corporate Valuations

17 min read· Updated 25 August 2026 · By TechDirectory Editorial Team

Share with your friends:

The public story of the AI boom is a software story — chatbots, copilots, agents, model releases arriving weeks apart. The economics underneath it are a hardware story, and a narrower one than most narratives allow. Large language models, multimodal systems and agentic workloads have converted intelligence into a manufactured commodity whose input costs are floating-point operations, memory bandwidth, interconnect capacity and megawatts. Those inputs are scarce, they are priced accordingly, and access to them now shapes which companies can compete at the frontier at all.

That scarcity, not model architecture, is the most economically consequential fact of the current cycle. It explains why four US technology companies have guided roughly US$695–720 billion of 2026 capital expenditure, although their accounting definitions differ; why a memory manufacturer just reported a 76% operating margin; why the world's most valuable company designs accelerators; and why the market capitalisation league table has been rebuilt around the AI hardware supply chain in under three years.

This essay develops that argument end to end: the scale of the demand shock, how workloads translate into hardware requirements, where supply bottlenecks, how scarcity transmits into valuations, the limits of the build-out, and a base, bull and bear path for 2026–2030. Quantitative claims rest on company disclosures, industry forecasts and published measurements, with sources listed at the end and uncertainty flagged where the evidence is contested.

The argument in brief: AI compute — accelerator FLOPs, high-bandwidth memory, scale-up interconnect and the power to run them — has become the scarce, meterable, forward-contracted input of the AI economy: functionally, a currency. Demand from training and inference has grown faster than any prior compute cycle, while supply is gated by four chokepoints — HBM output, advanced packaging, leading-edge wafers and grid power. That imbalance, not sentiment alone, is what repriced chip designers, memory suppliers, foundries, server vendors and the power chain. The valuations are defensible where scarcity rents are contracted for years ahead, and fragile wherever efficiency gains, power limits or depreciation reality bite faster than expected.

Table of contents

The AI compute demand shock, quantified

Start with training. Epoch AI estimates that the compute used to train notable frontier models grew at roughly four to five times per year for more than a decade — a compounding rate with no precedent in industrial input demand. The largest estimated runs approach the 1026 floating-point-operation class. Publicly disclosed frontier runs have used tens of thousands of accelerators, while some operating or announced clusters exceed 100,000; the distinction matters because estimates, procurement plans and installed capacity are not the same thing.

Inference scales with usage rather than research budgets, and the clearest public indicator is company-specific. Google disclosed in 2026 that its systems processed more than 3.2 quadrillion tokens a month as of May — up from roughly 480 trillion a year earlier and under ten trillion two years earlier. That is about sevenfold year-on-year growth at Google, faster than the historical frontier-training trend, while serving costs were falling. Whether that trajectory generalises across providers is unknown because the industry does not publish a consistent token census.

The capital response is proportionate. Microsoft, Amazon, Alphabet and Meta have guided to roughly US$695–720 billion of 2026 capital expenditure, depending on which lease and non-cloud items are included, versus about US$410 billion in 2025. Much of the increase is directed at AI data centres, accelerators, custom silicon and power. Gartner separately estimates AI semiconductors at about 30% of a US$1.32 trillion 2026 chip market, while WSTS's spring forecast runs higher at US$1.51 trillion on memory strength. Deloitte estimates fewer than 20 million generative-AI accelerators among roughly 1.05 trillion semiconductor units shipped in 2025. Those categories cannot be divided into one another, but together they show how narrowly the demand shock is concentrated on leading-edge logic, advanced memory and packaging.

That concentration is what separates this cycle from its predecessors. The cloud build-out of the 2010s bought commodity servers on mature process nodes from a competitive supplier base. The mobile cycle shipped billions of units, but horizontally, across many price points and fabs. Crypto mining consumed GPUs, but the compute produced nothing of intrinsic economic value — its demand collapsed with coin prices, and the hardware was fungible. AI compute is different on every axis: the output (tokens, model capability) has revenue attached; the hardware is a tightly integrated rack-scale system rather than a card; and substitution is gated by software ecosystems, memory supply and packaging slots measured in quarters. Demand that is concentrated, monetised and hard to substitute is precisely the kind that produces scarcity pricing.

What drives AI compute demand: models, tokens and agents

From model architecture to hardware bill

Each architectural trend of the past three years has translated into a specific hardware requirement. Dense scaling set the baseline appetite for FLOPs. Mixture-of-experts designs reduce arithmetic per token by activating only a subset of experts, but all weights still need storage and expert routing adds fabric pressure. Long-context and multimodal systems enlarge the key-value cache and make prefill attention expensive; autoregressive decode at modest batch sizes is often memory-bandwidth-bound, while prefill and highly batched inference can remain compute-bound. Reasoning and agent workflows add serial or parallel model calls, tool invocations and retrievals, but their multiplier is workload- and orchestration-dependent. AI compute demand is therefore denominated in four currencies at once — floating-point operations, bytes per second of memory bandwidth, bytes of capacity, and joules.

Layer diagram of the AI compute supply stack in 2026. Demand for training runs and inference tokens flows down through six layers: accelerators supplied by NVIDIA, AMD and hyperscaler ASICs; high-bandwidth memory from SK hynix, Samsung and Micron with supply substantially committed under long-term agreements; advanced packaging dominated by capacity-constrained TSMC CoWoS-class lines; leading-edge wafers concentrated at TSMC N3 and N2 nodes; rack power and liquid cooling at up to roughly 140 to 142 kilowatts in current reference designs; and grid interconnection with multi-year queues. Each layer lists its 2026 constraint status.
The AI compute stack in 2026. Demand passes through six layers; the tightest layer, not the average one, sets delivered capacity and price.

Accelerators: the rack is the product

The unit of AI compute procurement is no longer a chip but a rack. NVIDIA's GB200 NVL72 binds 72 Blackwell GPUs and 36 Grace CPUs into a single NVLink domain — 1.8 TB/s per GPU, roughly 130 TB/s of aggregate fabric and 13.4 TB of pooled HBM3e — in a liquid-cooled system. NVIDIA's preliminary Rubin specifications raise that to 288 GB of HBM4 and up to 22 TB/s per GPU, with 260 TB/s of rack-scale NVLink. NVIDIA expects partner shipments beginning in fall 2026. AMD's MI355X carries 288 GB of HBM3E, while the MI455X launched in July 2026 with 432 GB of HBM4 and 23.3 TB/s; Helios is AMD's 72-GPU rack reference platform. Google's 9,216-chip Ironwood superpod became generally available on 31 March 2026, AWS's Trainium3 UltraServer packs 144 chips and 362 FP8 petaflops, and Microsoft's Maia 200 was deployed in US Central in January while its SDK remained in preview. These are vendor peak or preliminary specifications, not cross-platform benchmarks, but the direction is uniform: more memory, more fabric and more facility engineering per rack. Our GPU architecture explainer and AI compute infrastructure guide cover the underlying mechanics.

High-bandwidth memory: the binding constraint

Autoregressive decode repeatedly streams active model weights and its key-value cache, so at modest batch sizes serving throughput is often constrained by memory bandwidth rather than peak arithmetic. Prefill, large batches and training can shift the balance back toward compute and network utilisation, which is why no single model-flop-utilisation percentage describes the fleet. High-bandwidth memory nevertheless remains a gating component. HBM4 doubles the per-stack interface to 2,048 bits and enables the bandwidth figures claimed for Rubin and MI455X. The market has exactly three suppliers: SK hynix, Samsung and Micron. SK hynix began HBM4 mass shipments in the second quarter of 2026 and disclosed long-term agreements with about ten key customers. Volumes are undisclosed, so the defensible conclusion is that supply is substantially committed, not that every future stack is sold. HBM is now among the largest line items in an accelerator's bill of materials.

Power, cooling and the 140 kW rack

Compute density has outrun the buildings. Legacy enterprise halls were engineered for 8–15 kW per rack, while current high-power AI reference designs reach roughly 140–142 kW, making direct-to-chip liquid cooling, redesigned busways and facility water loops part of the compute bill of materials. Rubin Ultra NVL576 is described as an eight-rack compute domain, not a single-rack specification; facility designs must be evaluated at both rack and domain level. The International Energy Agency estimates data centres used about 485 TWh in 2025 and projects roughly 950 TWh by 2030, with AI driving most of the increase. Interconnection queues in major markets now stretch years. A booked accelerator allocation without a dated grid connection is inventory, not capacity. How this reshapes facility design is covered in our AI data centre infrastructure explainer.

Storage, networking and data movement

The layers around the accelerator are being pulled along. Frontier training writes multi-terabyte checkpoints on tight intervals; inference fleets offload key-value caches to flash; data pipelines feed clusters at rates that push high-capacity SSDs into the AI budget — one reason Gartner forecasts NAND prices up 234% in 2026. Networking has become a first-class line item rather than plumbing: NVIDIA's data-centre networking revenue alone reached US$15 billion in its April-2026 quarter, nearly tripling year on year, while open scale-up standards such as UALink 1.0 (up to 1,024 accelerators in one domain) formalise the fabric layer the industry now competes over. Every additional metre a byte must travel costs joules and latency; the economics reward moving compute, memory and network closer together — which is exactly what advanced packaging does, and why it is scarce.

The supply side of AI compute: vendors, bottlenecks and contracts

One clock-setter, faster-growing challengers

NVIDIA still holds roughly three-quarters or more of AI-accelerator revenue through the first half of 2026, and its annual platform cadence — Hopper, Blackwell, Blackwell Ultra, Rubin — sets the industry's clock. The moat is systemic: CUDA's two-decade software accumulation and the NVLink fabric mean competitors must beat a shipping rack, not a chip. The fastest-growing segment, however, is silicon NVIDIA never sells: industry trackers project custom-ASIC shipments growing materially faster than merchant GPUs as the largest clouds scale first-party accelerators. Broadcom, a major supplier of custom-accelerator and networking technology to hyperscalers, reported fiscal second-quarter 2026 AI semiconductor revenue of US$10.8 billion, up 143% year on year. Marvell, its principal rival, grew fiscal-2026 revenue 42% to US$8.2 billion. A specialist tier led by firms such as Cerebras, Groq and Etched attacks inference cost per token with architectures that trade generality for memory locality or workload-specific efficiency. The fuller vendor mechanics are set out in our NVIDIA stack explainer and the companion AI chip manufacturers report.

The memory oligopoly

Three firms supply every HBM stack on earth, and 2026 made the consequences explicit. SK hynix reported second-quarter revenue of KRW 79.3 trillion, up 257% year on year, and operating profit of KRW 60.5 trillion — a 76% operating margin, its fifth consecutive record quarter. A commodity memory manufacturer printing software-company margins is the single clearest datum of scarcity pricing in this cycle. The tightness radiates outward: Gartner forecasts DRAM prices up 125% and NAND up 234% in 2026, with memory revenue between US$633 billion (Gartner) and more than US$800 billion (WSTS) depending on forecast vintage, and no broad supply relief expected before late 2027. Long-term agreements lock allocation years forward, which stabilises suppliers' revenue and transfers forecast risk to buyers — a structure closer to power purchase agreements than to historical DRAM spot markets.

Foundry concentration and advanced packaging

Beneath every brand sits a concentrated manufacturing base. Counterpoint Research estimates TSMC held 73% of pure-play foundry revenue in the first quarter of 2026; its second quarter delivered US$40.2 billion of revenue at a 60.3% operating margin, and management raised its 2026 capital budget to US$60–64 billion after spending US$40.9 billion in 2025. N2 entered high-volume manufacturing in late 2025, with N2P and A16 scheduled for the second half of 2026. The subtler chokepoint is packaging: HBM stacks must be bonded beside logic dies on interposers with exacting thermal and yield requirements. TSMC says advanced-packaging capacity is so tight that it constrains customer growth; that supports "capacity-constrained," not a claim that every 2026 slot is sold. Rubin, MI455X, Ironwood, Trainium3 and Maia are many logos at the price sheet and a highly concentrated queue at the wafer and package. Equipment supply is scaling behind it — SEMI forecasts US$165.9 billion of 2026 equipment sales; ASML guides to €43–45 billion with EUV capacity up roughly 30% for 2027 — but cleanrooms, tools, recipes and qualification form a serial chain measured in years. Announced capacity and shippable systems are different things.

Pull-through, purchase commitments and compute as a service

The scarcity pulls a long tail of vendors with it: server OEMs and ODMs whose AI systems now dominate their revenue mix, liquid-cooling and power-train specialists, optics and switching suppliers. It has also restructured how compute is bought. Multi-year agreements — HBM long-term arrangements, foundry prepayments, accelerator contracts and cloud reserved-capacity deals reflected in ballooning remaining-performance-obligation balances — now sit alongside spot procurement at the frontier, although take-or-pay terms are rarely disclosed. Anthropic said it planned capacity of up to one million Google TPUs and expected well over a gigawatt in 2026; Stargate disclosures similarly mix operating, under-development, secured and planned sites. Preserving those verbs prevents announcements from being mistaken for online megawatts. In parallel, a neocloud tier led by CoreWeave rents GPU-hours against those same supply lines, often financed by debt secured on the chips themselves. When an input is scarce, meterable, sold forward under contract and accepted as loan collateral, describing it as a currency stops being a metaphor and starts being a fair account of how the market treats it.

How AI compute scarcity repriced the hardware stack

The transmission from scarcity to market value runs through a specific sequence, and it is worth stating mechanically rather than rhetorically. First, allocation power: when demand exceeds supply at every price yet observed, the bottleneck owners set terms, which appears in the data as margin expansion — NVIDIA's gross margins in the seventies, TSMC's 60.3% operating margin, SK hynix's 76%. Second, visibility: long-term agreements, packaging reservations and cloud backlogs convert next year's demand into contracted revenue, which lets investors capitalise earnings years forward with unusual confidence. Third, the re-rating itself: higher earnings multiplied by higher multiples. NVIDIA passed US$1 trillion of market value in mid-2023, US$4 trillion in July 2025 and US$5 trillion by late October 2025; in August 2026 it trades around US$5.0–5.3 trillion as the world's most valuable company, having reported April-quarter revenue of US$81.6 billion (up 85%) with US$75.2 billion from data centre and US$49 billion of quarterly free cash flow. Broadcom and TSMC each crossed the trillion-dollar line during 2024–2025 on the same logic.

CompanyPosition in the stack2026 evidenceWhat it demonstrates
NVIDIAAccelerators, fabric, softwareQ1 FY2027 revenue US$81.6B (+85%); data centre US$75.2B (+92%); ~US$5T market valueSystem-level scarcity rents; the rack as the product
SK hynixHBM (with Samsung, Micron)Q2 revenue KRW 79.3T (+257%); 76% operating margin — and the shares fell ~11% on the printCommodity memory repriced as a strategic input; expectations now embedded
TSMCLeading-edge wafers + packagingQ2 revenue US$40.2B; 60.3% operating margin; US$60–64B capexThe toll booth every roadmap crosses
BroadcomCustom ASICs + networkingAI revenue US$10.8B in FQ2 (+143%); 20 GW financing platform with Apollo and BlackstoneCustom silicon growing ~3× the merchant market
HyperscalersThe buyers~US$695–720B combined 2026 guidance; definitions differDemand underwritten by the strongest balance sheets in corporate history
Circular diagram of the transmission loop from AI compute scarcity to corporate valuations. Demand outruns supply, giving bottleneck owners allocation and pricing power, which produces margin expansion and multi-year contracted backlogs, which drive earnings upgrades and multiple re-rating, which fund larger capital expenditure, which pulls more HBM, wafers, packaging and power and renews the demand pressure. Two counter-forces brake the loop from outside: algorithmic efficiency gains that reduce compute per task, and power and grid limits that delay deployment.
The scarcity-to-valuation loop — and the two brakes acting on it from outside: efficiency gains and power limits.

Distinguishing structural re-rating from transient enthusiasm requires a test, and the market itself has begun applying one. Structural claims rest on contracted, multi-year demand visibility; margins that survive capacity additions; and end-buyers whose AI revenue grows into their capital spending. Transient claims rest on spot shortage alone. Mid-2026 price action shows discrimination, not mania: SK hynix fell 11% on a record quarter that missed expectations; hyperscalers faced pointed investor scrutiny of capex through July 2026; commentary on the widening gap between AI capital spending and AI revenue has moved from the sceptical fringe into mainstream financial coverage. A further caveat belongs on any balance sheet built of accelerators: most operators depreciate AI hardware over five to six years while the silicon cadence is annual, and if competitive life proves shorter, current earnings across the ecosystem are overstated — a question prominent bears raised publicly in late 2025 and that remains unresolved.

The secondary re-rating is broader than the chip stack. Utilities and independent power producers with contractable capacity have been repriced as AI counterparties — the restarts and long-term nuclear power agreements signed by US hyperscalers since 2024 mark the template — alongside grid-equipment makers with multi-year transformer and turbine backlogs, liquid-cooling specialists, and the materials chain beneath packaging and memory. Scarcity travels down the dependency tree until it finds the next thing that cannot be expanded quickly.

The limits of the AI compute build-out

Electricity is the hardest limit. The IEA's projection of data-centre consumption rising from about 485 TWh in 2025 to roughly 950 TWh by 2030 collides with interconnection queues, transformer lead times and local politics; gigawatt-scale campuses increasingly plan onsite generation. Power now gates deployment schedules in several major markets more tightly than chip supply does. Water and heat rejection follow directly: with reference designs reaching 140–142 kW per rack, cooling design is inseparable from compute planning, and jurisdictions are formalising efficiency requirements — Singapore conditions new capacity on PUE of 1.3 or better, a template other constrained grids are studying.

Geopolitics concentrates the risk. HBM supply is concentrated in three vendors: SK hynix and Samsung are based in South Korea, while Micron is US-headquartered and its manufacturing footprint spans several countries. Leading-edge logic and much advanced packaging remain concentrated in East Asia; lithography depends on one Dutch firm. US export controls have tightened and mutated — including HBM controls announced in December 2024 — while China accelerates domestic alternatives behind the wall. A globally portable software stack can be attached to legally non-portable hardware; export classification now belongs in the architecture decision, not the appendix.

Obsolescence compounds quietly. An annual silicon cadence can compress residual values: commodity H100 rental rates have fallen substantially from early scarcity peaks, but resale and rental curves differ by chip, region, contract, software support and available power. Fleets financed with debt against rapidly ageing accelerators are therefore the system's most fragile node — the place a demand pause or refinancing shock would surface first. A two-to-three-year value halving should be treated as a downside scenario, not a universal depreciation rule.

Efficiency is the deepest uncertainty. The January 2025 DeepSeek release — frontier-adjacent capability claimed at a fraction of assumed training cost — erased roughly US$589 billion of NVIDIA's market value in a single session, the largest one-day loss on record, precisely because it attacked the scarcity premise. Epoch AI estimates that historical algorithmic progress halved the pretraining compute needed to reach a given performance level about every eight months, with a wide confidence interval; that finding does not cover every inference or multimodal workload. The evidence so far is consistent with a rebound effect: Google's token volumes grew sevenfold in the year to May 2026 while serving costs fell, so aggregate use expanded faster than unit cost declined. The counterfactual is unknowable, however, and Jevons is not a law. If token demand saturates while efficiency keeps compounding, compute scarcity — and the valuations priced on it — would unwind faster than supply chains could adjust. Alternative substrates (photonic interconnect is commercialising now; analog and neuromorphic approaches remain outside frontier-scale deployment) are worth monitoring but are unlikely to relieve the 2026–2028 bottlenecks.

AI compute outlook 2026–2030: base, bull and bear

Any honest outlook is a function of three variables: how fast token demand compounds, how fast efficiency improves, and how much power gets connected. The scenarios below hold supply-chain behaviour constant and vary those three.

ScenarioCompute demandHardware investmentValuation regime
Base — scarcity normalises slowlyToken volumes keep growing but decelerate from ~7× toward 2–3× a year; inference dominates the mix by 2028Hyperscaler capex grows through 2027, then plateaus at a structurally higher level; memory stays tight into late 2027 (Gartner sees no broad relief sooner), then normalisesEarnings growth does the work; multiples drift down; NVIDIA's share erodes toward 60–70% of a much larger market; power-advantaged operators outperform
Bull — the agentic decadeAgentic and physical-AI workloads hold demand growth at 5×+ a year; sovereign programmes add price-insensitive buyers; realised demand keeps beating forecastsAnnual AI infrastructure spending passes US$1 trillion before 2030; HBM supply stays substantially committed and packaging remains capacity-constrained; the power chain re-rates againScarcity rents persist; the hardware stack compounds; the equity story broadens from silicon to electrons
Bear — absorption and disciplineEfficiency gains plus open-weight commoditisation collapse token prices faster than volumes grow; enterprise AI revenue disappoints; one major buyer cuts firstCapex discipline arrives in 2027; memory overshoots into a 2028 glut, as memory always eventually has; leveraged GPU fleets default; second-hand accelerators overhang the marketMultiple compression of 30–50% across the chain even where revenue holds — the post-2000 Cisco lesson, softened by today's far lower starting multiples and contracted backlogs

The base case deserves the most weight because both tails are already visibly restrained: realised demand remains strong against the bear, while financing markets and boards began interrogating the capex-to-revenue gap in 2026 rather than 2028 against the bull. The inflection points worth watching are concrete: N2P and A16 ramp allocation in late 2026; HBM4 yields and the HBM4E step in 2027; whether facilities can host 140 kW-class racks and multi-rack Rubin Ultra domains; the inference-versus-training mix in hyperscaler disclosures; token price indexes set against volume growth; dated grid-interconnection wins; depreciation-schedule revisions; and the maturity of Chinese domestic accelerators. Movements there will reveal the cycle's direction two to four quarters before headline revenue does.

A Singapore lens on the compute build-out

Singapore illustrates how a small-grid jurisdiction plays a compute-constrained era: by specialising rather than competing on gigawatts. Domestic data-centre capacity of roughly 1.4 GW grows only under the IMDA Green Data Centre Roadmap's rationed allocations — at least 300 MW, plus more for operators bringing green power, conditional on PUE ≤ 1.3 — which channels new builds toward dense, liquid-cooled AI halls. A reasonable regional forecast is that some training-scale capacity will locate in Johor's larger corridor or Batam while latency-sensitive inference remains in Singapore, but that is an outlook rather than an observed allocation rule. Upstream, the position is stronger than the grid suggests: Singapore produces about a fifth of global semiconductor equipment and roughly one in ten chips (EDB figures), and Micron's US$7 billion HBM advanced-packaging facility — operational from 2026, adding meaningful capacity from 2027 — places the country inside the very chokepoint this essay describes. For Singapore enterprises, the practical reading is that local AI capacity will stay scarce and priced accordingly; the accessible margin lies in efficient inference, cross-border training placement and the supply chain itself.

The bottom line: a currency, priced accordingly

Strip the cycle to its mechanics and the currency framing holds without embellishment. AI compute is scarce; it is metered and sold in standard units; it is contracted years forward; it is accepted as collateral; and it converts directly into a revenue-bearing product. Every layer of the hardware economy — accelerators, memory, wafers, packaging, racks, networks, cooling, electrons — is now priced off access to it, and the great re-ratings of 2023–2026 are, at core, the market capitalising the rents that scarcity creates. Those rents are real but conditional: conditional on demand continuing to outrun remarkable efficiency gains, on power arriving, on geopolitics holding, and on the accounting surviving contact with annual hardware cadences. The signals that will announce any turn are already published quarterly — HBM contract terms, packaging throughput, paid utilisation, interconnection dates. Until they move, compute remains what the past three years made it: the defining scarce resource, and therefore the defining currency, of the present AI cycle.

Frequently asked questions

What does it mean to call AI compute 'the new currency'?

It is a precise analogy, not a slogan. AI compute is scarce, metered in standard units (GPU-hours, tokens, FLOPs), contracted years forward like power purchase agreements, accepted by lenders as collateral, and directly convertible into a revenue-bearing product. Access to it — allocation — has become the negotiating instrument that decides which AI businesses can operate at scale, which is functionally how a currency behaves.

Why can't the industry simply build more AI compute?

Because supply is a serial chain with four hard chokepoints: high-bandwidth memory (three suppliers, with substantial volumes committed under long-term agreements), capacity-constrained advanced packaging, leading-edge wafers (TSMC holds about 73% of pure-play foundry revenue), and grid power (interconnection queues run years). Adding capacity at any one layer takes years of cleanrooms, tools and qualification — and delivered capacity is set by the tightest layer, not the average one.

How much electricity does AI compute consume?

The International Energy Agency estimates data centres consumed about 485 TWh in 2025 and projects roughly 950 TWh by 2030, with AI the dominant growth driver. Current high-power AI reference designs reach roughly 140–142 kW per rack versus 8–15 kW for legacy enterprise racks; next-generation systems also combine multiple racks into one compute domain. That is why liquid cooling and grid access have become procurement criteria alongside chip allocation.

Are AI hardware valuations a bubble?

The evidence is mixed and the market is already discriminating. Supporting the valuations: contracted multi-year demand (HBM long-term agreements, cloud backlogs), extraordinary realised margins, and demand that has beaten every forecast since 2023. Against them: the widening gap between AI capital spending and AI revenue, depreciation schedules that may overstate earnings if hardware ages faster than five to six years, and single-day events like the January 2025 DeepSeek selloff showing how much scarcity premium is embedded. Record results that still miss expectations — SK hynix's 76%-margin quarter greeted with an 11% share-price fall — show expectations, not just fundamentals, are stretched.

Will efficiency gains reduce demand for AI compute?

The evidence is consistent with a rebound effect, not proof of a universal rule. Historical algorithmic progress sharply reduced pretraining compute for a given measured performance, while Google's token volumes grew about sevenfold in the year to May 2026 as serving costs fell. Aggregate use therefore rose faster than unit cost declined at Google, but the counterfactual is unknowable. The bear case for the hardware stack is that token demand saturates while efficiency keeps compounding.

What does the AI compute race mean for Singapore businesses?

Three practical things. Local AI capacity will remain rationed and premium-priced — new data-centre allocations under the Green Data Centre Roadmap are conditional and total only hundreds of megawatts — so efficient inference and workload placement matter more than raw capacity. Some training-scale capacity may increasingly locate in Johor or Batam while latency-sensitive inference stays in Singapore; that is a regional forecast, not a settled allocation rule. And the supply chain runs through the country: about a fifth of global semiconductor equipment is made here, and Micron's US$7 billion HBM packaging plant puts Singapore inside the industry's tightest bottleneck, creating opportunities well beyond running models.

Sources and further reading