What changed in this guide
- — Added an original diagram
- — Published
Executive Summary
GPU procurement is a hardware-standardisation and ownership decision before it is a vendor decision. The two questions that determine cost and risk are which accelerator you standardise on, and whether you rent it by the hour, reserve it for years, or own it in a colocation hall. For most Singapore enterprises in 2026 the correct answer to the second question is still to rent — but the organisations that get the first question wrong overpay for years, because an accelerator generation commits you to a memory ceiling, a power envelope and a software ecosystem long after the invoice cadence has been forgotten.
The silicon question has a small number of real answers. NVIDIA's H100 and the memory-expanded H200 are the mature workhorses; the Blackwell generation (B200, and the rack-scale GB200 NVL72) is where new large-scale deployments are heading; AMD's Instinct MI300X, MI325X and MI355X are a credible second source on memory capacity, with a maturing software stack. The right choice depends on model size, context length, the training-versus-inference mix and — increasingly — how many kilowatts and how much liquid cooling you can actually get. It does not depend on peak TFLOPS.
The ownership question turns on one number: utilisation. Owned or long-reserved capacity beats on-demand rental only when it runs busy — as a rule of thumb, sustained utilisation above roughly 60–70% over a two-to-three-year horizon. Real enterprise utilisation on shared clusters, absent mature scheduling, frequently lands far below that. This guide gives you the silicon comparison, the build-versus-buy break-even, the hosting-option trade-offs, a total-cost-of-ownership model that counts power and cooling rather than just the sticker rate, and the RFP and contract terms that decide whether the deal survives the next silicon generation. For the parallel decision — which cloud provider, and what its published rate cards actually say — see the companion AI Cloud Computing in Singapore guide.
Key Takeaways
- Choose the accelerator before the vendor. The GPU generation you standardise on sets your memory ceiling, power envelope and software ecosystem. It is the decision most likely to be made by default and regretted at scale.
- Most enterprises should rent, not own. Renting inference or on-demand GPU capacity avoids capacity risk, cluster operations and idle-time waste. Ownership needs a specific trigger: proven high utilisation, an isolation or residency requirement, real custom training, or a measured cost case.
- Utilisation decides build-versus-buy, not unit price. Owned capacity pays only above roughly 60–70% sustained utilisation over two to three years. Measure a full quarter of production load before converting variable cost into fixed cost.
- You are procuring kilowatts and cooling, not just chips. AI racks can exceed 40–100 kW and often need liquid cooling. In Singapore, power and data-centre capacity — not silicon allocation — are frequently the binding constraint.
- Memory and bandwidth matter more than peak FLOPS. For large-model and long-context inference, HBM capacity and bandwidth decide how many GPUs a model needs and how fast it runs. Peak throughput figures rarely predict real cost per token.
- Never compare a networked cluster SKU to a single-GPU SKU. Cluster SKUs bundle InfiniBand-class fabric and cost materially more per GPU-hour. That premium is either necessary for distributed training or wasted — it is never neutral.
- Software portability is a procurement term. NVIDIA CUDA is the default; AMD ROCm is a real alternative that can need porting effort. Treat framework, kernel and toolchain portability as a negotiated requirement, not an afterthought.
- Treat Singapore's AI governance frameworks as procurement inputs. IMDA's Model AI Governance Frameworks, CSA's Guidelines on Securing AI Systems and AI Verify are voluntary, but large buyers and regulated counterparties use them as de facto requirements. The PDPA and, for CII, the Cybersecurity Act are binding.
- Run counterparty diligence on specialist providers. Capital intensity, revenue concentration and pivots from other industries are legitimate procurement considerations for any multi-year GPU commitment.
- Negotiate refresh, egress and exit before signing. Generation-upgrade rights, capacity reallocation, data and fine-tuned-weight export, and exit assistance are cheap at contract and expensive to retrofit — in a market where the chip changes about once a year.
Quick Facts
| Question | Practical answer for buyers |
|---|---|
| What is being procured? | An accelerator standard (which GPU), an ownership model (rent, reserve or own), and the physical envelope to run it (power, cooling, space, network). |
| Default enterprise choice in 2026 | Rent — inference APIs or on-demand/reserved cloud GPUs — until utilisation and use case are proven. Own only against a measured trigger. |
| Mainstream accelerators | NVIDIA H100 (80 GB), H200 (141 GB), B200 (Blackwell, up to 192 GB); AMD Instinct MI300X (192 GB), MI325X (256 GB), MI355X (288 GB). |
| Best unit of comparison | Cost per unit of useful work — per million tokens or per training run — not cost per GPU-hour or peak TFLOPS. |
| Build-versus-buy break-even | Owning tends to win only above ~60–70% sustained utilisation across a 2–3 year horizon; below that, on-demand or reserved rental usually wins. |
| Singapore in-country capacity | Hyperscaler SG regions (AWS ap-southeast-1, Azure southeastasia, Google Cloud asia-southeast1, Oracle Singapore), Singtel RE:AI, Firmus/SMC Cloud, and private GPU in Singapore colocation. |
| Blackwell in Singapore | Scarce in-country as at August 2026 — newest silicon concentrates in a few mostly-US regions; H100/H200 are more readily available locally. |
| The real local constraint | Power and cooling. AI racks run 40–100+ kW; data-centre vacancy is below ~1.4% and new capacity is state-rationed. |
| Binding Singapore law | PDPA (including the Transfer Limitation Obligation) and, for designated critical information infrastructure, the Cybersecurity Act. |
| Voluntary but expected | IMDA Model AI Governance Framework for Agentic AI (January 2026, updated 20 May 2026); CSA Guidelines on Securing AI Systems; AI Verify; ISO/IEC 42001; NIST AI Risk Management Framework. |
| Financial-sector direction | MAS consulted on proposed Guidelines on AI Risk Management for financial institutions; consultation closed 31 January 2026. |
| Singapore funding | Enterprise Compute Initiative: up to SGD 150 million; consulting capped at SGD 150,000 with 70% funded up to SGD 105,000; separate cloud credits published by participating providers. |
What GPU Procurement Involves
GPU procurement is the decision about which AI accelerators to standardise on and how to obtain them. It sits one level below AI strategy and one level above a cloud contract. It is not the same as buying cloud computing: renting cloud GPU capacity is one option inside GPU procurement, alongside reserving capacity for years and owning hardware in a colocation facility. The distinguishing questions are hardware selection (which chip), ownership model (rent, reserve or own), and the physical envelope — power, cooling, space and networking — that a chosen accelerator demands.
The access-to-ownership spectrum
Most procurement disappointment comes from treating this as a single buy. It is a spectrum of five rungs, each with different economics, lock-in and operational burden. The correct default is the highest rung that meets the requirement — the one with the least commitment — and you descend only when a specific trigger justifies it.
| Rung | What you obtain | Priced by | Who it suits | Principal risk |
|---|---|---|---|---|
| 1. Inference API | Access to hosted models; no infrastructure. | Per million tokens or per request. | The majority of enterprises, for the majority of use cases. | Model deprecation and per-token cost growth at scale. |
| 2. Managed ML platform | Training and serving on a managed stack. | Platform fee plus compute. | Teams with data scientists but no platform engineers. | Platform-specific abstractions that are costly to unwind. |
| 3. On-demand GPU | GPU instances rented by the hour. | Per GPU-hour or node-hour. | Bursty training, experimentation, variable inference. | Headline rate is high; idle time is billed. |
| 4. Reserved / committed capacity | GPU blocks reserved for 1–3 years. | Committed term, discounted rate. | Predictable, sustained workloads. | Fixed cost if strategy, model or generation changes. |
| 5. Owned / colocated | Hardware you buy and run in colo or on-prem. | CapEx plus power, cooling, space, staff. | High, steady utilisation with residency or isolation needs. | Capacity, power, refresh and operations risk transfer to you. |
The rungs are not a maturity ladder to climb. A large, sophisticated buyer may correctly sit on rung 1 for most workloads and rung 5 for one. The mistake is assuming that scale or ambition requires ownership; ownership is a response to measured utilisation and specific constraints, not a sign of seriousness.
Who should use this guide
Use it if you are selecting an accelerator standard, sizing a first AI-infrastructure budget, deciding between renting and owning GPUs, evaluating colocation for AI, or writing a GPU or AI-infrastructure RFP. It complements two companion guides: for the cloud-provider decision, published rate cards and the residency shortlist, use AI Cloud Computing in Singapore; for the facility decision — power, cooling, colocation pricing and operators — use Data Centres in Singapore.
Do not use it as an AI strategy document. If the use case, data readiness, success metric and accountable owner are not defined, no hardware decision will rescue the project. For background on how the hardware actually works, see the Knowledge Base explainers on how GPUs work for AI and NVIDIA's AI infrastructure stack.
Why It Matters: Capacity and Power, Not Software, Are the Constraint
AI infrastructure procurement has become a board-level decision because the scarce inputs are physical, not digital. For most of the cloud era, compute was effectively unlimited and capacity questions were pricing questions. Dense AI workloads broke that assumption. The scarce resources are now accelerators, high-speed networking, power and cooling — and in Singapore, the last two bind first.
The four largest United States hyperscalers guided to combined capital expenditure approaching USD 700 billion for 2026, well above 2025 levels. Estimates vary by source and by what is counted, but the direction is consistent: capital at that scale is being deployed to secure power, land, accelerators and networking, not to win a price war. Three consequences follow, and each changes how an enterprise should buy.
- Availability is negotiated, not self-service. The newest accelerators are allocated, not listed. Region, generation, quantity and lead time are contract terms you must secure in writing, not assumptions you can make from a pricing page.
- Silicon depreciates fast. Accelerator generations turn over roughly every 12–24 months. A multi-year ownership or reservation decision is a bet that this generation stays economic for the life of the commitment.
- Power is the ceiling. A cluster you can pay for is not a cluster you can necessarily power and cool — especially in a supply-constrained market. Facility readiness, not purchase order, sets the true lead time.
Singapore Market Landscape
In Singapore, the decisive constraint on GPU procurement is power and in-country capacity, not price. Singapore is one of Asia-Pacific's most connected data-centre hubs and one of its most supply-constrained. More than 1.4 GW of capacity is operational, vacancy sits below roughly 1.4%, and data centres already consume a significant share of national electricity — commonly estimated at around 7% today, with projections toward 12% by 2030. That is why capacity is actively rationed rather than freely available, and why an AI cluster's true lead time is often set by megawatts, not by silicon allocation.
Capacity, power and the state-rationed pipeline
After a moratorium on new data centres from 2019 to 2022, Singapore releases capacity through competitive, sustainability-conditioned government calls. The 2024 Green Data Centre Roadmap set out at least 300 MW of near-term capacity with more available for projects using green energy and meeting efficiency targets. The second Data Centre Call for Application (DC-CFA2) opened at least 200 MW, required at least 50% green energy, and closed to applications on 31 March 2026, with awards pending. A roughly 700 MW low-carbon data-centre park has been announced for Jurong Island. For AI buyers this means dense capacity must be secured early and is increasingly planned as a two-market strategy, with Johor and Batam absorbing bulk workloads while Singapore keeps the latency-critical, regulated core. The facility-level detail is covered in Data Centres in Singapore.
Regulatory drivers and compliance context
Singapore's approach separates binding law from voluntary-but-expected guidance. The Personal Data Protection Act (PDPA), administered by the PDPC, is binding and imposes the Transfer Limitation Obligation on overseas data transfers; the Cybersecurity Act binds designated critical information infrastructure. On top of these, IMDA and GovTech publish frameworks that large buyers treat as procurement requirements: the Model AI Governance Framework family, including the version for Agentic AI first published in January 2026 and updated on 20 May 2026, and the AI Verify testing framework. For financial institutions, MAS consulted on proposed Guidelines on AI Risk Management, with the consultation closing on 31 January 2026; once finalised these operate as supervisory expectations. What binds and what is voluntary is set out in Singapore AI Regulations.
Funding, maturity and challenges
The Enterprise Compute Initiative, administered by Digital Industry Singapore, pairs eligible companies with participating cloud providers: up to SGD 150 million in total, consulting capped at SGD 150,000 per company with the government funding 70% up to SGD 105,000, and separate cloud credits published by each provider. It can offset a meaningful share of first-year cost — but credits expire while an accelerator standard persists, so it should not choose the architecture. Adjacent grants are covered in IT & Digital Grants in Singapore. On maturity, Singapore has strong foundations — connectivity, regulation, talent density — but many enterprises are ahead on governance frameworks and behind on production operations: utilisation discipline, scheduling and cost allocation. The recurring challenges buyers should expect are constrained in-country capacity for the newest silicon, high power costs, a shortage of GPU-cluster operations skills, and vendor claims that outrun in-region availability.
GPU Silicon Compared: Choosing the Accelerator
Choose silicon by memory capacity, memory bandwidth and workload fit — not by peak TFLOPS. For large-language-model training and inference, the amount and speed of high-bandwidth memory (HBM) usually decides how many GPUs a model needs, how long its context can be, and how fast it responds. NVIDIA dominates enterprise AI silicon and its CUDA software ecosystem is the default; AMD is a credible second source on memory capacity; Google and AWS offer their own accelerators inside their clouds as alternative supply lines. The table below summarises the mainstream 2026 options.
| Accelerator | Architecture | Memory | Bandwidth | Notable | Best-fit workload |
|---|---|---|---|---|---|
| NVIDIA H100 | Hopper | 80 GB HBM3 | ~3.35 TB/s | FP8; mature, widely available; broadest software support. | Established training and inference; the safe default for known workloads. |
| NVIDIA H200 | Hopper | 141 GB HBM3e | ~4.8 TB/s | Same compute as H100 with far more memory and bandwidth. | Larger models and longer context without adding GPUs; memory-bound inference. |
| NVIDIA B200 | Blackwell | up to 192 GB HBM3e | ~8 TB/s | Native FP4/FP6; higher throughput and efficiency; scarcer, higher power. | New large-scale training and high-volume inference. |
| NVIDIA GB200 NVL72 | Blackwell (rack) | ~13.4 TB aggregate | NVLink domain | 72 Blackwell GPUs + 36 Grace CPUs in one liquid-cooled, NVLink-connected rack. | Frontier training and trillion-parameter-class models. |
| AMD Instinct MI300X | CDNA3 | 192 GB HBM3 | ~5.3 TB/s | Large memory per GPU; ROCm software stack. | Memory-bound inference; multi-vendor and cost-sensitive strategies. |
| AMD Instinct MI325X | CDNA3 | 256 GB HBM3e | ~6.0 TB/s | Memory-capacity upgrade over MI300X. | Very large models served on fewer accelerators. |
| AMD Instinct MI355X | CDNA4 | 288 GB HBM3e | ~8 TB/s | Adds FP4/FP6; competes with Blackwell on memory and data types. | Large-scale training and inference where ROCm is validated. |
Specifications are vendor-published and simplified for comparison; SXM and PCIe variants differ. A newer NVIDIA Blackwell Ultra tier (B300/GB300), ramping through 2026, raises memory further (around 288 GB per GPU), and AMD's CDNA-Next MI400 is expected in 2026. Verify the exact SKU, form factor and regional availability with the provider.
How to read the table for your workload
- Model size and context length → memory. If a model or its context does not fit in one GPU's HBM, it must be sharded across GPUs, adding cost and interconnect dependency. More memory per GPU (H200, MI325X, MI355X) can shrink the cluster a given model needs.
- Training versus inference → different priorities. Distributed training is sensitive to interconnect bandwidth and favours NVLink/InfiniBand-connected clusters; high-volume inference is often memory-bandwidth- and capacity-bound and can be the stronger case for large-memory parts.
- Energy and cooling → a real constraint. Blackwell-class and MI355X parts draw more power and increasingly assume liquid cooling. In a power-constrained facility, the accelerator you can cool may matter more than the one you can buy.
The software dimension: CUDA, ROCm and lock-in
Accelerator choice is also a software choice. NVIDIA CUDA and its libraries are the default target for most frameworks, kernels and tooling, which is a genuine advantage and a genuine lock-in. AMD ROCm has matured and is supported by major frameworks, but porting and performance-tuning can require engineering effort, and some third-party kernels assume CUDA. Google TPUs and AWS Trainium and Inferentia are competitive within their own clouds but tie you to that platform's stack. Treat framework, kernel and toolchain portability as a procurement requirement: the ability to move a workload to a second accelerator is worth negotiating for, even if you never exercise it.
Build, Buy or Rent: The Ownership Decision
The ownership decision is an economics decision, and utilisation is the deciding variable. Enterprises typically choose among three paths, and the right one depends on how busy the capacity will actually be, not on how large the ambition is.
- Build (own, on-premises or colocation): full control over hardware, data residency and configuration. High capital expenditure, long lead times (often 12–24+ months for multi-megawatt capacity), and ongoing responsibility for power, cooling, networking and operations.
- Buy (public cloud or specialist GPU hosting): fast time-to-capacity, elastic scaling and managed services. Higher unit cost, potential lock-in, and data-sovereignty considerations to check per region.
- Partner / hybrid: reserved capacity with a hyperscaler or specialist for training bursts, combined with owned or dedicated capacity for steady-state inference or regulated data. The common answer for larger buyers.
Utilisation is the decisive economic variable. On-demand GPU-hours are expensive per unit; owned or multi-year reserved capacity is cheaper per unit only if it runs busy. As a rule of thumb, sustained utilisation above roughly 60–70% over an 18-to-36-month horizon favours ownership or committed capacity; below that, on-demand or short-reserved rental usually wins. The trap is that realistic enterprise utilisation on shared clusters, without mature scheduling, queueing and right-sizing, often lands in the 15–40% range — which quietly makes owned capacity more expensive per unit of useful work than rental at a higher headline rate. Baseline your actual utilisation before any large commitment.
Decision matrix
| Factor | Own / colocate | Hyperscaler cloud | Specialist (neocloud) | Marketplace / spot |
|---|---|---|---|---|
| Cost at high utilisation | Lowest long-term | Highest on-demand | Competitive | Lowest short-term |
| Time to capacity | Months–years | Hours–weeks (quotas) | Minutes–days | Minutes |
| Compliance / sovereignty | Full control | Strong, certified | Good, improving | Variable |
| Scalability | Planned | Elastic | Elastic within fleet | Highly elastic |
| Interconnect performance | Custom | Good | Often excellent (IB focus) | Variable |
| Operational burden | High | Low | Medium | Low–medium |
| Singapore in-region availability | Yes (colo) | Yes | Rare (as at Aug 2026) | Rare |
GPU-Hosting Options Compared
The hosting market divides into four tiers, and for Singapore buyers the shortlist is shortened first by in-country availability, then by fit. This section frames the tiers through the procurement lens; for provider-by-provider depth, published rate cards and the residency finding, see AI Cloud Computing in Singapore.
- Hyperscalers (AWS, Microsoft Azure, Google Cloud, Oracle Cloud Infrastructure). Broadest ecosystems, deepest compliance certifications, global reach, and managed ML platforms. All operate Singapore regions with H100-class capacity. Best for regulated workloads, teams already embedded in a hyperscaler, and buyers who value managed services over raw density. On-demand high-end GPU pricing is highest, but reserved and committed rates can undercut specialist on-demand.
- Specialist GPU clouds / neoclouds (CoreWeave, Lambda, Crusoe, Nebius and peers). Purpose-built for GPU-dense workloads: bare-metal or Kubernetes-native clusters, high-speed interconnects, rapid provisioning and simpler pricing. The critical caveat for Singapore buyers is that most have no Singapore region as at August 2026 — CoreWeave's first Asia-Pacific sites were announced for Indonesia, expected online in 2028. Strong for sustained training and inference at scale where in-country hosting is not required.
- Singapore specialists (Singtel RE:AI, Firmus / SMC Cloud). In-country GPU-as-a-service from Singapore data centres — the practical route to specialist-style capacity under a residency constraint. Confirm accelerator generation and scale against your requirement.
- Marketplaces and spot providers (RunPod, Vast.ai and others). Lowest headline prices and high flexibility, but variable reliability, limited SLAs and less enterprise support. Suitable for experimentation and non-critical burst work, not production-critical clusters.
Pricing Models and Total Cost of Ownership
The sticker rate — per GPU-hour or per server — is a fraction of total cost of ownership. A defensible GPU cost model counts the inputs that dominate at scale: power and cooling, networking, storage, staff, and the refresh cycle. Modelling only the hardware or the hourly rate is the most common way GPU business cases go wrong.
| Cost line | What it covers | Why it is often missed |
|---|---|---|
| Accelerator / server | The GPUs and the host system (CPU, memory, chassis). | Treated as the whole cost; it is often under half of a five-year TCO for owned kit. |
| Power | Electricity for compute at 40–100+ kW per rack, billed continuously. | High in Singapore; scales with utilisation and never sleeps for owned capacity. |
| Cooling | Liquid or advanced air cooling for dense racks; PUE overhead. | Newer or retrofitted facilities only; a capacity and cost constraint, not a detail. |
| Networking | InfiniBand or high-speed Ethernet fabric for distributed training. | Cluster SKUs bundle it; comparing against single-GPU SKUs understates cost. |
| Storage & egress | Parallel/high-throughput storage; data egress on cloud. | Egress and idle storage accumulate quietly across a project. |
| Staff & software | Platform engineering, MLOps, scheduling, monitoring, licences. | Ownership transfers this burden to you; scarce and expensive locally. |
| Refresh & residual | Replacement every ~2–3 years; resale/residual value. | GPUs depreciate fast; a multi-year bet on one generation staying economic. |
| Idle / utilisation | The gap between capacity paid for and work done. | The single largest hidden cost; low utilisation inverts the whole business case. |
A defensible cost model in five steps: (1) forecast the workload in units of useful work — tokens per month, training runs per quarter; (2) measure or estimate utilisation honestly, sized to the low case; (3) build the fully loaded cost for each option including power, cooling, network, storage, staff and refresh; (4) express the result as cost per unit of useful work, not per GPU-hour; (5) re-run the model when the next accelerator generation or a price change lands. The winner frequently changes between steps (1) and (4).
Evaluation Framework for Enterprise Buyers
Score providers on a weighted scorecard built from your workload, not on headline price. Characterise the workload first — model sizes, context lengths, training-versus-inference mix, growth, residency and latency needs — then filter on in-region availability, then score the survivors.
Buying criteria and weights
- Availability and lead time (gate first): accelerator generation, region, quantity and committed delivery date.
- Architecture and performance fit: measured throughput and latency on your workload, interconnect, storage performance.
- Cost behaviour at scale: fully loaded cost per unit of useful work across commitment models, not per GPU-hour.
- Security and compliance: certifications with scope and reporting period; data-processing and residency terms.
- Portability and exit: data, weight and log export; software portability; reallocation and generation-upgrade rights.
- Commercial durability: supplier financial strength, revenue concentration, references at your size and sector.
Vendor evaluation questions
Common mistakes
- Selecting silicon on peak TFLOPS instead of memory and real workload fit.
- Committing to capacity before utilisation is measured in production.
- Comparing a networked cluster SKU against a single-GPU SKU on price per GPU-hour.
- Underestimating power, cooling and facility lead times — the true constraint locally.
- Accepting CUDA lock-in without pricing the loss of portability.
- Signing multi-year terms without refresh, reallocation and exit rights.
Compliance, Security and Risk Management
Governance obligations follow your data and your use case, not the location of the GPU. Map both the binding rules and the voluntary frameworks buyers are expected to meet, and treat provider certifications as support for your position rather than a replacement for it.
- Binding law. The PDPA, including the Transfer Limitation Obligation, governs personal data and overseas transfer; the Cybersecurity Act governs designated critical information infrastructure. For financial institutions, MAS Technology Risk Management and outsourcing requirements apply, and the proposed MAS Guidelines on AI Risk Management (consultation closed 31 January 2026) will operate as supervisory expectations once finalised.
- Voluntary but expected. IMDA's Model AI Governance Framework for Agentic AI (January 2026, updated 20 May 2026), CSA's Guidelines on Securing AI Systems, and the AI Verify testing framework are widely used as procurement requirements by large buyers and regulated counterparties.
- Certifications to verify by scope. MTCS SS 584 (Singapore's multi-tier cloud security standard), ISO/IEC 27001, ISO/IEC 42001 (AI management systems), SOC 2, and where relevant ISO/IEC 27017 and 27018. Check the services covered and the reporting period, not just the logo.
- Frameworks for your own governance. The NIST AI Risk Management Framework and ISO/IEC 42001 give structure to your AI inventory, risk classification, testing evidence and oversight design.
- Supply-chain and counterparty risk. Several specialist providers are capital-intensive, financed against long contracts with a few large customers, and some entered from other industries. Run standard supplier due diligence: financial durability, revenue concentration, contract seniority, data-export rights, and what happens to your workload if the provider is restructured.
Implementation Roadmap and Pitfalls
Sequence the work so that the irreversible commitments come last. Prove the workload and the utilisation before you convert variable cost into fixed cost, and negotiate the exit before you onboard.
| Phase | Indicative timing | What happens |
|---|---|---|
| 1. Characterise & baseline | Weeks 0–6 | Define workloads, residency and latency needs; measure current utilisation; forecast in units of useful work. |
| 2. RFP & shortlist | Weeks 4–10 | Filter on in-region availability, then issue an evidence-based RFP; score against weighted criteria. |
| 3. Proof of concept | Weeks 8–16 | Run on your actual models; measure throughput, latency, failure/restart, storage, egress and cost per unit of useful work. |
| 4. Contract & commit | Weeks 14–22 | Negotiate commitment sized to the low case, with refresh, reallocation, egress and exit terms. |
| 5. Deploy & operate | Ongoing | Stand up scheduling, queueing, monitoring and cost allocation; review utilisation quarterly. |
Contract terms to secure
Failure modes to design against
- Buying ahead of proof. Reserving capacity before a production utilisation baseline exists is the most expensive and most common failure.
- Power and cooling surprise. The facility, not the purchase order, sets the true lead time; discover this before signing, not after.
- Generation stranding. A multi-year commitment to one accelerator with no upgrade right ages badly as the annual cadence continues.
- Utilisation neglect. Without scheduling and right-sizing, owned capacity idles and the business case inverts.
- Silent lock-in. Framework and kernel choices that assume one accelerator's software stack make a later switch costly.
Future Outlook: 2026–2029
Plan for a market where power, not silicon, is the visible constraint and inference is the dominant spend. The direction of travel is consistent across vendor roadmaps and market reporting, even where specifics remain uncertain.
- Power and land become the ceiling. Facility readiness, grid capacity and cooling — not chip allocation — increasingly determine what an enterprise can actually deploy, especially in Singapore.
- Inference overtakes training as the dominant cost. As models move into production, per-token inference spend grows faster than training, shifting the accelerator priority toward memory capacity and bandwidth.
- Rack-scale systems redefine the unit. Rack-scale designs such as GB200 NVL72 make the meaningful unit a networked rack, not a single GPU, changing how capacity and cost are compared.
- Accelerator diversity increases, and portability becomes a procurement term. AMD, Google TPUs, AWS silicon and others expand real choice; the ability to move workloads becomes a negotiated asset.
- Sovereign and regional models mature. Regional efforts such as AI Singapore's SEA-LION make in-region hosting and locally-governed models a more realistic option.
- Commercial structures mature, and concentration risk gets board attention. Multi-year GPU commitments, capital intensity and supplier consolidation move counterparty risk onto the board agenda.
Frequently Asked Questions
What is GPU procurement, and how is it different from buying cloud computing?
GPU procurement is the enterprise decision about which AI accelerators to standardise on and how to obtain them — rented by the hour, reserved for one to three years, or owned in a colocation facility. It is broader than buying cloud computing because it includes hardware selection, the ownership model, and the physical constraints of power, cooling and space. Buying cloud GPU capacity is one option within GPU procurement, not the whole of it.
Which NVIDIA GPU should an enterprise choose in 2026 — H100, H200 or B200?
The H100 (80 GB) is the mature, widely available option with strong price-performance for known workloads. The H200 adds memory (141 GB) and bandwidth, helping larger models and longer context without adding GPUs. The B200 (Blackwell, up to 192 GB) offers higher throughput and native FP4 for new large-scale training and high-volume inference, but is scarcer and draws more power. Choose by model size, context length and the training-versus-inference mix — not by peak throughput.
Is AMD Instinct a viable alternative to NVIDIA for enterprise AI?
Yes, for the right workloads. The MI300X (192 GB), MI325X (256 GB) and CDNA4 MI355X (288 GB) lead on memory capacity, which suits memory-bound inference and large models. The main consideration is software: CUDA remains the default, while ROCm has matured but can need porting effort. A second accelerator source improves availability and negotiating position — validate your own models on the target stack first.
Should we buy and own GPUs or rent them?
Most enterprises should rent, at least until demand is proven. Renting avoids capacity risk, cluster operations and idle-time waste. Owning becomes defensible with sustained high utilisation, a residency or isolation requirement rental cannot meet, substantial custom training, or a measured cost case at scale. Ownership also transfers power, cooling, networking, refresh and operations to you.
What utilisation level justifies buying GPUs instead of renting?
As a rule of thumb, owned or multi-year reserved capacity beats on-demand rental only above roughly 60–70% sustained utilisation over a two-to-three-year horizon, because GPUs depreciate in about two to three years and idle owned capacity still costs power, cooling and space. Many shared clusters run well below that. Measure a full quarter of production utilisation before committing, and size to the low case.
Can we buy the latest Blackwell GPUs in Singapore?
Not the newest ones easily. As at August 2026, Blackwell-class capacity concentrates in a few mostly-US regions, and specialist GPU clouds that list it generally have no Singapore region. In-country options are the hyperscaler Singapore regions (AWS, Azure, Google Cloud, Oracle), Singapore providers such as Singtel RE:AI and Firmus, or private GPU in Singapore colocation, where H100 and H200 are more available than Blackwell. Confirm generation, region and lead time in writing.
How much does an enterprise GPU server cost?
Hardware prices are volatile and quote-driven, and the unit that matters is a server, not a chip. An eight-GPU H100 or H200 system is a substantial capital item, and a production deployment adds high-speed networking, parallel storage, and the power and cooling for racks that can exceed 40–100 kW. Treat any single figure as indicative and build a total-cost-of-ownership model, not just a purchase price.
What is the real cost of running GPUs beyond the hardware?
The purchase price is only part of total cost of ownership. Dense AI racks draw far more power than traditional servers, so electricity and cooling are major recurring costs, amplified in Singapore by high power costs and constrained capacity. Add networking, storage, space, platform and operations staff, software, and a two-to-three-year refresh. Idle time is the hidden cost: capacity you own but do not use still consumes power, cooling and depreciation.
What Singapore rules and grants apply to AI infrastructure procurement?
The binding law is the PDPA, including the Transfer Limitation Obligation, and the Cybersecurity Act for critical information infrastructure. IMDA's Model AI Governance Frameworks (including for Agentic AI, January 2026, updated 20 May 2026), CSA's Guidelines on Securing AI Systems and AI Verify are voluntary but widely required. MAS consulted on proposed AI Risk Management Guidelines (consultation closed 31 January 2026). The Enterprise Compute Initiative pairs eligible companies with cloud providers, with consulting support and published cloud credits.
What should a GPU or AI-infrastructure RFP require?
It should require evidence, not adjectives: the specific accelerator generation, region and quantity with committed lead times; interconnect specification; whether pricing is on-demand, reserved or owned, with commitment terms; the power density and cooling method contractually delivered; data-processing and residency terms; security certifications with scope and reporting period; refresh and generation-upgrade rights; observability and cost allocation; support and escalation; and exit assistance with rights to export data, fine-tuned weights and logs.
What are the biggest GPU procurement mistakes?
Choosing silicon on peak throughput rather than memory and workload fit; buying or reserving before utilisation is proven; comparing a cluster SKU against a single-GPU SKU on price; underestimating power, cooling and facility lead times, which locally are often the true constraint; ignoring software portability and accepting avoidable lock-in; and signing multi-year commitments without refresh, reallocation and exit rights in a market where the chip changes about once a year.
Final Recommendations
Decide the accelerator standard first, then the ownership model, then the vendor. The chip sets your memory ceiling, power envelope and software ecosystem for years; make it a deliberate choice on memory, bandwidth and workload fit, not a default inherited from whatever was available.
Rent by default; own against a measured trigger. Start on the highest rung of the access-to-ownership spectrum that meets the requirement. Descend to reserved or owned capacity only when a production utilisation baseline, a residency or isolation requirement, real custom training, or a defensible cost case justifies it.
Model total cost of ownership, not the sticker rate. Count power, cooling, networking, storage, staff and refresh, and express the result as cost per unit of useful work. Utilisation will move the answer more than the headline rate does.
Secure the physical envelope early. In Singapore, power, cooling and in-country capacity are frequently the binding constraint. Confirm megawatts, cooling method and lead time before you plan around a chip.
Negotiate the exit while you have leverage. Refresh and generation-upgrade rights, capacity reallocation, egress terms, and data, weight and log export are cheap at contract and expensive to retrofit — and decisive in a market that re-bases roughly every year.
Primary Sources and Further Reading
Source links last checked 6 September 2026. This records that each link resolved, not that its content was re-read.
This guide was compiled and reviewed against the public sources below on 18 August 2026. Accelerator specifications, availability, pricing, capacity, regulatory guidance and contractual terms change frequently — sometimes within weeks. Every figure should be re-verified against the primary source, and performance validated on your own workload, before it informs a procurement, budget or compliance decision.
Accelerator vendors and specifications
- NVIDIA H100, H200 and GB200 NVL72 product and datasheet pages.
- AMD Instinct MI300X, MI325X and MI355X accelerator specifications.
- Google Cloud TPU and AWS Trainium and Inferentia for non-NVIDIA supply lines.
Singapore government and regulators
- IMDA — Artificial Intelligence in Singapore, including the Model AI Governance Framework for Agentic AI and AI Verify.
- AI Verify Foundation — AI Verify testing framework and toolkit.
- CSA — Guidelines and Companion Guide on Securing AI Systems.
- MAS — Consultation Paper on Proposed Guidelines on AI Risk Management (consultation closed 31 January 2026).
- PDPC — PDPA advisory guidelines, including on the use of personal data in AI systems.
- IMDA — Cloud computing standards, including MTCS SS 584.
- Digital Industry Singapore — Enterprise Compute Initiative and participating cloud service providers.
Market and infrastructure context
- Singapore's Green Data Centre Roadmap and the DC-CFA2 capacity allocation call.
- ISO/IEC 42001 — AI management systems and NIST AI Risk Management Framework.
- AI Singapore — SEA-LION regional language models.
- TechDirectory Knowledge Base: how GPUs work for AI, NVIDIA AI infrastructure and AI data-centre infrastructure.
Browse AI Infrastructure Providers in Singapore
TechDirectory lists directory records for AI computing companies, cloud providers, data-centre operators, system integrators and semiconductor firms serving Singapore. Profiles may show recorded capabilities, certifications and approved reviews where available. Verify accelerator availability, region, delivery capability, commercial relationship and references directly with the provider before contracting.
Browse AI Computing Providers →- AI Cloud Computing in Singapore — providers, published rate cards and the residency shortlist
- Data Centres in Singapore — power, cooling, colocation pricing and operators
- Large Language Models in Singapore — model selection, RAG and AI agents
- AI Computing in Singapore — grants, governance and vendors
- Singapore AI Regulations — what binds and what is voluntary
- Cloud Migration Cost and Vendor-Selection Frameworks
- Enterprise Cybersecurity in Singapore
- IT & Digital Grants in Singapore
- IT Procurement Templates — RFP, scorecard and diligence prompts