What changed in this guide
- — Replaced a dead MAS outsourcing source link
- — Published
Executive Summary
“AI cloud computing” is not a market. It is four separate purchases that vendors present as one. An enterprise can rent accelerated infrastructure by the GPU-hour, buy a managed training and machine-learning platform, consume model inference priced per token, or license an AI application platform with agents and connectors. These have different unit economics, different lock-in profiles, different regulatory exposure and different failure modes. Most procurement disappointment in this category traces back to running one evaluation across four products.
For Singapore buyers, the decisive constraint is rarely price. It is what can actually be provisioned in-country, in the accelerator generation you need, on the timeline you have. The three largest hyperscalers and Oracle operate Singapore regions with H100-class GPU capacity, and there are Singapore-based specialist providers. But the widely cited specialist “neoclouds” — CoreWeave, Nebius, Lambda and their peers — largely do not have Singapore regions as at August 2026. CoreWeave's first Asia-Pacific facilities were announced for Indonesia in August 2026 and are expected online in 2028. If your workload carries an in-country hosting constraint, the realistic shortlist is materially shorter than any global vendor comparison suggests.
This guide also corrects a claim that circulates widely in vendor and analyst commentary: that specialised GPU clouds are two to seven times cheaper than hyperscalers. Compared like-for-like on published rate cards, the gap for the same accelerator generation is closer to 1.1× to 3×, and hyperscaler multi-year reserved rates can undercut specialist on-demand rates. The larger multiples come from comparing hyperscaler on-demand list prices against the cheapest marketplace, spot or preemptible capacity — which is not the same product. The variable that dominates AI cloud cost is not the hourly rate at all. It is utilisation, and utilisation is an organisational capability, not a contract term.
Key Takeaways
- Decide which layer you are buying before you shortlist. Accelerated infrastructure, managed ML platform, inference API and AI application platform are four markets with four sets of competitors. Mixing them produces scorecards that cannot be scored.
- Most enterprises should not buy GPU capacity. Renting inference per token avoids capacity risk, cluster operations and idle-time waste. Direct GPU procurement needs a specific trigger: sustained high utilisation, an isolation or residency requirement an API cannot meet, real custom training, or latency that requires local placement.
- Region and SKU availability decides the shortlist, not the vendor logo. Blackwell-class capacity is concentrated in a small number of regions, mostly in the United States. Even within a single specialist provider, the newest accelerators are available in only one or two of its regions.
- Verify the price claim, then ignore the price. Published on-demand rate cards for the same accelerator generation cluster within a factor of about three. Utilisation, batching efficiency, storage throughput and egress typically move total cost more than the headline rate.
- Do not compare a cluster SKU against a single-GPU SKU. Cluster SKUs bundle InfiniBand-class fabric and cost materially more per GPU-hour. That premium is either necessary for your workload or wasted — it is never neutral.
- Treat Singapore's AI governance frameworks as procurement inputs even where they are voluntary. IMDA's Model AI Governance Framework family, CSA's Guidelines on Securing AI Systems and AI Verify are not binding law, but large buyers and regulated counterparties use them as de facto requirements.
- Financial services buyers should plan against MAS's proposed AI Risk Management Guidelines now. The consultation closed on 31 January 2026. Once finalised, these operate as supervisory expectations for all MAS-regulated institutions.
- Run supplier due diligence on specialist providers, not just technical evaluation. Capital intensity, revenue concentration in a handful of customers, and pivots from other industries are all legitimate procurement considerations for a multi-year commitment.
- Use the grant, but do not let it choose the architecture. The Enterprise Compute Initiative can offset a meaningful share of first-year cost, and participating providers publish credit allocations. Credits expire; the architecture persists.
- Negotiate exit before onboarding. Data export, fine-tuned weight portability, log access, model deprecation notice periods and capacity reallocation rights are cheap to secure at contract and expensive to retrofit.
Quick Facts
| Question | Practical answer for buyers |
|---|---|
| What is actually being bought? | One of four things: accelerated infrastructure (per GPU-hour), a managed ML platform (per platform unit plus compute), inference (per token or request), or an AI application platform (per seat, agent or action). |
| Best unit of comparison | Cost per unit of useful work — cost per million tokens, per training run, or per completed business transaction. Not cost per GPU-hour. |
| Published on-demand range | Roughly USD 2–14 per GPU-hour across accelerator generations and providers on August 2026 list prices; committed and spot rates sit materially below, capacity-reservation products above. |
| Singapore in-country GPU options | AWS ap-southeast-1, Azure southeastasia, Google Cloud asia-southeast1, Oracle Cloud Singapore, Singtel RE:AI (Nxera data centres), Firmus/SMC Cloud, plus private GPU deployments in Singapore colocation. |
| Specialist providers without a Singapore region | CoreWeave (first APAC sites announced for Indonesia, expected online 2028), Nebius (published regions in Finland, France, Israel, UK, US), Lambda, Crusoe, Nscale, RunPod, Together AI, Vast.ai — verify before assuming APAC availability. |
| Binding Singapore law | PDPA (including the Transfer Limitation Obligation for overseas transfers) and, for designated critical information infrastructure, the Cybersecurity Act. Sector rules may add more. |
| Voluntary but expected | IMDA Model AI Governance Framework for Generative AI (May 2024) and for Agentic AI (January 2026, updated 20 May 2026); CSA Guidelines on Securing AI Systems (October 2024); AI Verify testing framework and toolkit. |
| Financial-sector direction | MAS consulted on proposed Guidelines on AI Risk Management for all regulated financial institutions; consultation closed 31 January 2026. |
| Relevant certifications | MTCS SS 584:2020 (three tiers, 535 controls), ISO/IEC 27001, ISO/IEC 42001 (AI management system), SOC 2, ISO/IEC 27017 and 27018. Check scope, services covered and reporting period. |
| Singapore funding | Enterprise Compute Initiative: up to SGD 150 million total; consulting capped at SGD 150,000 per company with 70% government funding up to SGD 105,000; separate cloud credits published by participating providers. |
| Capacity context | Singapore data-centre capacity exceeds 1.4 GW. The second Data Centre Call for Application (DC-CFA2) opened at least 200 MW, applications closed 31 March 2026, with a requirement for at least 50% green energy. |
| Decision evidence to collect | Workload profile, utilisation forecast, region and SKU availability confirmation, proof-of-concept measurements, data-flow map, security evidence, scored RFP, commercial model with commitment terms, and an exit plan. |
What Is AI Cloud Computing?
AI cloud computing is the delivery of accelerated compute, high-throughput storage, high-performance interconnect and the software layers above them as a rented service, used to train, adapt or run artificial-intelligence models. The definition is simple. The procurement problem is that the term is applied to at least four distinct products with different buyers, budgets and risks.
The four layers — and why the distinction is the whole decision
| Layer | What you rent | Priced by | Who it suits | Principal risk |
|---|---|---|---|---|
| 1. Accelerated infrastructure (GPU-as-a-service, bare metal, clusters) | GPU or accelerator instances, optionally with InfiniBand-class fabric, parallel storage and cluster orchestration (Kubernetes, Slurm). | Per GPU-hour, per node-hour, or per reserved block. | Organisations training or fine-tuning at scale, running high-volume self-hosted inference, or with isolation requirements. | Idle capacity. You pay for allocated GPUs whether or not work is running. |
| 2. Managed ML and training platform | Experiment tracking, distributed training orchestration, feature stores, pipelines, model registry, managed endpoints. | Platform fee plus underlying compute, sometimes a surcharge on compute. | Teams with data scientists but no platform engineering capacity. | Platform-specific abstractions that are costly to unwind. |
| 3. Model inference and model APIs | Access to hosted foundation models, or your own model served on managed infrastructure. | Per million input and output tokens, per request, or per provisioned throughput unit. | The majority of enterprises, for the majority of use cases. | Model deprecation, silent behaviour changes, and per-token cost growth as usage scales. |
| 4. AI application and agent platform | Agent frameworks, connectors to business systems, retrieval, evaluation, guardrails, human-in-the-loop tooling. | Per seat, per agent, per action, or consumption-based. | Buyers automating defined business processes rather than building models. | Deep coupling to a single vendor's data and identity model. |
A useful test: if two shortlisted vendors cannot be compared on the same unit — if one quotes GPU-hours and the other quotes tokens — you are running one evaluation across two markets. Split it. Then, if the layers must be bought together, evaluate the bundle explicitly, including what happens if you later want to change one layer without changing the others.
Hyperscalers and neoclouds: a real distinction, imprecisely drawn
Industry commentary divides supply into hyperscalers — general-purpose clouds with broad service catalogues that also sell AI infrastructure — and neoclouds, sometimes called specialised GPU cloud or GPUaaS providers, whose business is predominantly renting high-end accelerators. The distinction is genuine and useful. It is also frequently overstated in both directions.
What is accurate: specialised providers typically provision faster, impose fewer platform constraints, expose newer accelerator generations earlier in some cases, and offer simpler commercial structures. Several are NVIDIA Cloud Partners and several receive NVIDIA investment. They also function as capacity suppliers to the hyperscalers and large AI labs, which means part of the “competition” between the two groups is actually a supply relationship.
What is overstated: the price gap (see Pricing and Cost Model), the breadth of geographic coverage, the maturity of compliance evidence, and the depth of managed services around the raw compute. A specialised provider that gives you a GPU cluster in fourteen days has not given you identity federation, a landing zone, a data platform, an audited control environment or a support model equivalent to a hyperscaler's. Whether that matters depends entirely on what you are building.
Who should use this guide — and who should not
Use it if you are selecting infrastructure or platform for AI workloads, sizing a first AI infrastructure budget, deciding between an inference API and self-hosted models, assessing whether to place AI workloads in Singapore, negotiating a multi-year AI compute commitment, or writing an AI cloud RFP.
Do not use it as an AI strategy document. If the use case, data readiness, success metric and accountable owner are not defined, no infrastructure decision will rescue the project. It is also the wrong guide for buying an off-the-shelf AI feature inside a SaaS product you already own — that is a software renewal conversation, and the relevant question is what the vendor does with your data, not which GPUs it rents.
Why It Matters: Capacity, Not Software, Is the Constraint
For most of the cloud era, compute was effectively unlimited from the buyer's perspective. Capacity questions were pricing questions. AI infrastructure broke that assumption, and the break is structural rather than temporary.
The four largest United States hyperscalers guided to combined capital expenditure approaching USD 700 billion for 2026, an increase of more than 60% on 2025 levels, with Amazon alone indicating around USD 200 billion. Estimates vary by source and by what is counted, but the direction and magnitude are consistent across reporting. Capital at that scale is not being deployed to win a price war. It is being deployed to secure power, land, shells, accelerators and networking — the inputs that are actually scarce.
Three consequences follow, and each changes how an enterprise should buy.
1. Availability is a negotiated outcome, not a self-service one
Newer accelerator generations are frequently subject to quotas, capacity reservations, waitlists or account-team allocation. In the Asia-Pacific region specifically, capacity is being built at record pace — the region's development pipeline reached 26.5 GW in the first half of 2026 according to Cushman & Wakefield research — but pipeline is not operational capacity, and most of it lands after 2027. Treat “can we get it, where, and when” as the first evaluation criterion rather than a fulfilment detail.
2. Commitment is how capacity is allocated — which transfers risk to you
Providers allocate scarce capacity preferentially to committed demand. That is rational, and it means the buyer who wants guaranteed access is generally the buyer who signs a one- or three-year commitment. The commitment converts a variable cost into a fixed one at exactly the moment when the technology, model strategy and workload profile are least stable. Accelerator generations have been superseding each other roughly annually. A three-year commitment on a current-generation accelerator is a bet that your utilisation will hold and that the generation gap will not make the economics uncompetitive before the term ends.
3. Cost is dominated by utilisation, not by unit price
This is the single most consequential and least discussed point in AI cloud procurement. Consider two organisations running the same training workload:
| Organisation A | Organisation B | |
|---|---|---|
| Contracted rate | USD 3.00 per GPU-hour (three-year reserved) | USD 6.00 per GPU-hour (on-demand) |
| Cluster size | 64 GPUs, reserved continuously | 64 GPUs, provisioned only when jobs run |
| Achieved utilisation | 25% | 85% |
| Monthly spend (730 hours) | USD 140,160 — paid regardless of use | USD 217,464 at full occupancy; ~USD 184,844 at 85% run-time |
| Effective cost per GPU-hour of useful work | USD 12.00 | USD 7.06 |
The organisation with the better rate card has the worse economics by a wide margin. This is an illustrative model rather than a benchmark, but the mechanism is exact and it recurs constantly in practice: idle reserved capacity is the most common source of AI infrastructure waste, and it is invisible on a rate-card comparison. Before negotiating price, establish whether the organisation has scheduling, queueing, checkpointing, multi-tenancy and shutdown discipline good enough to keep expensive accelerators busy. If it does not, buy flexibly and fix the capability first.
The competitive pressure is real, but asymmetric
Specialised providers have taken meaningful share of GPU-intensive training and inference by being faster to provision and simpler to buy. Hyperscalers retain the advantage where the AI workload has to live inside an existing enterprise estate: identity, networking, data governance, procurement vehicles, regional breadth, compliance evidence and an established support relationship. Hyperscalers have also responded on price — AWS reduced on-demand rates for its P4 and P5 NVIDIA GPU instance families by up to 45%, effective 1 June 2025 — which compresses the differential that made specialised providers attractive on cost alone.
Singapore Market Overview
Singapore is a credible location for enterprise AI workloads and a constrained one. Both facts matter, and neither cancels the other.
Infrastructure and the capacity allocation model
Singapore hosts total data-centre capacity exceeding 1.4 GW, with all three largest hyperscalers plus Oracle operating full regions, dense subsea and terrestrial connectivity, and a deep operator ecosystem. Capacity growth is deliberately managed rather than market-allocated. Following a pause on new data-centre capacity, the pilot Data Centre Call for Application in July 2022 awarded about 80 MW across four operators. The Green Data Centre Roadmap published in 2024 set out at least 300 MW of additional near-term capacity, with further capacity available through green energy pathways, and a power usage effectiveness target below 1.3 for new builds. A second call, DC-CFA2, was announced on 1 December 2025 by EDB and IMDA, opened at least 200 MW, required at least 50% of power from green sources such as biomethane, hydrogen or solar, and closed to applications on 31 March 2026.
The practical implication for buyers is straightforward: in-country AI capacity is administratively rationed and awarded against efficiency and sustainability criteria. That produces a market where Singapore capacity commands a premium, where high-density AI deployments compete for the same allocated power, and where a provider's ability to serve you locally depends on an allocation process you cannot influence. Operators are investing heavily nonetheless — Digital Realty has announced around SGD 7 billion of total Singapore investment including SGD 4.3 billion for new developments — but lead times reflect the constraint.
Two responses have emerged. The first is engineering: high-density liquid and immersion cooling to extract more compute per allocated megawatt. Singtel's Nxera data centres are designed for advanced liquid cooling supporting up to 150 kW per rack for next-generation NVIDIA systems, against a traditional industry average nearer 10 kW. Firmus Technologies operates its SMC Cloud in Singapore on a proprietary single-phase immersion platform it reports running at 1.05 PUE. The second response is regional: placing training capacity in Johor, Indonesia or Australia while keeping latency-sensitive inference and regulated data in Singapore. Both are legitimate. Both need to be reflected in the data-flow map rather than assumed away.
Government programmes and funding
Singapore's public-sector support for enterprise AI is unusually concrete for buyers, because it takes the form of named provider allocations rather than generic subsidy.
The Enterprise Compute Initiative (ECI), announced at Budget 2025 with up to SGD 150 million set aside and administered by Digital Industry Singapore (DISG), pairs eligible companies with participating cloud service providers for cloud credits, AI tools, training and consulting. Consulting costs are capped at SGD 150,000 per company, with government funding of 70% of actual costs up to SGD 105,000; the company funds the remaining 30%. The programme runs for one year from each provider's respective launch date, with staggered cohorts.
| Participating provider | Published allocation under ECI | Extended support noted |
|---|---|---|
| Amazon Web Services | Up to SGD 350,000 in AI Springboard cloud credits and training | Up to SGD 600,000 with complementary programmes such as Migration Acceleration |
| Microsoft | Up to SGD 250,000 of Azure cloud credits, AI training and tools | Co-development services valued up to SGD 700,000 for selected enterprises pursuing agentic AI strategies |
| Oracle | Up to SGD 250,000 per company in Oracle Cloud Universal Credits | Up to SGD 1.9 million for enterprises requiring private cloud infrastructure and Exadata access |
| Google Cloud | Up to SGD 200,000 worth of on-demand training licences | Workshops, certification programmes and an AI showcase opportunity |
Baseline eligibility includes a Singapore-registered or incorporated entity with physical presence, CEO-level sponsorship, at least ten Singapore-based employees, a technical team of at least two professionals such as software engineers, AI engineers or data scientists, prior experience building custom AI solutions, access to relevant datasets, and the financial ability to complete the project. Complex and scalable use cases are prioritised.
Adjacent programmes remain relevant. The Productivity Solutions Grant and Enterprise Development Grant continue to support digitalisation and capability projects, with different eligibility and claim mechanics. Public research compute is available through the National Supercomputing Centre for eligible research and public-sector users, and AI Singapore's SEA-LION family of open Southeast Asian language models provides a regionally adapted alternative to frontier commercial models for some use cases.
Market maturity: foundations ahead of operations
IMDA's official statistics show AI adoption among Singapore SMEs more than tripling to 14.5% in 2024 from 4.2% in 2023, against 62.5% among non-SMEs, with the digital economy reaching 18.6% of GDP. Vendor research points the same way while highlighting the gap that matters for infrastructure buyers: ServiceNow's 2026 Enterprise AI Maturity Index, which surveyed 200 senior leaders in Singapore among 4,500 globally, reports agentic AI adoption rising from 22% to 51% year on year but only 10% of enterprises redesigning processes end to end, and 58% citing data privacy and security as a top challenge. Vendor-sponsored survey data should be read with that provenance in mind; the direction is nonetheless consistent with the official statistics.
For a buyer, the operational reading is this: Singapore has strong AI foundations, a governance framework ahead of most jurisdictions, and a shortage of organisations that have converted either into production workloads with measured returns. That shortage is why infrastructure commitments so often outrun utilisation. It is an argument for buying flexibly first and committing later — not for waiting.
Challenges a Singapore buyer should expect
- Power-constrained local capacity for high-density AI deployments, with allocation decided through a competitive process and long lead times for new supply.
- Accelerator generation lag between Singapore regions and the largest United States regions for the newest SKUs.
- Premium pricing for Singapore-region capacity relative to large United States regions, on both hyperscaler list prices and specialist quotes.
- A short in-country specialist shortlist, discussed in the next section.
- Talent competition for platform engineering, ML operations and AI security skills, which raises the real cost of self-managed infrastructure.
- Governance workload: overlapping voluntary frameworks, sector rules and customer-imposed requirements produce genuine assurance effort even where nothing is legally mandatory.
The Residency Constraint: Why the Global Shortlist Does Not Apply
This section exists because it is the most frequent and most expensive misconception in Singapore AI infrastructure procurement. Comparison articles, vendor decks and analyst notes list the same specialist providers — CoreWeave, Lambda, Nebius, Crusoe, Nscale, Together AI, RunPod, Vast.ai — without stating where those providers actually operate. For a buyer with an in-country hosting requirement, most of that list is not selectable.
As at August 2026, based on the providers' own published documentation and announcements:
| Provider | Singapore region? | Published position |
|---|---|---|
| AWS | Yes | ap-southeast-1 with H100-class P5 instances. Verify newest-generation availability against the EC2 instance-types-by-region table. |
| Microsoft Azure | Yes | southeastasia with H100-class ND and NC series. Verify per-SKU regional availability. |
| Google Cloud | Yes | asia-southeast1 with accelerator-optimised A-series machine types. Verify per-machine-type availability. |
| Oracle Cloud Infrastructure | Yes | Singapore region; GPU shape availability varies and is often capacity-reserved. |
| Singtel RE:AI | Yes | Sovereign AI-as-a-service and GPUaaS delivered from Singtel's liquid-cooled Nxera data centres, orchestrated via Paragon. |
| Firmus Technologies / SMC Cloud | Yes | Singapore-headquartered; operates an SMC Cloud Singapore H100 region on its HyperCube immersion platform. Its large new capacity (Project Southgate) is in Australia. |
| CoreWeave | No | First Asia-Pacific data centres announced August 2026 in Indonesia — three facilities, 360 MW contracted IT power, expected online 2028. Operated 49 data centres globally as of March 2026. |
| Nebius | No | Public regions: Finland, France, Israel, United Kingdom, United States (Kansas City), plus private regions in Iceland and France. Its March 2026 Asia-Pacific announcement was a commercial expansion led from Singapore, not a Singapore region. |
| Lambda, Crusoe, Nscale, Together AI, RunPod, Vast.ai | Generally no in-country region | Footprints concentrated in North America and Europe as at publication. Confirm current regions directly — several have announced APAC intentions. |
Two refinements matter. First, the absence of a Singapore region does not disqualify a provider for workloads with no residency constraint — training on synthetic, public or de-identified data, for instance, or research workloads where latency is irrelevant. Second, generation availability is region-specific even inside a single provider. Nebius's own documentation shows B300 available only in its UK and one private European region, B200 in Israel and Kansas City, H200 across four regions, and H100 in only one. Assuming a provider's newest accelerator is available wherever the provider operates is a reliable way to lose a quarter to a re-plan.
- Does this workload have a hard in-country or in-region hosting constraint, from law, contract, sector rule or internal policy? If yes, filter the market by region first and price second.
- Which specific accelerator SKU does the workload need, and is it available in that region today, with what quota and lead time?
- Where do the secondary data paths go — backups, logs, telemetry, evaluation datasets, support access, model endpoints, human review? A compliant compute region with a non-compliant log destination is not compliant.
How Enterprise Buyers Should Evaluate AI Cloud Providers
A feature checklist cannot distinguish AI cloud providers usefully. They mostly rent the same accelerators from the same supplier. What differs is availability, interconnect, operational behaviour under load, commercial structure, evidence quality and what happens when something fails at 3 a.m. during a fourteen-day training run.
Step 1: Characterise the workload before the market
Record, for each candidate workload: whether it is training, fine-tuning, batch inference or interactive inference; model size and precision; required context length and batch size; throughput and latency targets; dataset size, growth and read pattern; expected duty cycle across a week; data classification; residency constraints; and the accountable business owner. Workloads with duty cycles below roughly 40% are usually poor candidates for reserved capacity. Interactive inference with strict latency targets is usually a poor candidate for spot or preemptible capacity.
Step 2: Filter on availability, not price
Confirm in writing: the specific SKU, the region, current quota, the lead time to first allocation, whether capacity is reserved or best-effort, the interconnect specification, and the notice period if the provider needs to reallocate capacity. Ask what happens to your workload if the provider's own capacity is contracted to a larger customer. A provider that cannot answer these precisely is not offering capacity; it is offering an intention.
Step 3: Score against weighted criteria
Weights should reflect the workload. The illustrative weighting below suits a regulated Singapore enterprise placing production inference plus periodic fine-tuning; a research team optimising for cost per experiment would weight differently.
| Criterion | Illustrative weight | What to actually test or verify |
|---|---|---|
| Capacity availability and lead time | 18% | Written SKU, region, quota, reservation type, lead time, reallocation rights, capacity roadmap. |
| Measured performance on your workload | 15% | Achieved tokens/second or samples/second at your batch size and context length; not vendor benchmark figures. |
| Total cost per unit of useful work | 14% | Cost per million tokens or per training run, including storage, egress, idle time and platform fees. |
| Data handling and residency | 12% | Processing locations for prompts, outputs, embeddings, logs, evaluation data and support access; training-use terms; sub-processor list. |
| Security and compliance evidence | 10% | MTCS tier and scope, ISO/IEC 27001, ISO/IEC 42001, SOC 2 type and period, penetration-test summaries, isolation model. |
| Operational maturity | 9% | Failure and restart behaviour on long jobs, checkpoint support, node-health handling, maintenance windows, incident history. |
| Observability and cost governance | 7% | Per-job and per-team attribution, utilisation metrics, budget alerts, exportable cost data, tagging enforcement. |
| Support and escalation | 6% | Named escalation path, response targets for capacity and cluster failures, Singapore or APAC coverage hours. |
| Portability and exit | 5% | Data and weight export, container and orchestration portability, log access, exit assistance, no proprietary checkpoint format lock. |
| Commercial flexibility | 4% | Commitment scope, generation-upgrade rights, reallocation between SKUs, term, currency, termination terms. |
Step 4: Run a proof of concept that measures the right thing
The purpose of the proof of concept is not to confirm that the GPUs work. It is to expose the operational and cost behaviour that the rate card conceals. Measure: time from order to usable capacity; achieved throughput at production batch size and context length; storage read throughput against your actual dataset; behaviour when a node fails mid-run and whether checkpointing recovers cleanly; egress volume generated by your real data movement pattern; quality of cost attribution; and support responsiveness to a genuine incident. Then compute cost for one fixed unit of work — a full fine-tuning run, or one million tokens at your prompt profile — and compare that number, not the hourly rate.
Questions to ask vendors
Common procurement mistakes at the evaluation stage
- Comparing rate cards across incomparable SKUs. A single-GPU instance and an eight-GPU fabric-connected node are different products; per-GPU-hour comparison between them is meaningless.
- Letting the grant or credit allocation pick the provider. Credits are one-off; the architecture and the switching cost are not.
- Accepting a region name as a residency answer. Verify the secondary data paths.
- Evaluating on vendor benchmarks. MLPerf and vendor figures are useful context and poor predictors of your throughput at your batch size, precision and context length.
- Skipping supplier financial due diligence on specialised providers because the technology evaluation went well.
- Buying the platform layer to solve a capability gap. A managed ML platform does not substitute for platform engineering and ML operations skills; it changes what those skills do.
- Running the evaluation without security, data protection and legal involved early. AI-specific data-processing terms are frequently the longest part of the negotiation.
Vendor Landscape
The assessments below describe positioning, strengths and limitations as at August 2026. They are not rankings, and they are not endorsements. TechDirectory does not receive vendor payment for placement in buyer's guides. Every provider listed has customers for whom it is the right choice and customers for whom it is not; the trade-offs, rather than the verdicts, are the useful part.
Tier 1: hyperscalers with Singapore regions
Amazon Web Services
Position. The largest cloud provider by revenue and service breadth, with a Singapore region (ap-southeast-1) carrying H100-class P5 instances, newer P6 Blackwell-class families in selected regions, SageMaker for the ML platform layer, Bedrock for managed model access, and its own Trainium and Inferentia accelerators as an alternative to NVIDIA supply.
Strengths. Broadest service catalogue and partner ecosystem in Singapore; mature identity, networking and governance primitives; multiple purchasing routes including on-demand, Savings Plans, Capacity Blocks for ML and private pricing agreements; custom silicon gives a second supply line and a different price point; substantial ECI credit allocation.
Limitations. GPU access commonly involves quotas, capacity reservations or account-team allocation rather than straightforward self-service. Pricing structure is complex, and the interaction between Savings Plans, Capacity Blocks and reservations is a genuine source of buyer error. Capacity Block pricing has been repriced upward during 2026 according to practitioner reporting, without accompanying announcements — a reminder that reservation products are not price-stable. Trainium adoption requires framework and code adaptation that is real work, not a flag.
Best fit. Enterprises already standardised on AWS; buyers needing breadth beyond AI; organisations that can exploit custom silicon economics. Weaker fit. Teams wanting immediate, self-service, large-cluster GPU access with minimal commercial process.
Microsoft Azure
Position. Enterprise-centric, with a Singapore region (southeastasia), ND-series GPU virtual machines with InfiniBand for distributed training, NC-series for single- and dual-GPU work, Azure AI Foundry as the model and agent platform, and deep integration with Microsoft 365, Entra and the wider Microsoft estate.
Strengths. Strongest fit where AI must integrate with an existing Microsoft identity, data and productivity estate; substantial model catalogue including OpenAI models under Azure terms; enterprise agreement and licensing leverage for existing Microsoft customers; well-documented reserved and savings-plan mechanics; large ECI allocation including co-development services for selected agentic AI projects.
Limitations. On published list prices, Azure's fabric-equipped eight-GPU ND SKUs are among the more expensive per GPU-hour — the ND96isr H100 v5 lists at roughly USD 98 per hour, about USD 12.29 per GPU-hour, while the single-GPU NC40ads H100 v5 lists near USD 6.98 per hour. That is a defensible premium for the InfiniBand fabric, but it means buyers who do not need the fabric can materially overpay by choosing the wrong SKU. Reserved pricing is where Azure becomes competitive: roughly USD 7.93 per GPU-hour at one year and USD 5.47 at three years for the same ND SKU. Capacity for the newest generations has been constrained.
Best fit. Microsoft-centric enterprises; regulated organisations valuing a single accountable vendor across productivity and AI. Weaker fit. Cost-sensitive on-demand experimentation.
Google Cloud
Position. Differentiated by custom TPU accelerators alongside NVIDIA A-series accelerator-optimised machine types, Vertex AI as the platform layer, and strong data and analytics adjacency. Singapore region asia-southeast1.
Strengths. TPUs offer a genuinely different price-performance curve for suitable workloads and reduce dependence on NVIDIA supply; strong data platform integration where the AI workload is downstream of BigQuery-resident data; Kubernetes maturity; competitive position for foundation-model and large-scale training work.
Limitations. Published NVIDIA GPU list prices have been at the higher end — an eight-GPU H100 machine type has listed around USD 88 per hour in a large United States region, roughly USD 11 per GPU-hour — before sustained-use and committed-use discounts. TPU adoption requires framework compatibility and is not a drop-in substitute for CUDA workloads; assess portability honestly before committing. Enterprise partner depth in Singapore is smaller than AWS or Microsoft in some segments. ECI allocation is framed around training licences rather than large infrastructure credits.
Best fit. Data-platform-led AI programmes; teams able to exploit TPUs; Kubernetes-native organisations. Weaker fit. Buyers requiring a broad existing enterprise-agreement relationship or CUDA-only portability.
Oracle Cloud Infrastructure
Position. Competitive bare-metal GPU offerings with high-performance RDMA cluster networking, aggressive commercial positioning, and strong ties to Oracle database and application estates. Singapore region available.
Strengths. Bare-metal GPU access avoids virtualisation overhead for some workloads; RDMA cluster networking is well regarded for distributed training; commercially flexible, frequently the most aggressive on price in competitive situations; the largest published ECI allocation, extending substantially for enterprises requiring private cloud infrastructure and Exadata access.
Limitations. Narrower general-purpose service catalogue and smaller third-party ecosystem than the other three; GPU shape availability by region is variable and frequently capacity-reserved; the commercial aggressiveness that makes OCI attractive at signature can concentrate the estate in a way that weakens leverage at renewal. Assess Oracle licensing interactions carefully where the AI workload touches Oracle-licensed software.
Best fit. Oracle-centred estates; bare-metal and RDMA-sensitive training; buyers with a defined commercial case. Weaker fit. Organisations wanting the widest managed-service catalogue.
Tier 2: Singapore and regional specialists
Singtel RE:AI
Position. A Singapore-operated AI-as-a-service platform combining GPU-as-a-service, carrier networks and orchestration, delivered from Singtel's liquid-cooled Nxera data centres and marketed on sovereignty for Southeast Asia. Publicly announced work includes a partnership with Mistral AI for a sovereign offering in Singapore, hosting for AI Singapore's SEA-LION models, and a capacity arrangement with Nscale for global GPU access.
Strengths. In-country delivery with a Singapore-regulated counterparty, which materially simplifies residency and sovereignty positions; integration of connectivity, including 5G and fixed networks, with compute; high-density cooling designed for next-generation accelerators; a plausible answer for buyers who need Singapore placement and cannot use a global specialist.
Limitations. A newer platform with a shorter production track record than the hyperscalers; managed-service catalogue and developer tooling are narrower; capacity depends on the Nxera build programme, so verify what is operational versus announced for your required timeline; ask directly which SKUs are available in Singapore versus fulfilled through partner capacity elsewhere, because a sovereign platform delivering from a non-Singapore partner region does not solve a residency constraint.
Best fit. Singapore enterprises and public-sector-adjacent buyers with residency or sovereignty requirements; organisations already contracting Singtel for connectivity. Weaker fit. Buyers needing the widest managed AI platform or the deepest global footprint.
Firmus Technologies / Sustainable Metal Cloud
Position. A Singapore-headquartered NVIDIA Cloud Partner operating SMC Cloud on a proprietary single-phase immersion cooling platform (HyperCube), with a Singapore H100 region and large new capacity under construction in Australia through Project Southgate. It has published MLPerf Training results including power-consumption figures, and reports operating at 1.05 PUE.
Strengths. Genuine engineering differentiation on performance per watt, which matters specifically in a power-allocated market like Singapore; published third-party benchmark participation rather than self-reported figures alone; in-country Singapore capacity; sustainability metrics that can be used in an organisation's own reporting, subject to verifying methodology.
Limitations. Smaller scale than the hyperscalers and the largest specialists; the substantial new capacity is Australian rather than Singaporean, so a Singapore residency requirement constrains you to the existing local region; efficiency claims are vendor-published and should be validated against your own measurement basis before being used in external reporting; managed platform services above the infrastructure layer are limited relative to a hyperscaler.
Best fit. Buyers with a Singapore placement requirement, a sustainability mandate and workloads that fit H100-class capacity. Weaker fit. Buyers needing the newest accelerator generation in Singapore, or a broad managed-service catalogue.
Regional and adjacent options
Alibaba Cloud, Tencent Cloud and Huawei Cloud operate Singapore regions with AI services and can be relevant for China-facing workloads — a decision with its own regulatory and data-flow considerations covered in Connectivity to China from Singapore. Regional carriers and data-centre operators including StarHub, SPTel, ST Telemedia Global Data Centres, Digital Edge and Equinix provide colocation, interconnection and in some cases managed GPU hosting, which is the route for organisations that want to own accelerators while renting facility capacity. For that path, see Data Centres in Singapore and Equinix vs Digital Realty.
Tier 3: global specialised GPU clouds (no Singapore region as at August 2026)
These providers are frequently the right answer for workloads without in-country constraints, and frequently the wrong answer for regulated Singapore data. Assess them on that basis.
| Provider | Positioning | Strengths | Limitations for a Singapore buyer |
|---|---|---|---|
| CoreWeave | Largest pure-play specialised GPU cloud; Kubernetes-native platform for large-scale clusters; public company. | Scale, availability of current NVIDIA generations, strong orchestration, transparent published rate card, substantial contracted backlog. | No APAC region until the Indonesian facilities come online, expected 2028. Revenue concentrated among a small number of very large customers — assess allocation protection. Eight-GPU minimum on published SKUs. |
| Nebius | AI cloud with owned and leased European and United States capacity, plus managed Kubernetes, Slurm and serverless inference. | Transparent per-GPU-hour and preemptible pricing, competitive published rates, integrated ML tooling, expanding footprint. | No APAC region; the March 2026 APAC announcement was commercial. Newest accelerators are region-limited within its own footprint. European and United States placement only. |
| Lambda | Developer- and researcher-oriented AI cloud with one-click clusters and dedicated options. | Among the lowest published on-demand rates for H100 and B200 class; good developer experience; straightforward provisioning. | No Singapore region; enterprise compliance evidence and managed services thinner than hyperscalers; suited to training and fine-tuning more than regulated production inference. |
| Crusoe | Energy-oriented infrastructure using stranded and renewable power with modular data centres; builds facilities as well as selling cloud. | Sustainability positioning with a real engineering basis; involvement in very large AI campus projects; competitive rates. | No Singapore presence; capacity concentrated in North America; sustainability claims need verification against your reporting methodology. |
| Nscale | UK-based full-stack AI cloud with Kubernetes and Slurm, serverless inference and bare metal; large deployments contracted with hyperscalers. | Sovereign and AI-optimised positioning in Europe; scale contracts; a Singtel capacity partner, which is how some Singapore buyers encounter it indirectly. | Primary footprint in Europe and the United States. Where accessed through a partner, confirm which party is the data controller, processor and contracting counterparty. |
| Together AI | Software- and inference-oriented: open-model hosting, fine-tuning and token-based APIs alongside GPU rental. | Strong fit for running open-weight models at scale without operating clusters; per-token pricing; fast model onboarding. | Inference-layer decision rather than infrastructure; residency and data-handling terms need the same scrutiny as any model API. See the LLM guide. |
| RunPod | Developer-centric flexible pods, serverless inference and clusters at competitive on-demand and community rates. | Excellent for burst workloads, experimentation and cost-sensitive production inference; low friction. | Community-tier reliability and host quality vary; enterprise assurance, contractual and residency positions are limited relative to enterprise requirements. |
| Vast.ai | Marketplace model aggregating third-party and community GPU capacity. | Frequently the lowest available cost for interruption-tolerant work. | Variable host quality and reliability; heterogeneous underlying operators make data-handling and security assurance difficult. Unsuitable for regulated data. |
| Others | Vultr, FluidStack, Hyperstack, Paperspace (DigitalOcean), DataCrunch, Genesis Cloud; capacity-layer participants including IREN, Core Scientific and Applied Digital. | Niche cost, regional or capacity advantages; some supply campuses and capacity to larger providers rather than selling directly. | Varying maturity; several entered the market from other industries such as cryptocurrency mining. Standard supplier due diligence applies with more weight, not less. |
Comparison Tables
Provider category comparison
| Dimension | Hyperscaler | Singapore specialist | Global specialised GPU cloud | Own hardware in colocation |
|---|---|---|---|---|
| Singapore in-country capacity | Yes, all four majors | Yes — core proposition | Generally not as at Aug 2026 | Yes, subject to allocated power |
| Time to first capacity | Days to months depending on SKU and quota | Days to weeks | Hours to weeks | Months — procurement, delivery, installation |
| Published on-demand rate (same generation) | Higher end of range | Negotiated, often not published | Lower to mid range | Not applicable — capital plus facility |
| Cost at high sustained utilisation | Competitive with reserved and committed pricing | Negotiated | Competitive | Frequently lowest, if utilisation is genuinely high |
| Newest accelerator generation access | Region-dependent, quota-managed | Depends on build programme | Often earliest, region-limited | Depends on supply allocation; you carry obsolescence risk |
| Managed platform breadth | Broadest | Narrower | Narrow to moderate | None — you build it |
| Compliance evidence maturity | Most extensive (MTCS, ISO, SOC 2) | Varies — verify scope | Varies — often thinner | Yours to establish |
| Operational burden on your team | Low to moderate | Moderate | Moderate to high | Highest |
| Vendor lock-in profile | Highest at platform layer; moderate at raw compute | Moderate | Lowest — largely standard containers and orchestration | Lowest technically; highest capital commitment |
| Counterparty durability | Strongest | Varies — assess | Varies — assess carefully | Your own balance sheet plus facility provider |
Deployment and consumption models
| Model | How it is bought | Suits | Cost characteristic | Main risk |
|---|---|---|---|---|
| Model API / managed inference | Per million tokens or per request; no infrastructure | Most enterprise use cases, especially first deployments | Zero idle cost; grows linearly with usage and can surprise at scale | Model deprecation, behaviour drift, per-token cost at volume |
| Serverless / provisioned throughput inference | Reserved throughput units or autoscaled endpoints | Production inference with predictable latency needs | Predictable, higher than on-demand tokens at low volume | Over-provisioned throughput sitting idle |
| On-demand GPU instances | Per GPU-hour, self-service subject to quota | Experimentation, fine-tuning, bursty batch work | Highest unit rate, lowest commitment | Capacity unavailability at the moment of need |
| Spot / preemptible GPU | Per GPU-hour at a discount, interruptible | Checkpointed training, batch inference, research | Materially lower — frequently 40–70% below on-demand | Eviction mid-job; unsuitable without robust checkpointing |
| Capacity reservation / capacity blocks | Fixed block of capacity for a defined window | Scheduled training runs needing guaranteed capacity | Premium for certainty; has been repriced upward during 2026 | Paying for the window whether or not the job is ready |
| Reserved / committed use (1–3 years) | Discount for a term and scope commitment | Steady-state production inference or continuous training | Lowest rate; highest exposure to utilisation shortfall | Idle commitment, generation obsolescence, strategy change |
| Dedicated / private cluster | Isolated capacity, often with a managed service wrap | Regulated workloads, isolation requirements, very large training | Highest absolute cost; negotiable | Under-utilisation of an expensive dedicated estate |
| Own hardware in colocation | Capital purchase plus facility, power and cooling | Very high sustained utilisation with a multi-year horizon | Lowest marginal cost; highest capital and obsolescence risk | Power allocation, refresh cycle, in-house operations capability |
Decision Matrix: Which Model Fits Your Situation
| If this describes you | Start here | Avoid | Why |
|---|---|---|---|
| First AI deployment; use case not yet proven; no ML platform team | Model API or managed inference from a provider with a Singapore region and clear data terms | Any GPU commitment; dedicated clusters | You do not yet know your utilisation, and idle accelerators are the dominant waste mode. |
| Production inference at growing volume; per-token bill becoming material | Model your break-even against provisioned throughput or self-hosted open-weight inference; pilot both | Migrating to self-hosting on assumption alone | Self-hosting only wins above a utilisation threshold that must be measured, not estimated. |
| Regular fine-tuning; bursty, checkpointable jobs | On-demand plus spot or preemptible capacity, with disciplined checkpointing | Multi-year reserved commitment | Duty cycle is too low to justify fixed capacity; interruption tolerance is a genuine cost lever. |
| Continuous large-scale distributed training | Cluster SKUs with InfiniBand-class fabric; compare hyperscaler reserved against specialist reserved | Single-GPU SKUs; providers who cannot specify fabric performance | Interconnect determines achieved scaling efficiency and therefore true cost per run. |
| Regulated data with an in-country hosting constraint | Hyperscaler Singapore region, Singtel RE:AI, Firmus/SMC, or private deployment in Singapore colocation | Global specialised GPU clouds without a Singapore region; marketplace capacity | Compute location is only part of it — you also need controllable log, backup and support paths. |
| MAS-regulated financial institution | Provider with MTCS and SOC 2 evidence in scope, contractual audit and notification rights, and a documented AI risk position | Providers unable to supply outsourcing-grade contractual terms | Accountability stays with the regulated entity; MAS's proposed AI Risk Management Guidelines raise the documentation bar. |
| Strong sustainability mandate and power-constrained placement | High-efficiency providers; verify PUE and energy methodology; consider regional placement for training | Accepting vendor efficiency claims without methodology review | Singapore capacity is allocated partly on efficiency; efficiency claims may enter your own reporting. |
| Very high sustained utilisation, multi-year horizon, in-house operations capability | Model own hardware in colocation against three-year reserved cloud | Buying hardware without a verified power allocation and refresh plan | Ownership can be cheapest at the margin but transfers obsolescence and operations risk entirely to you. |
| Singapore SME or mid-market with a defined AI use case | Check Enterprise Compute Initiative eligibility; model three-year cost without credits, then apply them | Choosing the provider with the largest credit number | Credits are one-off; switching costs are not. |
| Research, non-sensitive data, cost is the priority | Global specialised providers, spot and marketplace capacity | Enterprise-grade dedicated capacity you will not utilise | Without a residency constraint, the full global market is available and price competition is real. |
Pricing and Cost Model
This section states published list prices with dates, then explains why they are the least important number in the decision.
Published on-demand list prices, August 2026
The figures below are published list prices captured in August 2026, normalised to United States dollars per GPU-hour for comparability. They are indicative only. Rates change without notice, Singapore-region rates are typically higher than large United States regions, and negotiated enterprise pricing can differ substantially. Verify against the provider's own pricing page before using any figure in a business case.
| Provider and SKU | Accelerators | Published on-demand | Per GPU-hour | Notes |
|---|---|---|---|---|
AWS p5.48xlarge | 8 × H100 80GB | USD 55.04 / hour | ~USD 6.88 | After the up-to-45% reduction on P4 and P5 families effective 1 June 2025 |
AWS p6-b200.48xlarge | 8 × B200 | ~USD 113.93 / hour | ~USD 14.24 | Blackwell-class; region availability limited |
Azure ND96isr H100 v5 | 8 × H100 + InfiniBand | ~USD 98.32 / hour | ~USD 12.29 | ~USD 7.93 at 1-year reserved; ~USD 5.47 at 3-year; spot ~USD 2.25–3.69 |
Azure NC40ads H100 v5 | 1 × H100 NVL 94GB | ~USD 6.98 / hour | ~USD 6.98 | No cluster fabric — not comparable to ND SKUs |
Azure ND96isr H200 v5 | 8 × H200 141GB | ~USD 110.24 / hour | ~USD 13.78 | Fabric-equipped cluster SKU |
Google Cloud a3-highgpu-8g | 8 × H100 | ~USD 88.49 / hour | ~USD 11.06 | Large US region; before sustained-use and committed-use discounts |
| CoreWeave HGX H100 | 8 × H100 | USD 49.24 / hour | USD 6.16 | Spot USD 19.71 / hour (USD 2.46 per GPU-hour); 8-GPU minimum |
| CoreWeave HGX H200 | 8 × H200 | USD 50.44 / hour | USD 6.31 | Spot USD 20.93 / hour |
| CoreWeave HGX B200 | 8 × B200 | USD 68.80 / hour | USD 8.60 | Spot USD 34.11 / hour (USD 4.26 per GPU-hour) |
| CoreWeave GB200 NVL72 | 4 Blackwell GPUs per instance | USD 42.00 / hour | USD 10.50 | Rack-scale platform; a full 72-GPU rack requires 18 instances |
| Nebius H200 NVLink | Per GPU | USD 4.50 / GPU-hour | USD 4.50 | Preemptible USD 2.45 |
| Nebius B200 NVLink | Per GPU | USD 7.15 / GPU-hour | USD 7.15 | Preemptible USD 3.95 |
| Lambda H100 SXM | Per GPU | USD 3.99–4.29 / GPU-hour | USD 3.99–4.29 | Excludes applicable tax |
| Lambda B200 SXM6 | Per GPU | USD 6.69–6.99 / GPU-hour | USD 6.69–6.99 | Cluster pricing quoted separately |
The “two to seven times cheaper” claim does not survive a like-for-like check
A widely repeated claim holds that specialised GPU clouds rent capacity at rates two to seven times below hyperscalers. Compared properly, that is not what published rate cards show.
- Same accelerator, both on-demand list: AWS H100 at about USD 6.88 per GPU-hour against CoreWeave HGX H100 at USD 6.16. That is a gap of roughly 12%, not a multiple.
- Cheapest specialist against most expensive hyperscaler SKU: Lambda H100 at USD 3.99 against Azure ND96isr H100 v5 at about USD 12.29 gives roughly 3×. But the Azure SKU includes an InfiniBand fabric that the Lambda single-GPU rate does not, so part of that gap is a different product.
- Hyperscaler committed against specialist on-demand: Azure's three-year reserved rate of about USD 5.47 per GPU-hour for the fabric-equipped ND H100 SKU is below CoreWeave's published on-demand H100 rate of USD 6.16.
- Where the large multiples come from: comparing hyperscaler on-demand list against marketplace, community, spot or preemptible capacity. Those are real options with real savings — and different reliability, isolation, support and interconnect. Comparing them to on-demand enterprise capacity is not a price comparison; it is a product substitution.
The honest summary: specialised providers are typically cheaper on published on-demand list prices for the same accelerator generation, by roughly 1.1× to 3× depending on which SKUs you compare, and the advantage narrows or reverses against hyperscaler committed pricing. They also often provision faster and with less commercial friction, which has real value. But a business case built on a claimed 5× saving will not survive contact with an actual quote.
Prices move in both directions
Buyers frequently assume GPU pricing declines monotonically. It has not. AWS reduced on-demand pricing for P4 and P5 NVIDIA GPU instances by up to 45% effective 1 June 2025, with corresponding SageMaker reductions. In the other direction, practitioner reporting indicates AWS repriced EC2 Capacity Blocks for ML upward during 2026 — approximately 15% in January and a further approximately 20% on 1 July 2026 — without an accompanying announcement. We have not seen an official AWS statement confirming those increases, so treat the specific percentages as reported rather than verified; the general point stands, and it is that capacity-reservation products are priced on scarcity and are not price-stable. Build sensitivity analysis into any multi-year model, and prefer contractual price protection over an assumption of continued declines.
The cost lines that are usually missed
| Cost domain | What to model | Why it gets missed |
|---|---|---|
| Idle and low-utilisation capacity | Allocated-but-unused GPU hours; reserved capacity outside job windows; over-provisioned inference throughput | Invisible on a rate card; typically the single largest waste line |
| High-performance storage | Parallel or high-throughput file storage sized for training read patterns, plus checkpoint volume and retention | Priced separately from compute, and checkpoints for large models are substantial |
| Data transfer and egress | Dataset ingestion, cross-zone and cross-region movement, inference response egress, multi-cloud traffic | Depends on architecture, not on the compute choice, so it is often modelled last or not at all |
| Interconnect premium | The delta between fabric-equipped cluster SKUs and standalone GPU SKUs | Buyers compare per-GPU-hour rates without checking what fabric is included |
| Platform and orchestration fees | Managed ML platform charges, sometimes as a percentage surcharge on underlying compute | Quoted separately from the compute rate that anchored the comparison |
| Evaluation, guardrails and observability | Model evaluation runs, red-teaming, guardrail inference calls, tracing and logging volume | Guardrail and evaluation calls are themselves inference, and can add a material percentage to token spend |
| People | Platform engineering, ML operations, AI security, FinOps; recruitment and retention in a competitive Singapore market | Frequently excluded from an infrastructure business case entirely |
| Model and SKU lifecycle | Re-testing, re-tuning and re-validating when a model or accelerator generation is deprecated | Deprecation notice periods are short relative to enterprise change cycles |
| Commitment shortfall | The expected value of unused committed capacity under realistic utilisation scenarios | Business cases model the discount, not the probability of not earning it |
| Exit | Data and weight export, parallel running during migration, re-tuning on a new platform | Deferred to renewal, when leverage is lowest |
A defensible cost model in five steps
- Define the unit of useful work. Cost per million tokens at your prompt and response profile; or cost per completed fine-tuning run; or cost per business transaction. Every comparison uses this unit.
- Forecast the duty cycle honestly, with a range. Model low, expected and high utilisation. If the low case cannot fund a commitment, do not commit yet.
- Price the full stack, not the accelerator. Compute, storage, checkpoints, egress, platform fees, evaluation and guardrail inference, observability, support and people.
- Compare purchase modes at your measured utilisation — on-demand, spot, capacity block, one-year and three-year commitment — and identify the break-even utilisation for each. Buy the mode that wins at your low case, not your expected case.
- Re-run quarterly. Accelerator generations, list prices, model prices and your own usage all move. A cost model reviewed annually is a cost model that is wrong for most of the year.
Compliance and Security Considerations
Singapore's AI governance environment is comparatively well developed and predominantly voluntary. That combination confuses buyers. The practical position is that a small number of instruments are binding law, a larger number are guidance that functions as a de facto procurement requirement, and one significant instrument is moving toward supervisory expectation in the financial sector.
What binds, and what does not
| Instrument | Issuer | Status | What it means for an AI cloud purchase |
|---|---|---|---|
| Personal Data Protection Act (PDPA) | Parliament / PDPC | Binding law | Consent, purpose limitation, protection, accuracy and the Transfer Limitation Obligation for overseas transfers all apply to personal data used in AI systems, including in prompts, training data and outputs. |
| Advisory Guidelines on Use of Personal Data in AI Recommendation and Decision Systems | PDPC | Advisory guidance interpreting binding law | Sets out PDPC's position on consent, legitimate interests, business improvement and notification where personal data is used to develop or deploy AI systems. |
| Cybersecurity Act | Parliament / CSA | Binding for designated critical information infrastructure | If the AI workload touches designated CII, obligations on incident reporting, audits and risk assessments apply to the owner regardless of provider. |
| Model AI Governance Framework for Generative AI | IMDA (May 2024) | Voluntary | Nine dimensions covering accountability, data, testing, incident reporting, security, provenance, safety research and public good. Widely used as an internal governance baseline. |
| Model AI Governance Framework for Agentic AI | IMDA (January 2026; updated 20 May 2026) | Voluntary | Four-pillar structure addressing bounding risks upfront, making agents controllable, and operating them safely. Directly relevant if you are buying an agent platform or granting AI systems tool access. |
| AI Verify testing framework and toolkit | IMDA / AI Verify Foundation | Voluntary | Eleven governance principles aligned with EU, US and OECD standards; open-source AI Verify Toolkit for testing. Increasingly requested as evidence in enterprise and public-sector procurement. |
| Global AI Assurance Sandbox | IMDA / AI Verify Foundation | Voluntary programme | Pairs builders and deployers of generative AI applications with specialist technical testers — a route to independent testing evidence. |
| Guidelines and Companion Guide on Securing AI Systems | CSA (15 October 2024) | Voluntary, strongly encouraged | Lifecycle security approach covering supply-chain and adversarial machine-learning risks; the Companion Guide curates controls and references MITRE ATLAS and the OWASP Top 10 for ML and generative AI. |
| Proposed Guidelines on AI Risk Management | MAS — consultation closed 31 January 2026 | Proposed; supervisory expectation once finalised | Applies to all MAS-regulated financial institutions. Covers AI oversight, AI inventory, risk-materiality assessment, lifecycle controls, and required capabilities and capacity. |
| Technology Risk Management Guidelines and outsourcing requirements | MAS | Supervisory expectations | Existing obligations on material outsourcing, due diligence, audit rights, business continuity and exit apply to AI cloud arrangements. |
| FEAT principles and Veritas | MAS | Voluntary sector guidance | Fairness, ethics, accountability and transparency in financial-sector AI and data analytics. |
Certifications and what they actually evidence
| Certification / report | What it evidences | What it does not evidence |
|---|---|---|
| MTCS SS 584:2020 (Singapore) | Cloud security controls at one of three tiers — 535 controls covering basic security, governance and tenancy, then reliability and resilience for high-impact systems | That your specific service is in scope, or that the tier matches your data classification. Check the certified service list and tier. |
| ISO/IEC 27001 | An information security management system, independently certified | Which services, regions and processes are inside the statement of applicability |
| ISO/IEC 42001 | An AI management system — governance of AI development and use | That any individual model is safe, accurate or fit for your purpose |
| SOC 2 (Type 1 / Type 2) | Controls design (Type 1) or operating effectiveness over a period (Type 2) against selected trust services criteria | Anything outside the stated period, criteria or system boundary. Read the exceptions, not just the opinion. |
| ISO/IEC 27017 / 27018 | Cloud-specific controls and protection of personal data in public cloud | PDPA compliance for your particular processing |
| NIST AI Risk Management Framework | A voluntary structure for identifying and managing AI risk | Certification of any kind — it is a framework, not an attestation |
AI-specific security and data-handling questions
Standard cloud security due diligence is necessary and insufficient. AI workloads introduce concerns that conventional questionnaires do not cover.
Governance you own regardless of provider
No cloud contract discharges accountability. Establish before production: an AI system inventory with risk classification; named accountable owners per system; documented use-case approval; data-provenance records for training and fine-tuning data; evaluation and acceptance criteria with pre-deployment testing evidence; guardrail and human-oversight design; monitoring for drift, quality and cost; incident response covering AI-specific failure modes; and periodic review against your governance framework of choice. MAS's proposed guidelines make the inventory and risk-materiality assessment explicit for financial institutions; the same structure is good practice everywhere.
Implementation Considerations
Indicative phasing
Durations below are indicative for a mid-sized Singapore enterprise deploying a first production AI workload. They vary widely with data readiness, regulatory scope and internal decision speed. Use exit criteria per phase rather than announcing a date before discovery.
| Phase | Indicative duration | Exit criteria | Common failure |
|---|---|---|---|
| Use-case definition and data readiness | 3–8 weeks | Named business owner, measurable success metric, data available and classified, baseline performance of the current process | Starting infrastructure selection before the success metric exists |
| Governance and risk assessment | 2–6 weeks, parallel | Risk classification, data-flow map, approval recorded, residency position agreed with legal and compliance | Treating governance as a post-build sign-off |
| Provider evaluation and proof of concept | 4–10 weeks | Measured throughput and cost per unit of work on the real workload, capacity confirmed in writing, support tested | Proof of concept on a benchmark rather than the workload |
| Foundation build | 4–12 weeks | Landing zone, identity, network, key management, logging, cost allocation, model registry, evaluation harness, guardrails | Deferring cost allocation and evaluation tooling until after go-live |
| Pilot in production conditions | 4–8 weeks | Real users or real traffic, measured quality and latency, cost per transaction, incident dry-run completed | A pilot with synthetic traffic that never surfaces real failure modes |
| Scale and commercial optimisation | Continuous | Utilisation measured over a full quarter before any commitment; commitments sized to the low case | Committing at pilot enthusiasm rather than measured baseline |
Stakeholders and who must be involved when
- Business owner — owns the success metric and the decision to proceed. Involved from day one.
- CIO / CTO and enterprise architecture — own target architecture, integration and platform standards.
- CISO and security engineering — own the AI-specific threat model, tenancy and key-management requirements, and guardrail design. Involved before shortlisting, not at contract review.
- Data protection officer and legal — own PDPA position, data-processing terms, transfer assessments and sub-processor review.
- Risk and compliance — own regulatory mapping; in financial services, own the outsourcing and AI risk assessment.
- Procurement — owns the scoring model, commercial structure, commitment terms and exit provisions.
- Platform engineering and ML operations — own scheduling, utilisation, checkpointing, observability. Their capability determines whether the economics work at all.
- FinOps or finance business partner — owns cost allocation, forecast accuracy and commitment utilisation tracking.
- Internal audit — validates that the governance you documented is the governance you operate.
Integration, migration and change
Integration is usually where AI programmes overrun. Model access is easy; the work is in identity and authorisation for AI systems, retrieval pipelines over enterprise data with permissions preserved, connectors to business systems, evaluation harnesses, guardrail placement, observability that traces a request through retrieval, model and tool calls, and cost attribution per team and per use case. Budget for this explicitly. It is frequently larger than the compute line in year one.
Migration between AI providers is not equivalent to a virtual-machine migration. Prompts and agent configurations are tuned to a specific model's behaviour; fine-tuned weights may not be exportable in a usable form; embeddings are model-specific and re-embedding a large corpus has real cost; evaluation baselines must be re-established. Assume re-tuning and re-validation effort on any model change, including one forced by deprecation, and secure export rights and notice periods contractually.
Change management and training determine adoption. The evidence from Singapore enterprises is consistent on this point: organisations that redesigned processes around AI report materially better outcomes than those that layered AI onto existing workflows. That is a change-management finding, not an infrastructure one, and it is where most of the realised value sits. Plan for role and workflow redesign, not only tool rollout. Related capability funding may be available through Singapore's grant programmes.
Common Mistakes
- Buying GPU capacity when the requirement was an inference API. The most expensive mistake in the category, and the most common. Establish the utilisation case first.
- Comparing per-GPU-hour rates across SKUs with different interconnect. A fabric-equipped cluster node and a standalone GPU are different products.
- Assuming a global provider can serve Singapore. Verify the region, not the logo. Several widely recommended specialists have no Singapore region.
- Assuming a provider's newest accelerator is available everywhere it operates. Generation availability is region-specific, including within specialist providers' own footprints.
- Committing before utilisation is measured. Reserved capacity at low utilisation is more expensive than on-demand at high utilisation, often by a factor of two or more.
- Building the business case on a claimed price multiple. Check the claim against published rate cards for the same accelerator generation and purchase mode.
- Letting credit allocations choose the architecture. Model the three-year cost without credits first.
- Treating a region name as a data-residency answer. Map prompts, outputs, embeddings, logs, telemetry, backups, evaluation data and support access.
- Omitting evaluation and guardrail inference from the cost model. Guardrails and evaluations are themselves model calls and can add materially to token spend.
- Ignoring storage and egress. Training read throughput, checkpoint volume and inference egress are frequently underestimated and priced separately.
- Skipping counterparty due diligence on specialist providers because the technical evaluation was strong.
- No plan for model or SKU deprecation. Notice periods are short relative to enterprise validation cycles.
- Deferring FinOps until the first large bill. Tagging, per-job attribution, utilisation reporting and budget alerts are foundation-phase requirements.
- Assigning governance to the provider. Certifications are evidence. Accountability remains with the deploying organisation.
- Running the programme without platform engineering capability. Utilisation discipline is the difference between a working business case and a failed one, and it cannot be bought as a feature.
- Leaving exit terms to renewal. Weight and data export, log access and assistance obligations are negotiable only before you depend on the provider.
Future Trends: 2026–2029
Power and land, not chips, become the visible constraint
Accelerator supply has dominated commentary; increasingly the binding constraints are electrical capacity, grid connection and suitable buildings. Singapore has made this explicit by allocating data-centre capacity administratively against efficiency and green-energy criteria. Expect placement decisions to become more regional — training in markets with available power, latency-sensitive inference and regulated data in Singapore — and expect energy efficiency and performance-per-watt to move from a sustainability annex into the commercial evaluation.
Inference overtakes training as the dominant spend
As deployments move from experimentation to production, recurring inference rather than episodic training will consume most AI infrastructure budget. This shifts the buying centre of gravity: cost per million tokens, batching and caching efficiency, quantisation and precision choices, model routing between smaller and larger models, and latency at percentile targets become the economic levers. Lower-precision formats materially change cost per token on newer accelerators, which means the same workload can get cheaper without a price change — if your stack can exploit it.
Accelerator diversity increases, and portability becomes a procurement term
NVIDIA remains central, but AWS Trainium and Inferentia, Google TPUs, and Microsoft's own silicon give the hyperscalers a second supply line and a different price curve. AMD accelerators continue to gain deployment. For buyers this creates genuine optionality and a new due-diligence question: how portable is your stack across accelerator architectures, and what is the cost of exercising that portability? Expect portability to appear explicitly in RFPs rather than being assumed.
Rack-scale systems change what a “GPU-hour” means
Rack-scale platforms such as NVIDIA's GB200 NVL72 are sold and priced as coherent units rather than as independent GPUs, with liquid cooling and dense interconnect designed in. Comparing them on per-GPU-hour rates against a conventional eight-GPU node understates what is being bought. Interconnect topology, memory coherence and rack-level power will need to appear in evaluation criteria, and facility readiness — 100 kW-plus per rack — becomes a gating factor for in-country placement.
Agentic AI governance moves from framework to procurement requirement
Singapore published a governance framework specifically for agentic AI in January 2026 and updated it within four months, which is an indicator of how quickly this area is moving. As AI systems are granted tool access, credentials and the ability to take actions, buyers will need contractual and technical answers on permission bounding, action logging, reversibility and human confirmation points. Expect these to become standard RFP content, and expect the security review of an agent platform to look more like a privileged-access review than a software review.
Sovereign and regional models become a real option
Sovereign AI investment is growing across jurisdictions, and Southeast Asia now has regionally adapted open models — AI Singapore's SEA-LION family being the local example, hosted with commercial GPU partners. For some use cases involving regional languages, regulated data or cost-sensitive deployment, a smaller regionally adapted model running on in-country infrastructure will outperform a frontier model on total cost and compliance position, if not on raw capability. Evaluate it as a genuine alternative rather than a policy gesture.
Commercial structures mature — and concentration risk gets board attention
Expect more sophisticated contracting: generation-upgrade rights, capacity-reallocation clauses, utilisation-linked commitments, price protection and clearer exit assistance. In parallel, the circularity of the current market — specialist providers financed against contracts with the same hyperscalers and AI labs they nominally compete with — is drawing scrutiny. Concentration risk, counterparty durability and exit readiness are moving from procurement checklists to board risk registers, particularly in regulated sectors.
Frequently Asked Questions
What is AI cloud computing?
It is the delivery of accelerated compute, storage, high-speed interconnect and the software above them as a rented service, used to train, fine-tune or run AI models. Practically, buyers are choosing between four different purchases: raw accelerated infrastructure, a managed training and machine-learning platform, model inference sold per token or per request, and an AI application or agent platform. Identifying which one you need is the first and most consequential decision.
Are specialised GPU clouds really two to seven times cheaper than hyperscalers?
Not on a like-for-like comparison. In August 2026, CoreWeave listed NVIDIA HGX H100 at USD 6.16 per GPU-hour on-demand while AWS listed p5.48xlarge at USD 55.04 per hour, about USD 6.88 per GPU-hour — roughly a 12% gap. The widest multiples compare hyperscaler on-demand list prices against the cheapest marketplace, spot or preemptible capacity, which differs in interconnect, reliability, isolation and support. Azure's three-year reserved rate for its fabric-equipped H100 cluster SKU, around USD 5.47 per GPU-hour, sits below CoreWeave's published on-demand rate. Specialists are generally cheaper on published on-demand pricing, by roughly 1.1× to 3× — not by a factor of five.
Can Singapore enterprises buy GPU capacity from CoreWeave, Nebius or Lambda in Singapore?
Not from an in-country region as at August 2026. CoreWeave announced its first Asia-Pacific data centres for Indonesia in August 2026, with three facilities totalling 360 MW of contracted IT power expected online in 2028. Nebius publishes regions in Finland, France, Israel, the United Kingdom and the United States, plus private regions in Iceland and France; its March 2026 Asia-Pacific announcement was a commercial expansion led from Singapore rather than a Singapore region. If you have an in-country hosting constraint, shortlist the hyperscaler Singapore regions, Singapore-based providers such as Singtel RE:AI or Firmus/SMC, or a private GPU deployment in Singapore colocation.
Which providers offer GPU instances in a Singapore region?
AWS (ap-southeast-1), Microsoft Azure (southeastasia) and Google Cloud (asia-southeast1) all operate Singapore regions with H100-class GPU instances, and Oracle Cloud Infrastructure has a Singapore region. Singtel's RE:AI provides GPU-as-a-service from its Nxera data centres in Singapore, and Firmus Technologies operates an SMC Cloud Singapore H100 region. Blackwell-class accelerators are concentrated in a small number of regions, mostly in the United States. Verify current availability against each provider's own region and instance tables before committing — this changes frequently.
How much does AI cloud computing cost?
The unit differs by layer. Accelerated infrastructure runs roughly USD 2–14 per GPU-hour on published on-demand list prices as at August 2026, depending on accelerator generation, interconnect and provider; committed and spot rates sit below that, and capacity-reservation products above. Managed inference is priced per million tokens or per request. The dominant driver is not the rate but utilisation: a reserved cluster running at 30% utilisation can cost more per unit of useful work than on-demand capacity at twice the hourly rate.
Should most enterprises buy GPUs at all?
Most should not, at least initially. Renting inference through an API or managed platform avoids capacity risk, cluster operations and idle-time waste, and it lets you learn your real usage profile before committing capital. Direct GPU procurement becomes defensible when there is sustained high utilisation, an isolation or residency requirement that inference APIs cannot satisfy, meaningful custom training or fine-tuning, latency requirements that need local placement, or a measured cost case at volume.
Does the PDPA require AI workloads to stay in Singapore?
No. The PDPA does not impose a universal in-country hosting rule. Organisations transferring personal data overseas must meet the Transfer Limitation Obligation, and sectoral rules, contracts, public-sector policies and internal risk requirements may be stricter. For AI specifically, verify where prompts, outputs, embeddings, vector indexes, inference logs, evaluation datasets, fine-tuning data and human-review processes are stored and processed — not only where the model executes.
What Singapore AI governance rules apply to buyers?
The PDPA is binding law, as is the Cybersecurity Act for designated critical information infrastructure. IMDA's Model AI Governance Framework family is voluntary, including the Model AI Governance Framework for Generative AI (May 2024) and the Model AI Governance Framework for Agentic AI, first published in January 2026 and updated on 20 May 2026. CSA's Guidelines and Companion Guide on Securing AI Systems (October 2024) are voluntary but strongly encouraged. MAS consulted on proposed Guidelines on AI Risk Management for financial institutions, with the consultation closing on 31 January 2026. Voluntary instruments still function as de facto procurement requirements when large buyers and regulated counterparties ask for evidence against them.
What Singapore funding supports AI cloud adoption?
The Enterprise Compute Initiative, announced at Budget 2025 with up to SGD 150 million and administered by Digital Industry Singapore, pairs eligible companies with participating cloud service providers. Consulting costs are capped at SGD 150,000 per company with the government funding 70% up to SGD 105,000. Participating providers publish separate allocations — AWS up to SGD 350,000 in AI Springboard credits and training, Microsoft up to SGD 250,000 in Azure credits, Oracle up to SGD 250,000 in Universal Credits, Google Cloud up to SGD 200,000 in training licences — with larger figures cited for specific extended programmes. Baseline eligibility includes Singapore registration with physical presence, CEO-level sponsorship, at least ten Singapore-based employees, at least two technical professionals, prior custom-AI experience and access to relevant datasets.
Is InfiniBand necessary for AI workloads?
It depends on the workload. Large distributed training across many nodes is sensitive to interconnect bandwidth and latency, which is why cluster SKUs bundle InfiniBand or equivalent fabrics and cost materially more per GPU-hour than single-GPU SKUs. Single-node fine-tuning and most inference do not need it. Comparing a fabric-equipped cluster SKU against a standalone GPU SKU on price per GPU-hour is not a valid comparison — you are pricing two different products.
What is the main financial risk in AI cloud contracts?
Committing to capacity before utilisation is proven. Reserved capacity, private cloud blocks and multi-year terms convert a variable cost into a fixed one at the point when model strategy, workload profile and accelerator generation are least stable. If the project is descoped, the model strategy changes, or a newer generation makes the economics uncompetitive, the payments continue. Commit against a measured baseline of at least one quarter, keep a flexible tranche, and negotiate generation-upgrade and reallocation rights before signing.
What counterparty risk applies to specialised AI cloud providers?
Several are capital-intensive businesses financed against long-term contracts with a small number of very large customers, and some entered the market from adjacent industries such as cryptocurrency mining. That is not disqualifying — the largest are substantial, publicly scrutinised companies. It is a reason to apply normal supplier due diligence: financial durability, revenue concentration, the seniority of your contract relative to larger customers, what protects your allocation if those customers expand, data-export rights, and what happens to your workload if the provider is restructured.
How should an AI cloud proof of concept be scoped?
Run it on the actual workload, not a benchmark. Measure achieved throughput and latency at your production batch size, precision and context length; time from order to usable capacity; failure and restart behaviour during long jobs and whether checkpointing recovers cleanly; storage read throughput against your real dataset; egress volume from your real data movement; quality of per-job cost attribution; and support responsiveness to a genuine incident. Then compute cost for one fixed unit of work — a full fine-tuning run, or one million tokens at your prompt profile — and compare that.
Does a Singapore region guarantee low latency for AI applications?
No. A Singapore region reduces network distance, but end-to-end latency also depends on model size, batching, context length, retrieval steps, guardrail checks, agent tool calls and any cross-region dependency such as a global control plane or a model endpoint hosted elsewhere. Measure the full request path under representative load, at the percentile your users actually experience.
What should an AI cloud RFP require as evidence?
Region and SKU availability with capacity commitments and lead times; interconnect specification and measured cluster performance; published and negotiated rates with commitment scope and term; data-processing terms covering training use of customer data, retention and sub-processor locations; security certifications with scope and reporting period; model and SKU lifecycle with deprecation notice periods; observability and cost-allocation capability; support and escalation model with Singapore or APAC coverage; exit assistance; and rights to export data, configurations, fine-tuned weights and logs.
How do the four hyperscalers differ for AI specifically?
In broad terms, and subject to testing against your workload: AWS offers the widest service and partner breadth plus custom silicon as a second supply line; Azure is the strongest fit where AI must integrate with an existing Microsoft identity and productivity estate, with competitive reserved pricing but expensive on-demand cluster SKUs; Google Cloud differentiates on TPUs and data-platform adjacency, with higher published NVIDIA list prices; Oracle competes on bare-metal GPU access, RDMA cluster networking and commercial aggressiveness, with a narrower general-purpose catalogue. None is universally better. Test the specific services, region availability, commercial terms and operating model you will actually use.
Final Recommendations
Decide which of the four layers you are buying, and buy the highest one that meets the requirement. Inference APIs and managed platforms remove capacity risk, operations burden and idle-time waste. Descend to raw accelerated infrastructure only when a specific trigger justifies it: measured sustained utilisation, an isolation or residency requirement an API cannot meet, real custom training, or latency that needs local placement. The default should be the layer with the least commitment, not the most control.
Filter on region and SKU availability before price. For Singapore buyers with residency constraints, this single step eliminates most of the vendor list that generic comparisons recommend, and it does so before you spend evaluation effort. Get the SKU, region, quota, lead time and reallocation terms in writing, and separately confirm where logs, backups, telemetry and support access land.
Verify every price claim against a published rate card for the same accelerator generation and purchase mode. The gap between specialised providers and hyperscalers is real but considerably narrower than commonly asserted, and it inverts against multi-year committed pricing. Then set the rate aside and model cost per unit of useful work, because utilisation will move the answer more than the rate does.
Build utilisation discipline before buying commitment. Scheduling, queueing, checkpointing, right-sizing and shutdown hygiene are what make AI infrastructure economics work. They are organisational capabilities, not contract terms, and no discount compensates for their absence. Measure a full quarter of production utilisation before converting variable cost into fixed cost, and size commitments to the low case.
Treat Singapore's governance frameworks as procurement content, and own the accountability. Use IMDA's Model AI Governance Frameworks, CSA's Guidelines on Securing AI Systems and AI Verify as the structure for your own AI inventory, risk classification, testing evidence and oversight design. If you are MAS-regulated, plan against the proposed AI Risk Management Guidelines now rather than after they are finalised. Provider certifications support your position; they never replace it.
Negotiate the exit while you still have leverage. Data and fine-tuned-weight export, log access, deprecation notice periods, generation-upgrade rights, capacity reallocation and exit assistance are inexpensive to secure at contract and very expensive to retrofit. In a market where accelerator generations turn over annually and providers are still consolidating, the ability to change your mind is a substantive commercial asset.
Who should consider this category: organisations with a defined, owned AI use case, measurable success criteria and data that is available and classified. Who should wait: organisations whose AI programme is still a set of experiments without an accountable owner or a success metric. For them, the correct next purchase is not infrastructure — it is a bounded pilot on consumption-priced inference, with a decision gate attached.
Primary Sources and Further Reading
Source links last checked 6 September 2026. This records that each link resolved, not that its content was re-read.
This guide was compiled and reviewed against the public sources below on 17 August 2026. AI cloud pricing, region and SKU availability, capacity, regulatory guidance and contractual terms change frequently — in this category, sometimes within weeks. Every figure in this guide should be re-verified against the primary source before it is used in a procurement, budget or compliance decision. Where a claim rests on practitioner reporting rather than an official statement, we have said so in the text.
Singapore government and regulators
- IMDA — Artificial Intelligence in Singapore, including the Model AI Governance Framework for Agentic AI and AI Verify.
- AI Verify Foundation — AI Verify testing framework and toolkit and the Global AI Assurance Sandbox.
- CSA — Guidelines and Companion Guide on Securing AI Systems (15 October 2024).
- MAS — Consultation Paper on Proposed Guidelines on AI Risk Management for Financial Institutions (consultation closed 31 January 2026).
- MAS Technology Risk Management Guidelines and MAS outsourcing requirements.
- PDPC — PDPA advisory guidelines, including those on the use of personal data in AI recommendation and decision systems.
- IMDA — Cloud computing and services standards, including MTCS SS 584:2020.
- Digital Industry Singapore — Enterprise Compute Initiative and the participating cloud service providers.
- IMDA — Singapore's digital economy and enterprise AI adoption statistics.
Provider documentation and pricing
- AWS — Amazon EC2 instance types by Region, P5 instances, P6 instances and P6e UltraServers and EC2 Capacity Blocks for ML pricing.
- AWS — up to 45% price reduction for EC2 NVIDIA GPU-accelerated instances (effective 1 June 2025).
- Microsoft Azure — virtual machine series pricing (ND and NC series).
- Google Cloud — accelerator-optimised machine type pricing and GPU machine types.
- Oracle Cloud Infrastructure price list.
- CoreWeave pricing and CoreWeave's first Asia-Pacific expansion (Indonesia, August 2026).
- Nebius — AI Cloud regions and per-region platform availability, compute pricing and its March 2026 Asia-Pacific announcement.
- Lambda — GPU cloud pricing.
- Singtel RE:AI — AI cloud services and the RE:AI launch release.
- Firmus Technologies and Sustainable Metal Cloud.
Market and infrastructure context
- Singapore's Green Data Centre Roadmap and the DC-CFA2 capacity allocation call.
- Hyperscaler 2026 capital expenditure guidance — estimates vary by source and basis.
- Asia-Pacific data-centre pipeline (Cushman & Wakefield research, H1 2026).
- Network World — neoclouds challenge hyperscalers for AI workloads.
- AI Singapore — SEA-LION regional language models.
- ISO/IEC 42001 — AI management systems and NIST AI Risk Management Framework.
Browse AI and Cloud Providers in Singapore
TechDirectory lists directory records for cloud providers, AI computing companies, system integrators, managed-service providers, data-centre operators and cybersecurity vendors serving Singapore. Profiles may show recorded capabilities, certifications and approved reviews where available. Verify delivery capability, region availability, commercial relationship, service scope and references directly with the provider before contracting.
Browse AI Computing Providers →