Chapter 2.4 — Cloud & the Software Moat
The cloud is where models actually run, and it is where the AI build turns into revenue. Three companies still own most of it, but the more important fact is that renting GPUs has spawned a second tier of "neoclouds" whose backlogs dwarf their revenue and whose balance sheets are stacked with GPU-backed debt. The durable moat in this layer is not the data center; it is the software, above all Nvidia's CUDA, and the question that decides the next few years is whether that moat holds at the inference layer or cracks.
Running an AI model at scale requires renting, or building, an enormous cluster of accelerators, the networking to bind them, and the power to feed them. The cloud is the business of providing that on demand. It splits into two tiers. The hyperscalers — Amazon Web Services, Microsoft Azure, and Google Cloud — are the diversified giants that sell everything from storage to AI. The neoclouds — led by CoreWeave and Nebius — are specialists that do essentially one thing, rent out Nvidia GPUs, and have grown from nothing to a meaningful share of the market in two years. Sitting above both is the layer that actually locks customers in: the software stack, and in particular Nvidia's CUDA, the programming environment that fifteen years of developer investment have made the default way to run AI.
Cloud Infrastructure Provider Landscape
The roster spans the diversified giants, the GPU-rental specialists, the inference-optimized upstarts, and the Chinese incumbents building on a separate stack.
| Provider | Ticker | Tier | What to know |
|---|---|---|---|
| AWS | AMZN | hyperscaler | #1, ~28% share; Trainium silicon |
| Microsoft Azure | MSFT | hyperscaler | AI run-rate ~$37B, +123% |
| Google Cloud | GOOGL | hyperscaler | fastest-growing (+82%); TPUs |
| Oracle | ORCL | hyperscaler | Stargate; debt-funded |
| CoreWeave | CRWV | neocloud | ~$99B backlog; GPU-backed debt |
| Nebius | NBIS | neocloud | "sold out"; $7–9B ARR guide |
| Crusoe | PVT | neocloud | ~$30B valuation; ~4.9 GW contracted |
| Lambda | PVT | neocloud | $9B val; IPO planned H2 2026 |
| Together AI | PVT | neocloud | ~$1B ARR; inference serving |
| Groq, Cerebras, SambaNova | PVT / CBRS | inference specialist | custom inference silicon-as-a-service |
| Alibaba Cloud | BABA / 9988.HK | China | ~36% China share; AI run-rate ~$5.2B |
| Huawei Cloud | PVT | China | Ascend/CANN "CloudMatrix" stack |
| Tencent Cloud | 0700.HK | China | ~9% share; GPU-supply-constrained |
The giants, and how fast AI is growing inside them
The cloud market is large and accelerating. Total enterprise cloud infrastructure spending reached $128.6B in the first quarter of 2026, up 35% year on year, the fastest growth in more than four years, on an annualized run-rate above half a trillion dollars. The three leaders hold roughly two-thirds of it.
The headline market share understates how quickly AI is reshaping these businesses. Microsoft's AI operation reached a roughly $37B annual run-rate, up 123% year on year, and its commercial book of committed future revenue hit $627B, nearly doubling. Google Cloud grew 82% in the second quarter of 2026 to $24.8B, and its backlog roughly doubled in a single quarter to about $514B. Amazon's AWS, the largest at 28% share and $37.6B in quarterly revenue, grew 28%, its fastest in fifteen quarters, though its operating margin slipped from 39.1% to 37.7% as the depreciation on all those AI servers began to bite. That margin detail is a preview of a theme in Chapter 2.11: the capital intensity of AI is starting to show up in the income statement.1
The neoclouds, and the backlog that dwarfs the revenue
The more novel part of this layer is the neocloud, and it is best understood through one company's numbers. CoreWeave reported first-quarter 2026 revenue of about $2.08B, up 112% year on year, which annualizes to roughly $8B. Against that, it disclosed a revenue backlog of $99.4B, including a $21B commitment from Meta and multi-year deals with Anthropic and others. The backlog is more than ten times the revenue run-rate.2
That gap is the whole model, and the whole risk. CoreWeave spent about $7.7B on capital in the quarter, close to four times its revenue, funded by roughly $25B of debt, much of it collateralized by the Nvidia GPUs it buys, and it carries over 1 GW of active power against more than 3.5 GW contracted. The rest of the field shows the same shape at different scales: Nebius guides to $7–9B of committed annual revenue on a base that was under $2B and tells investors it is "sold out"; Crusoe is raising at around a $30B valuation with roughly 4.9 GW contracted; Lambda, at a $9B valuation, plans an IPO in the second half of 2026 on the back of a large Microsoft deployment; and Together AI has reached roughly $1B of annualized revenue serving inference. The economics work if the GPUs stay highly utilized — a debt-financed cluster roughly breaks even around 70% utilization ⚠️ — and if the take-or-pay contracts hold; they break if either fails. GPU rental prices are themselves falling, with H100 rates down sharply as Blackwell supply floods in, which squeezes the margin on older fleets. Nvidia sits at the center of it all, having invested around $2B each in CoreWeave and Nebius, whose purchases of Nvidia chips its own money helps fund, the circular arrangement Chapter 2.11 dissects.
The CUDA moat, and where it is cracking
If the data center is a commodity, the software is not. Nvidia's CUDA has roughly six million developers and first-class support in every major machine-learning framework, and the switching cost is real: years of CUDA-optimized training pipelines and custom kernels would have to be rewritten to move off it. This is the true source of Nvidia's pricing power, more durable than any single chip.
The moat is not uniform, though, and the interesting question is where it thins. For training at the frontier, it holds firmly. For inference, it is eroding. AMD's ROCm software has closed enough of the gap that AMD claims parity by the end of 2026, and its data-center revenue grew 57%. Cross-hardware abstraction layers — OpenAI's Triton and PyTorch 2.0 among them — increasingly let a model run on AMD or a custom chip with little rewriting, which shifts the buying decision toward raw cost per token, exactly where the custom ASICs of Chapter 2.5 compete. Huawei's CANN, the Chinese alternative, claims that around 80% of standard PyTorch inference workloads run with only minor adjustments. The consensus of the evidence is that the moat is intact for training and genuinely eroding at inference, which matters because inference is now the larger share of compute.
The inference specialists
A distinct group of companies is betting that inference specifically will run best on purpose-built silicon rather than general GPUs. Groq (valued around $6.9B) builds deterministic low-latency inference chips; Cerebras (which staged a large 2026 IPO) makes wafer-scale processors and signed a multi-hundred-megawatt OpenAI deal; and SambaNova (around $11B) offers full-stack inference systems. Their pitch is inference several times cheaper per token than a general GPU ⚠️, and their relevance rises as the compute mix shifts toward inference. They are mostly private or newly public, higher-risk, and dependent on a few anchor contracts, but they are the names to recognize when the conversation turns to inference-optimized compute.
America's clouds, China's parallel
The national split repeats the pattern of the layers above. The United States has both tiers, the hyperscalers and the neoclouds, running on Nvidia and CUDA. China has a domestic cloud market roughly a tenth the size, concentrated among three players: Alibaba Cloud at about 36% share, Huawei Cloud at 16%, and Tencent Cloud at 9%. Alibaba is the standout, its external cloud revenue growing 40% with AI products now about 30% of it and an AI run-rate near $5.2B, eleven consecutive quarters of triple-digit AI growth, which makes it the investable expression of both the Chinese cloud and the open-weight model layer.
The constraint on the Chinese cloud is the one that runs through this whole atlas: chips. Huawei Cloud's external revenue actually fell 3.5% in 2025 despite a growing market, held back by US export controls on the compute it needs, and Tencent's cloud growth is capped by GPU supply. China's answer is the same vertical stack it is building everywhere else, Huawei's Ascend chips with the CANN software, deployed through Huawei's CloudMatrix systems that bind hundreds of Ascend chips into a single cluster. It is a genuine second ecosystem, years behind on software but improving, and it means a Chinese enterprise increasingly rents Ascend-powered compute running CANN rather than Nvidia running CUDA.
Utilization, backlog and cash conversion
Three things. The first is whether CUDA's inference moat holds or gives way, which will show up in AMD's and the custom-ASIC vendors' share of inference workloads and in whether the abstraction layers make switching genuinely painless. The second is the financial health of the neoclouds, where the signals to watch are utilization rates, any customer that fails to take contracted capacity, GPU-rental pricing, and stress in the GPU-backed debt that Chapter 2.11 treats as a systemic question. The third is margin: whether hyperscaler cloud profitability holds as AI-server depreciation mounts, the early sign of which is already visible in AWS's slipping margin.
Where to own it
Microsoft (MSFT), Alphabet (GOOGL), and Amazon (AMZN) combine cloud exposure with diversified cash generation; Oracle (ORCL) is more concentrated in large AI contracts and carries greater financing sensitivity. Nvidia (NVDA) owns a software and systems moat across cloud providers, while Broadcom (AVGO) and Arista (ANET) supply custom silicon and networking. CoreWeave (CRWV) and Nebius (NBIS) provide more direct GPU-cloud exposure with materially greater balance-sheet, customer and utilization risk. Cerebras, Groq and SambaNova are higher-risk bets on specialized inference. Alibaba (BABA, 9988.HK) is a listed expression of China's domestic cloud and open-weight model ecosystem. These securities share demand factors but have very different equity-capture and valuation profiles.
What would break the cloud thesis
The bullish case for the hyperscalers weakens if cloud margins compress faster than AI revenue grows, turning the capex into a drag rather than an engine, the early warning of which is already in AWS's margin line. The neocloud thesis breaks on a utilization or funding shock, a single large customer failing to take capacity, or a downgrade of the GPU-backed debt, either of which would expose the leverage quickly, especially with rental prices already falling. And Nvidia's software moat, the most valuable asset in this layer, would be revealed as thinner than assumed if a major lab or hyperscaler moved a frontier training run onto non-CUDA hardware without a painful rewrite, which has not yet happened but is the thing to watch.
Sources
Linked evidence for this chapter's figures and load-bearing claims: 2 1
Footnotes
-
Cloud Market Annual Revenue Run Rate Topped Half a Trillion Dollars in Q1. Synergy Research Group, undated; accessed 2026-07-25. ↩ ↩2
-
CoreWeave Reports Strong First Quarter 2026 Results. CoreWeave, 2026-05-07; accessed 2026-07-25. ↩ ↩2