Chapter 2.4 — Cloud & the Software Moat
The cloud's real product is not access to a GPU. It is utilization: combining customers, software, and scheduling well enough to keep an expensive, depreciating asset earning. Hyperscalers and neoclouds own similar machines but finance, reuse, and monetize them differently. That difference determines who compounds and who merely carries the hardware risk.
Running an AI model at scale requires renting, or building, an enormous cluster of accelerators, the networking to bind them, and the power to feed them. The cloud is the business of providing that on demand. It splits into two tiers. The hyperscalers — Amazon Web Services, Microsoft Azure, and Google Cloud — are the diversified giants that sell everything from storage to AI. The neoclouds — led by CoreWeave and Nebius — are specialists that do essentially one thing, rent out Nvidia GPUs, and have grown from nothing to a meaningful share of the market in two years. Sitting above both is the layer that actually locks customers in: the software stack, and in particular Nvidia's CUDA, the programming environment that fifteen years of developer investment have made the default way to run AI.

To understand the business, follow one accelerator through its economic life. A cloud operator commits capital before the machine earns anything. It reserves power, buys the server and network, installs the cluster, and begins recording depreciation once the asset is available for use. Customers then rent slices of time for model training, fine-tuning, or inference. The operator’s return depends on how many useful hours it can sell, the price of those hours, the cost of electricity and support, and the resale or reuse value of the machine after a newer generation arrives.
This is different from ordinary software. A software company can often serve another customer at very low incremental cost. A GPU cluster has a hard physical ceiling. An idle hour disappears forever, while a fully booked cluster may require another round of capital spending before revenue can grow. The cloud operator is therefore balancing two opposite risks: buy too little and lose customers; buy too much and own depreciating hardware that does not earn.
The same hardware can still produce very different returns in different hands. A hyperscaler can fill spare capacity with thousands of unrelated customers, bundle compute with storage and databases, move older GPUs to less demanding workloads, and fund the build from other businesses. A neocloud can offer scarce capacity quickly and optimize around one type of customer, but it has fewer alternative uses if a contract is delayed. The central question is not which company owns the newest GPU. It is which company can keep the whole fleet useful across several hardware generations.

Two business models carry the same depreciating asset
The roster spans the diversified giants, the GPU-rental specialists, the inference-optimized upstarts, and the Chinese incumbents building on a separate stack.
| Provider | Ticker | Tier | What to know |
|---|---|---|---|
| AWS | AMZN | hyperscaler | #1, ~28% share; Trainium silicon |
| Microsoft Azure | MSFT | hyperscaler | AI run-rate ~$37B, +123% |
| Google Cloud | GOOGL | hyperscaler | fastest-growing (+82%); TPUs |
| Oracle | ORCL | hyperscaler | Stargate; debt-funded |
| CoreWeave | CRWV | neocloud | ~$99B backlog; GPU-backed debt |
| Nebius | NBIS | neocloud | "sold out"; $7–9B ARR guide |
| Crusoe | PVT | neocloud | ~$30B valuation; ~4.9 GW contracted |
| Lambda | PVT | neocloud | $9B val; IPO planned H2 2026 |
| Together AI | PVT | neocloud | ~$1B ARR; inference serving |
| Groq, Cerebras, SambaNova | PVT / CBRS | inference specialist | custom inference silicon-as-a-service |
| Alibaba Cloud | BABA / 9988.HK | China | ~36% China share; AI run-rate ~$5.2B |
| Huawei Cloud | PVT | China | Ascend/CANN "CloudMatrix" stack |
| Tencent Cloud | 0700.HK | China | ~9% share; GPU-supply-constrained |
The categories in the table describe financing and customer mix, not merely size. Hyperscalers tend to own the land, buildings, power contracts, software services, and customer relationship. Their AI economics are mixed into a much larger cloud income statement. Neoclouds are closer to specialized equipment lessors: a few large customers sign long contracts, and the operator raises debt and equity to buy the machines that fulfill them. Inference specialists try to escape commodity GPU rental by controlling a purpose-built chip and serving stack.
That distinction changes what an investor must inspect. For a hyperscaler, the important questions are incremental cloud revenue, operating margin, depreciation, and whether AI increases demand for the rest of the platform. For a neocloud, the questions are customer concentration, contract enforceability, funding cost, utilization, collateral value, and the gap between the life of the debt and the useful life of the GPU. For an inference specialist, technical speed matters only if customers adopt its software and route a durable share of production traffic to it.
Hyperscalers can monetize old hardware twice
The cloud market is large and accelerating. Total enterprise cloud infrastructure spending reached $128.6B in the first quarter of 2026, up 35% year on year, the fastest growth in more than four years, on an annualized run-rate above half a trillion dollars. The three leaders hold roughly two-thirds of it.
The headline market share understates how quickly AI is reshaping these businesses. Microsoft's AI operation reached a roughly $37B annual run-rate, up 123% year on year, and its commercial book of committed future revenue hit $627B, nearly doubling. Google Cloud grew 82% in the second quarter of 2026 to $24.8B, and its backlog roughly doubled in a single quarter to about $514B. Amazon's AWS, the largest at 28% share and $37.6B in quarterly revenue, grew 28%, its fastest in fifteen quarters, though its operating margin slipped from 39.1% to 37.7% as the depreciation on all those AI servers began to bite. That margin detail is a preview of a theme in Chapter 2.11: the capital intensity of AI is starting to show up in the income statement.1
Neocloud backlog is both an asset and a financing obligation
The more novel part of this layer is the neocloud, and it is best understood through one company's numbers. CoreWeave reported first-quarter 2026 revenue of about $2.08B, up 112% year on year, which annualizes to roughly $8B. Against that, it disclosed a revenue backlog of $99.4B, including a $21B commitment from Meta and multi-year deals with Anthropic and others. The backlog is more than ten times the revenue run-rate.2
CoreWeave spent about $7.7B on capital in the quarter, close to four times its revenue, funded by roughly $25B of debt, much of it collateralized by Nvidia GPUs. It carries over 1 GW of active power against more than 3.5 GW contracted. Other providers show the same capital intensity at different scales: Nebius guides to $7–9B of committed annual revenue on a base below $2B; Crusoe is raising at around a $30B valuation with roughly 4.9 GW contracted; Lambda, valued near $9B, plans a second-half 2026 IPO following a large Microsoft deployment; and Together AI has reached roughly $1B of annualized inference revenue. A debt-financed cluster has been estimated to break even at around 70% utilization, making utilization and take-or-pay enforceability central sensitivities. Falling H100 rental prices squeeze older fleets. Nvidia's roughly $2B investments in each of CoreWeave and Nebius also connect supplier capital with customer purchases, a financing relationship examined in Chapter 2.11.
Backlog needs to be read as a schedule of obligations, not a pile of cash. A contract can cover several years, depend on the operator delivering capacity on time, and contain termination or performance provisions. The customer’s commitment may be strong while the operator still has to finance construction, absorb interest before the site opens, and replace hardware during the contract. A large backlog is evidence of demand; it is not proof of margin or free cash flow.
Consider a simplified cluster with annual fixed costs of 100. At 90% utilization, the operator spreads those costs across nine units of billable time and can earn an attractive return. At 60%, the same fixed cost is spread across six units, so cost per billable hour rises by half before considering any discount in rental price. If a newer GPU also pushes down the hourly price of the old one, revenue and collateral value can weaken together. Leverage turns what looks like a modest utilization change into an equity event.
Software determines the cost of changing chips
If the data center is a commodity, the software is not. Nvidia's CUDA has roughly six million developers and first-class support in every major machine-learning framework, and the switching cost is real: years of CUDA-optimized training pipelines and custom kernels would have to be rewritten to move off it. This is the true source of Nvidia's pricing power, more durable than any single chip.
The moat is not uniform, though, and the interesting question is where it thins. For training at the frontier, it holds firmly. For inference, it is eroding. AMD's ROCm software has closed enough of the gap that AMD claims parity by the end of 2026, and its data-center revenue grew 57%. Cross-hardware abstraction layers — OpenAI's Triton and PyTorch 2.0 among them — increasingly let a model run on AMD or a custom chip with little rewriting, which shifts the buying decision toward raw cost per token, exactly where the custom ASICs of Chapter 2.5 compete. Huawei's CANN, the Chinese alternative, claims that around 80% of standard PyTorch inference workloads run with only minor adjustments. The consensus of the evidence is that the moat is intact for training and genuinely eroding at inference, which matters because inference is now the larger share of compute.
Inference specialists need control of routing, not just faster silicon
A distinct group of companies is betting that inference will run more efficiently on purpose-built silicon than on general GPUs. Groq (valued around $6.9B) builds deterministic low-latency inference chips; Cerebras (which staged a large 2026 IPO) makes wafer-scale processors and signed a multi-hundred-megawatt OpenAI deal; and SambaNova (around $11B) offers full-stack inference systems. These companies claim several-fold lower cost per token, but the comparison depends on model, latency, utilization, and software support. Most are private or newly public, depend on a small number of anchor contracts, and carry greater execution risk.
China's constraint is fragmentation, not demand
The national split repeats the pattern of the layers above. The United States has both tiers, the hyperscalers and the neoclouds, running on Nvidia and CUDA. China has a domestic cloud market roughly a tenth the size, concentrated among three players: Alibaba Cloud at about 36% share, Huawei Cloud at 16%, and Tencent Cloud at 9%. Alibaba is the standout, its external cloud revenue growing 40% with AI products now about 30% of it and an AI run-rate near $5.2B, eleven consecutive quarters of triple-digit AI growth, which makes it the investable expression of both the Chinese cloud and the open-weight model layer.
Accelerator supply constrains the Chinese cloud. Huawei Cloud's external revenue fell 3.5% in 2025 despite a growing market, while Tencent's cloud expansion was limited by GPU availability. Huawei is responding with Ascend chips, CANN software, and CloudMatrix systems that connect hundreds of accelerators. The ecosystem remains behind CUDA in software maturity but is improving, so Chinese enterprises increasingly rent Ascend-powered compute rather than Nvidia systems.
Utilization converts a technology story into cash
Three things. The first is whether CUDA's inference moat holds or gives way, which will show up in AMD's and the custom-ASIC vendors' share of inference workloads and in whether the abstraction layers make switching genuinely painless. The second is the financial health of the neoclouds, where the signals to watch are utilization rates, any customer that fails to take contracted capacity, GPU-rental pricing, and stress in the GPU-backed debt that Chapter 2.11 treats as a systemic question. The third is margin: whether hyperscaler cloud profitability holds as AI-server depreciation mounts, the early sign of which is already visible in AWS's slipping margin.
Those three checks should be read in sequence. Software determines which hardware a customer is willing to use. Customer demand determines whether installed hardware is occupied. Utilization and rental price determine revenue. Depreciation, power, interest, and support determine how much of that revenue reaches equity. Skipping directly from “AI demand is strong” to a cloud-stock conclusion leaves out the entire mechanism that separates a good operator from an overleveraged owner of servers.
Who can retain the cloud economics
Microsoft (MSFT), Alphabet (GOOGL), and Amazon (AMZN) combine cloud exposure with diversified cash generation; Oracle (ORCL) is more concentrated in large AI contracts and carries greater financing sensitivity. Nvidia (NVDA) owns a software and systems moat across cloud providers, while Broadcom (AVGO) and Arista (ANET) supply custom silicon and networking. CoreWeave (CRWV) and Nebius (NBIS) provide more direct GPU-cloud exposure with materially greater balance-sheet, customer and utilization risk. Cerebras, Groq and SambaNova are higher-risk bets on specialized inference. Alibaba (BABA, 9988.HK) is a listed expression of China's domestic cloud and open-weight model ecosystem. These securities share demand factors but have very different equity-capture and valuation profiles.
What would falsify the cloud thesis
The bullish case for the hyperscalers weakens if cloud margins compress faster than AI revenue grows, turning the capex into a drag rather than an engine, the early warning of which is already in AWS's margin line. The neocloud thesis breaks on a utilization or funding shock, a single large customer failing to take capacity, or a downgrade of the GPU-backed debt, either of which would expose the leverage quickly, especially with rental prices already falling. And Nvidia's software moat, the most valuable asset in this layer, would be revealed as thinner than assumed if a major lab or hyperscaler moved a frontier training run onto non-CUDA hardware without a painful rewrite, which has not yet happened but is the thing to watch.
The conclusion is not that hyperscalers are safe and specialists are dangerous. It is that they are selling different claims on the same physical asset. The hyperscaler offers a diversified claim on customer relationships, software, and a fleet that can be repurposed. The neocloud offers a more concentrated claim on scarcity, contracted demand, and financial execution. The inference specialist offers a technology claim that still needs a distribution and software moat.
For the general investor, the quarterly checklist is straightforward. Compare AI-related revenue growth with total capital spending and depreciation. Track whether backlog converts into commissioned capacity on schedule. Watch rental prices by hardware generation rather than one blended price. Look for customer concentration and contract changes. Finally, ask where an aging GPU goes when its first workload moves to a newer system. A cloud business earns its premium by keeping expensive machines working; everything else is a supporting detail.
Sources
Linked evidence for this chapter's figures and load-bearing claims: 2 1
Footnotes
-
Cloud Market Annual Revenue Run Rate Topped Half a Trillion Dollars in Q1. Synergy Research Group, undated; accessed 2026-07-25. ↩ ↩2
-
CoreWeave Reports Strong First Quarter 2026 Results. CoreWeave, 2026-05-07; accessed 2026-07-25. ↩ ↩2