Chapter 2.5 — AI Chips & Accelerators
This is the layer everyone means when they say "AI chips," and it holds the central paradox of the whole build. By unit count, custom chips designed by the hyperscalers are about to overtake merchant GPUs. By revenue and profit, Nvidia still takes roughly four dollars in five and shows no sign of letting go. Both are true because units, revenue, and computing power are three different scoreboards. The durable question for an investor is not "who ships the most chips" but "who keeps the margin," and the answer runs through Nvidia's software and system integration, not the silicon alone.
An AI accelerator is a chip built to do the one kind of arithmetic that neural networks need, the multiplication of large matrices, thousands of times in parallel. A general-purpose CPU does a few things in sequence very fast; an accelerator does one thing across thousands of cores at once, which is exactly what training and running a model requires. The performance that matters is measured in floating-point operations per second (FLOPS), increasingly at low precision (the FP4 and FP8 formats that trade a little accuracy for a lot of throughput), and in how much high-bandwidth memory sits next to the compute. The most important shift in this layer is that the unit of sale is no longer the chip. Nvidia now sells the rack, an integrated "AI factory" of dozens of GPUs, CPUs, switches, and liquid cooling, and prices it accordingly.
AI Accelerator Market Structure
The roster of accelerators divides into three groups: the merchant GPUs anyone can buy, the custom ASICs the hyperscalers design for themselves, and the Chinese chips built behind the export-control wall.
| Chip / family | Maker | Type | What to know |
|---|---|---|---|
| Blackwell (GB300) → Rubin | Nvidia | merchant GPU | the standard; CUDA lock-in; ~75% of 2026 accelerator revenue (estimate) |
| Instinct MI300 → MI400 | AMD | merchant GPU | the credible #2; OpenAI & Meta anchors |
| Gaudi → Jaguar Shores | Intel | merchant GPU | the also-ran; relevance pushed to 2027 |
| TPU (Ironwood, v7) | custom ASIC | the most mature hyperscaler chip | |
| Trainium / Inferentia | AWS | custom ASIC | largest deployed ASIC fleet (500k+) |
| Maia | Microsoft | custom ASIC | in-house Azure silicon |
| MTIA | Meta | custom ASIC | in-house inference silicon |
| XPU (10 GW program) | OpenAI + Broadcom | custom ASIC | first volume ~2027 |
| Ascend 910C / 950 | Huawei | China champion | ~600k units planned 2026; self-HBM |
| MLU (Siyuan) | Cambricon | China (listed) | own architecture; ~500k target |
| BR-series | Biren | China | training-focused |
| MXC / C-series | MetaX | China | domestic general-purpose GPU / accelerator program |
| MTT | Moore Threads | China | GPU-style |
| DCU | Hygon | China (listed) | x86-linked; catalogue-cleared |
| (ASIC design partners) | Broadcom, Marvell | enablers | design ~95% of the hyperscaler ASICs |
The names to carry forward: Nvidia and AMD are the merchant GPUs; Broadcom and Marvell are the arms dealers who design everyone else's custom chips; Google, Amazon, Microsoft, Meta, and OpenAI are all building their own; and Huawei and Cambricon are China's champions, trailed by Biren, MetaX, Moore Threads, and Hygon.
The one-year cadence, and why it is a moat
Nvidia has put itself on a one-year architecture rhythm that the rest of the industry struggles to match. Its Vera Rubin NVL72 rack, ramping into production in 2026, combines 72 Rubin GPUs and 36 Vera CPUs and delivers 3.6 exaFLOPS of NVFP4 inference with 20.7 TB of HBM4 across the rack. Nvidia had not published a rack-power specification at the cutoff, so third-party power estimates should not be presented as product facts. Rubin Ultra follows on the roadmap in 2027, with rack power discussed near 600 kW ⚠️. The economics of selling systems rather than chips are enormous: a GB200 or GB300 NVL72 rack runs around $3M, a Vera Rubin rack an estimated $5–7M, against bare GPUs at $30–40k. The full physical BOM, evidence grades, and normalized 100 MW implications are in §2.13.12
The cadence is itself the moat. Shipping a materially better rack every twelve months, co-designed across compute, networking, and cooling, raises the switching cost for any customer and strains every rival to keep pace. The risk is the mirror image: a one-year cadence stresses the entire supply chain beneath it, from the advanced packaging in Chapter 2.7 to the power in Chapter 2.3, and any slip would ripple outward. Reports of a possible Rubin Ultra "Kyber" delay circulated and were denied in 2026, the kind of rumor that will recur precisely because the cadence leaves no slack.
AMD, Intel, and the merchant challengers
The merchant-GPU challenge to Nvidia is real but narrow. AMD is the credible number two, with its MI400 family on TSMC's 2nm process and its double-wide Helios rack, due in the third quarter of 2026, giving it a genuine rack-scale product for the first time; AMD claims the MI455X delivers many times the token throughput of its predecessor. The demand base finally arrived with the OpenAI agreement to deploy 6 GW of AMD chips, sweetened with warrants for up to 160M AMD shares, plus a multi-gigawatt Meta commitment. AMD's data-center revenue reached $5.8B in the first quarter of 2026, up 57%, though its share of AI accelerators remains in the mid-single digits.
Intel is the also-ran of this layer. Its Gaudi line missed even a modest $500M target, and its go-forward story is Crescent Island, an inference GPU in customer testing in the second half of 2026, and Jaguar Shores, a 2027 rack-scale design built around silicon photonics. Intel's relevance in accelerators is now a 2027 question, and the roughly 10% US government equity stake in it (Chapter 2.8) is a bet on its foundry, not its GPUs.
The custom-ASIC surge, and the arms dealers who enable it
The most important competitive shift in this layer is that the hyperscalers are increasingly designing their own chips. Custom application-specific chips (ASICs) are growing at roughly a 45% annual rate against about 16% for merchant GPUs. Google's TPU is the most mature, now on its seventh generation ("Ironwood"); Amazon has deployed over 500,000 Trainium chips, the largest ASIC fleet by unit count; Microsoft has Maia and Meta has MTIA; and OpenAI signed a 10 GW custom-silicon program with Broadcom in October 2025, first volume expected in 2027.
None of the hyperscalers designs these chips alone. Broadcom and Marvell are important enablers of custom silicon and can participate across several customer programs. That diversification reduces dependence on a single accelerator architecture but does not eliminate customer concentration, program timing or valuation risk. Industry forecasts that ASIC unit shipments overtake merchant GPUs should be treated as scenario inputs rather than a guarantee that design-service economics accrue equally to both suppliers.3
Units, revenue, and computing power — three scoreboards
The competitive map only makes sense once you separate the scoreboards, because unit share is not revenue share and neither is compute share. By units, ASICs are about to lead. By revenue and profit, Nvidia dominates, with J.P. Morgan estimating about 75% of data-center accelerator revenue in 2026, protected by the CUDA software ecosystem of Chapter 2.4. By computing power at the frontier, Nvidia is more dominant still, because the custom chips mostly serve captive internal inference rather than frontier training.
This is the sharpest bull-bear debate in the layer. The most aggressive bears see Nvidia's inference share falling from above 90% toward 20–30% by 2028 as inference shifts to cheaper custom chips; Morgan Stanley's blunt counter is that custom silicon does not threaten Nvidia's dominance at all. The more cautious reading, and the one this atlas holds, is that ASICs take the low-margin, internal-inference slice while Nvidia keeps the high-margin merchant market, the training frontier, and the software tax, so the "crossover" is a unit-count event that need not become a revenue event. It remains the single biggest swing factor for Nvidia sentiment regardless.
America's chips, China's substitutes
The national split in this layer is the sharpest in the atlas. The United States, through Nvidia, AMD, and Broadcom, owns the frontier of AI compute and the software that runs on it. China has been cut off from that frontier and is building a parallel one under duress, with a full roster of its own. Huawei's Ascend line is the national champion, with something like 600,000 Ascend 910C accelerators planned for 2026 and a roadmap (the Ascend 950 series and beyond) that adds in-house high-bandwidth memory. Cambricon (寒武纪, 688256.SH) is the listed pure-play, tripling output toward roughly 500,000 accelerators, and behind them sit Biren, MetaX, Moore Threads, and Hygon (海光, 688041.SH), whose x86-linked DCU cleared the domestic security review. Nvidia's share of the Chinese AI-chip market has fallen from about 95% in 2022 to effectively zero, not because Chinese chips are better but because Beijing has mandated domestic silicon and Washington has restricted the good American parts.
The gap is real and worth stating precisely. Chinese accelerators trail on a per-chip basis, and Huawei's claim that an Ascend cluster beats Nvidia's is a system-level argument that uses several times as many chips and far more power to get there. Its software stack, Huawei's CANN, lags CUDA's fifteen-year head start badly. And the true ceilings on the Chinese ramp are not design but manufacturing: SMIC's yields and the domestic supply of high-bandwidth memory, both covered in later chapters. This is the layer where the bifurcation of Chapter 2.12 is most concrete, two separate compute ecosystems that no longer interoperate.
The roadmap evidence to watch
Three questions matter. The first is whether the ASIC unit-crossover becomes a revenue-and-margin crossover or stays confined to captive internal inference; watch whether Google externalizes its TPU beyond its own cloud and whether OpenAI's Broadcom silicon ships at volume in 2027. The second is the durability of Nvidia's inference share, the single biggest swing factor for its sentiment, where the bear case is contested and probably too aggressive but is the number the market fixates on. The third is the Chinese ramp, gated not by ambition but by SMIC capacity and domestic memory, where any breakout past those ceilings would reshape the Asian compute map. Underneath all three runs the cadence-and-supply-chain question: whether Nvidia can hold its one-year rhythm without a slip that would ripple through packaging, memory, and power.
The trade: own the integrator and the arms dealers
Nvidia (NVDA) remains the system integrator to beat, protected by CUDA, networking and an annual platform cadence; its risks include inference-share pressure, customer internalization and policy exposure. Broadcom (AVGO) and Marvell (MRVL) provide custom-silicon exposure across hyperscaler programs while adding networking and optics exposure; program concentration and valuation differ materially between them. AMD (AMD) is a merchant second source with real rack-scale demand but a smaller installed software base; Intel's listed thesis increasingly depends on foundry execution. On the Chinese side, Cambricon and Hygon are listed expressions of mandated domestic demand, but access, policy and localization expectations can dominate near-term fundamentals.
What would break this
The Nvidia thesis breaks if inference-share loss becomes real in the revenue line rather than the rumor mill, which a genuine, externally-adopted hyperscaler ASIC operating at training scale would signal. It also breaks, from the other direction, if a second efficiency shock in the models layer cuts the compute intensity of inference faster than volume grows, or if the one-year cadence slips and the supply chain seizes. The Chinese thesis changes if SMIC and domestic high-bandwidth memory break their current ceilings, letting Ascend and Cambricon ship at true frontier scale, which would turn a mandated, subsidized market into a competitive one and pressure the American incumbents in the third-country markets that Chapter 2.12 calls the real prize.
Sources
Linked evidence for this chapter's figures and load-bearing claims: 3 1 2
Footnotes
-
NVIDIA Vera Rubin Pod: Seven Chips, Five Rack-Scale Systems, One AI Supercomputer. NVIDIA, undated; accessed 2026-07-25. ↩ ↩2
-
AI server rack power roadmap. TrendForce, 2026-06-25; accessed 2026-07-25. ↩ ↩2