Part IV — Where the Bottleneck Moves Next
A shortage is not a permanent moat. It attracts capacity, redesign, substitution, and policy until the constraint moves elsewhere. The forward-looking task is to identify which stage of that migration each layer has reached—and whether the market is pricing the right bottleneck at the right time.
4.1 A shortage passes through five stages
Most AI bottlenecks follow the same sequence:
- demand outruns qualified supply;
- lead times lengthen and customers reserve capacity early;
- prices, margins, and supplier capital spending rise;
- new capacity, substitution, and redesign arrive;
- lead times shorten before reported revenue peaks.
The best financial results often appear in stage three, when the market is most tempted to value temporary scarcity as a permanent moat. The useful forward indicator is therefore the first shortening of lead times, not the last quarter of revenue growth.
Consider a component that normally takes twelve weeks to deliver. Demand suddenly rises, customers place orders early, and lead time stretches to a year. The supplier raises price, expands capacity, and reports record backlog. Customers respond by qualifying a second source, redesigning the system, and carrying more inventory. By the time the new factory opens, the original supplier may still be reporting strong revenue from old orders even though current lead time has already fallen.
That gap between operating data and financial statements creates the investment opportunity and the trap. Lead time, spot price, inventory, cancellations, and customer qualification often turn before reported revenue and margin. A shortage can be strategically real at the same moment the security has begun to discount its end.
The scarce input has already moved several times. GPUs dominated the shortage discussion in 2023. By 2026, memory, advanced packaging, electrical equipment, and deliverable power constrain more projects. Nvidia has described more than $1 trillion of Blackwell-and-Rubin order visibility through 2027, with Blackwell sold out into late 2026. TrendForce projects HBM rising from 18% of DRAM wafer input at the end of 2025 to 30% by the end of 2027, with HBM contract prices possibly doubling in 2027. Farther downstream, transformer lead times run three to four years, and roughly 7 of 12 gigawatts of US data centers planned for 2026 have been canceled or delayed for lack of power and equipment.
Those cancellations establish that a chip order is no longer enough to create usable capacity. Over the next two years, electrical equipment and commissioned power deserve at least as much attention as accelerator supply. That conclusion must be reassessed when transformer lead times, project cancellations, and equipment book-to-bill begin to normalize.
4.2 Next-generation compute moves value off the logic die
Nvidia's roadmap shows how system content is changing. Vera Rubin ships in 2026, Rubin Ultra follows in the second half of 2027 (15 exaFLOPS of FP4 inference and 384 GB of HBM per GPU, according to TrendForce), and Feynman is scheduled for 2028. Feynman's disclosed emphasis on custom HBM and 3D die-stacking indicates that more performance improvement will depend on memory and packaging rather than the logic die alone.
Counterpoint and JPMorgan project custom ASIC shipments surpassing merchant GPUs in units in 2027—roughly 12.5 million ASICs versus 10.9 million GPUs out of 23.3 million accelerators—with Broadcom holding about 60% of AI-server ASIC design share. Unit share, revenue share, and delivered computing capacity are different measures. Nvidia can continue growing revenue in frontier training and the merchant market while captive ASICs absorb more stable inference workloads. Inference already represents about two-thirds of compute, and Gartner expects it to exceed 65% of AI-IaaS spending by 2029.
The GPU-versus-ASIC contest therefore does not settle the supplier question. Memory gains content per system as HBM grows, although high margins will eventually attract supply. Networking and optics gain from larger clusters; Broadcom spans custom ASICs, AI Ethernet, and co-packaged optics, while smaller component suppliers carry greater customer-concentration risk. Advanced packaging remains a near-term deployment constraint. For each group, capacity additions and the valuation already paid determine whether content growth becomes an attractive security return.
4.3 Model progress moves the commercial bottleneck into production reliability
Model capability is improving much faster than measured enterprise profit. Epoch AI measures frontier training compute growing about fivefold a year and algorithmic efficiency improving about threefold a year. METR finds that the length of tasks an AI agent can complete reliably is doubling roughly every three months. Yet MIT found that about 95% of enterprise AI pilots delivered no measurable profit. The commercial bottleneck is therefore shifting from model capability to reliable production use.
Long-horizon reliability connects those two observations. If multi-hour and multi-day tasks become dependable enough for production, more pilots can turn into recurring contracts and measured customer savings. In the meantime, infrastructure demand can still grow through volume: Goldman Sachs projects token demand rising roughly 24-fold by 2030 as inference cost falls 60–70% a year and agentic workflows consume 5–30 times the tokens of a chat query. Lower unit cost can expand total consumption, which helps explain why Gartner projects aggregate AI spending rising from $2.5 trillion in 2026 to $3.3 trillion in 2027 even as per-token prices fall.1
This shift creates several conditional exposures. Inference infrastructure and orchestration benefit if token volume grows faster than unit prices fall. Physical AI—simulation, robotics, actuators, sensors, and batteries—adds another source of compute demand, although market estimates vary enormously: Goldman sizes humanoids at $38 billion by 2035, while Morgan Stanley reaches $5 trillion by 2050. Open-weight convergence may compress closed-model margins while expanding demand for self-hosted compute. These conclusions reverse if production adoption stalls, token growth fails to offset price declines, or physical-AI deployments remain experimental.
4.4 Power scarcity becomes a siting and cost-allocation problem
Major forecasts agree that data-center electricity demand will rise sharply, although they differ on scale. The IEA sees demand roughly doubling from 415 TWh in 2024 to about 945 TWh by 2030. Goldman Sachs estimates global data-center power demand rises 165% by 2030 and US demand doubles from 31 GW in 2025 to 66 GW by 2027. Morgan Stanley estimates a roughly 49-gigawatt US shortfall and data centers reaching about 18% of US electricity by 2030.2
The immediate constraint is often equipment and delivery time rather than generation in the abstract. Transformer lead times reach three to four years, gas turbines are sold out through 2030, GE Vernova's backlog has reached 116 GW, and 7 GW of planned 2026 capacity has been canceled or delayed. Electrical-equipment and turbine suppliers such as Eaton, Siemens Energy, GE Vernova, Hitachi Energy, and Quanta have multi-year backlog visibility, but their re-rating makes lead-time and book-to-bill normalization important downside indicators. On-site generation can bridge grid delays: Bloom Energy signed roughly $7.65 billion of binding data-center fuel-cell contracts in about ninety days, alongside a $25 billion Brookfield commitment. Existing nuclear operators including Constellation, Vistra, and Talen can monetize current fleets through long-term PPAs, while most SMR economics depend on deployment after 2030. Liquid-cooling suppliers gain content as rack density rises. The distinction between contracted current assets and long-dated projects matters more than a generic “power” label.
4.5 Capital may become the final constraint after the physical ones
The financing question can be tested by comparing investment obligations with observable revenue and cash generation. JPMorgan estimates $5.5 trillion of AI capital spending through 2030, including $4.1 trillion financed with debt, while Morgan Stanley estimates a $1.5 trillion financing gap. Sequoia's David Cahn calculates that roughly $1.5 trillion of 2026 infrastructure spending would require about $3 trillion of revenue to justify it. Bain estimates that the industry needs about $2 trillion of annual AI revenue by 2030 and could fall roughly $800 billion short.
Realized GPU economic life is a major sensitivity. Michael Burry argues that hyperscalers understate depreciation by roughly $176 billion across 2026–28 by using five-to-six-year accounting lives for chips with two-to-three-year economic lives. Goldman's sensitivity analysis shows that shortening useful life from five years to three raises cumulative 2026–31 depreciation from about $3 trillion to about $4 trillion. Whether that becomes an earnings problem depends on utilization, rental prices, inference revenue, and enterprise adoption. Credit can reveal stress early: S&P has cut Oracle to one notch above junk, while its credit-default swaps reached multi-year highs. Single-name credit spreads should therefore be read alongside equity multiples and reported backlog.
Private-credit originators including Apollo, Blue Owl, and Blackstone earn fees and yield premiums by financing the build, but opaque collateral and borrower concentration can delay recognition of losses. Public-market leadership has also broadened: the Magnificent Seven remain about 34% of the S&P 500 but lagged the other 493 constituents in the first half of 2026, while the Russell 2000 Growth index rose 17%. That rotation shows that rising AI spending does not guarantee continued outperformance by the largest technology companies.
4.6 Geopolitics creates several parallel bottleneck timelines
The policy trajectory has settled into an unstable pattern: the US oscillates between denial and monetization (the diffusion rule rescinded, a 15% China-sales tax imposed, H200 sales nominally allowed but with essentially zero China data-center revenue actually flowing), while China builds a parallel stack. The forward evidence suggests the export-control debate may already be losing relevance: domestic accelerators reached 41% of China's market in 2025, and Morgan Stanley projects 86% by 2030. The reasoning is that restriction accelerates the very self-sufficiency it aims to prevent — Nvidia's own warning that continued curbs "hand China's market to Huawei permanently" is being borne out.
Three timelines matter. Sovereign AI demand adds a new customer group for the US stack: Saudi Arabia's HUMAIN has announced more than $100 billion and 2,200 MW, the UAE's Stargate reaches 5 GW, and sovereign AI-infrastructure deployment totaled $66 billion in 2025. China's parallel supply chain—Huawei, SMIC, and CXMT—gains protected domestic demand but remains limited by SMIC yields and domestic HBM. Taiwan concentration persists: about 70% of sub-2nm capacity is expected to remain in Taiwan through 2030 despite Arizona expansion. Packaging localization can reduce some delivery risk, but it does not replicate the complete leading-edge ecosystem. These are different exposures and should not be collapsed into a single geopolitical score.
4.7 History supplies base rates, not an answer
AI is technologically new; capital cycles are not. Four historical patterns provide useful—not mechanical—base rates.
| Analogue | What persisted | What disappointed investors | Relevant test for AI |
|---|---|---|---|
| 1990s telecom fiber | Data traffic and useful infrastructure | Utilization, pricing, leverage and duplicate networks | Can token growth fill commissioned capacity before debt and depreciation mature? |
| Cloud build-out, 2010s | Secular workload migration and hyperscaler scale | Falling unit prices and weak economics for undifferentiated hosts | Who retains margin as inference becomes cheaper? |
| Memory cycles | Rising long-run bit demand | Capacity additions, inventory correction and peak-earnings valuation | Does HBM qualification delay the normal supply response, and for how long? |
| China solar/EV localization | Rapid scale, falling global cost and strategic independence | Overcapacity, price wars and weak minority-shareholder returns | Does mandated AI capacity create durable economics or only strategic output? |
The common lesson is that useful infrastructure can coexist with poor security returns. Demand growth does not protect a supplier whose capacity, leverage or valuation assumed even faster growth. The relevant discipline is to estimate each shortage's scarcity half-life and compare it with announced capacity, qualification time and the multiple already paid.
4.8 Four futures constrain the base case
The analysis uses three weighted scenarios and one unweighted tail risk for the two years ending 25 July 2028. The probabilities express judgment; they are not measured frequencies.
| Scenario | Weight | Numerical or observable confirmation | Probability moves higher when | Portfolio implication |
|---|---|---|---|---|
| Bear: financial air pocket | 20% | At least one hyperscaler cuts planned capex; AI-linked credit spreads widen; utilization or backlog conversion falls materially | Two signals persist for a quarter | Reduce neoclouds and peak-multiple suppliers; favor balance-sheet strength and diversified cash generation |
| Base: slower monetization | 55% | Capex grows but decelerates; production adoption improves unevenly; shortages normalize at different speeds | Revenue and utilization improve without a broad agent breakthrough | Prefer qualified bottlenecks selectively; apply valuation discipline and monitor capacity response |
| Bull: production-agent adoption | 25% | Audited customer P&L benefits broaden; production conversion rises; disclosed AI revenue covers more depreciation | Two consecutive reporting periods confirm revenue and margin | Add workflow owners and higher-beta infrastructure while retaining physical chokepoints |
| Taiwan disruption tail | Unweighted tail | Material interruption to leading-edge production or logistics | Geopolitical indicators move beyond exercises and rhetoric | No listed basket fully hedges the physical interruption; reduce aggregate exposure and liquidity risk |
Probability updates should be rule-based. One company announcement is insufficient. A five-percentage-point shift requires either two independent indicators or one audited system-level change, such as a sustained hyperscaler capex reduction or a broad production-ROI dataset.
Base case: monetization broadens slowly. Enterprise returns improve in deep workflows and remain uneven elsewhere. Capital spending stays high but decelerates, policy truces hold, and no major financing node breaks. Power, memory, packaging, neutral chokepoints, and credit providers retain demand, while highly valued companies remain sensitive to any slowdown. Balance-sheet quality, earnings quality, and entry valuation determine the security outcome.
Bull case: agents reach dependable production use. If reliability improves enough for AI to compete for labor budgets, the addressable market expands and infrastructure demand is pulled forward. Confirmation requires broad production deployments, audited customer profit-and-loss benefits, and improving gross margins at model providers.
Bear case: financing weakens before the technology does. A hyperscaler cuts capital spending, GPU-backed debt comes under stress, an efficiency shock reduces capacity needs, or a model laboratory struggles to refinance. Circular financing can transmit the shock faster than a conventional supply chain. Neoclouds, pre-revenue companies, and suppliers valued on peak backlog are more exposed; diversified companies with positive free cash flow are better positioned.3
Tail risk: a Taiwan disruption. It is low probability in the scenario framework but severe enough that no listed basket offers a complete hedge.
4.9 Where the market may be mistiming the transition
Four timing mismatches deserve attention. Electrical constraints may last longer than current project schedules imply because transformers and interconnection queues take years to clear. Pre-revenue nuclear and humanoid companies may be valued on 2030s outcomes even though they cannot solve a 2027 capacity shortage. Chinese open-weight models and developer adoption may be further advanced than national-leadership comparisons suggest, affecting competition in third-country markets. Credit conditions may turn before equipment orders do as more of the build shifts from internal cash flow to debt. Each case requires a valuation check; an industrial trend can be correct and still be fully reflected in the security price.
4.10 The next two years have a sequence of proofs
No single metric settles the financing debate. The closest available test is whether disclosed hyperscaler AI revenue grows fast enough to cover depreciation, interest, power, and ongoing upgrades. Several near-term events provide additional evidence.
| Date / window | Catalyst | Why it matters |
|---|---|---|
| Late July 2026 | Microsoft, Meta, Amazon earnings | tests whether investors reward or penalize higher capital spending |
| H2 2026 | Vera Rubin ships; AMD Helios; co-packaged optics; HBM4 volume | tests whether the next system generation enters volume production on schedule |
| Late 2026 | AI IPO window: rumored OpenAI S-1, SpaceX, Databricks | tests whether private valuations hold under public-market pricing |
| 10 Nov 2026 | US "50% affiliates" export rule scheduled to snap back | re-extends Entity-List controls to majority-owned subsidiaries |
| ~27 Nov 2026 | US–China minerals truce expires | watch for re-tightening of rare-earth and gallium controls |
| Through 2026–27 | TSMC Arizona 2nm; SMIC yields; transformer/turbine lead-times; single-name AI-debt CDS | provides evidence on manufacturing, power, and credit constraints |
| 2027 | Rubin Ultra; ASIC unit-crossover; Intel 14A external test | whether the compute roadmap and the custom-silicon shift play out |
The next several quarters will clarify the durability of capital spending, the direction of export controls, and the pace at which physical constraints ease. Revenue coverage, useful life, utilization, and credit spreads should be assessed together; none is sufficient on its own.
4.11 A monitoring table is more useful than a permanent forecast
No single disclosed metric fully measures AI's economic return. The operating dashboard therefore uses a linked set of indicators.
| Indicator | Current interpretation at cutoff | Bull confirmation | Bear warning | Review cadence |
|---|---|---|---|---|
| Hyperscaler capex and FCF after capex | Spending remains historically high | Capex converts to cloud/AI revenue without sustained FCF deterioration | Planned capex cut or debt-funded acceleration without revenue coverage | Quarterly |
| Disclosed AI revenue | Partial and inconsistently defined | Comparable disclosure broadens and covers more depreciation | Bundled metrics disappear or margins weaken | Quarterly |
| Accelerator useful life | Accounting lives generally exceed rapid product cadence | Older fleets retain utilization and pricing | Rental prices or utilization fall before debt maturity | Quarterly |
| Neocloud utilization and credit spreads | Contracted demand is high; financing remains central | Utilization, cash conversion and spreads improve together | Backlog delays, customer concentration or spread widening | Monthly/quarterly |
| HBM price, qualification and wafer allocation | Scarcity supports pricing | HBM4 yield and qualification remain tight while demand grows | Capacity and inventory rise faster than qualified demand | Monthly/quarterly |
| Advanced-packaging capacity | Still a critical delivery constraint | Capacity additions are absorbed without lead-time collapse | Lead times and pricing normalize faster than expected | Quarterly |
| Transformer and turbine lead times | Multi-year physical constraint | Book-to-bill and margin remain firm as capacity expands | Cancellations, queue withdrawals or rapid lead-time compression | Monthly/quarterly |
| Agent pilot-to-production conversion | Broad deployment remains limited | Audited production conversion and P&L impact rise | Cancellation rates remain high and contracts stay experimental | Semiannual |
| Inference cost at fixed capability | Falling rapidly | Volume and useful-task demand grow faster | Price decline outruns demand and supplier margin | Quarterly |
| Ethernet versus proprietary fabric | Open networking is gaining relevance | Merchant Ethernet/optics take sockets without severe pricing pressure | Proprietary integration retains the frontier or merchant margins compress | Semiannual |
| SMIC yield and domestic HBM | China's principal physical ceiling | Verified yield, capacity and HBM qualification improve | Roadmap claims rise without shipment evidence | Quarterly |
| Export-control and ownership status | Transactional and entity-specific | Stable licensing and access rules | New entity, affiliate or ownership restrictions | Event-driven |
| Taiwan advanced-capacity distribution | Diversification is real but incomplete | Leading-edge and packaging capacity commissions outside Taiwan | Delays or concentration persists beyond disclosed plans | Semiannual |
Every security in Part V maps to at least two of these indicators and a stale-after date. The mapping allows conclusions to be updated when the evidence changes.
The table should be read as a chain rather than thirteen independent signals. Application conversion and inference demand determine whether outside cash is growing. Cloud revenue, utilization, and capex show how that demand reaches infrastructure. HBM, packaging, networking, and electrical lead times show where delivery is constrained. Credit spreads and free cash flow show who is financing the wait. Policy and Taiwan determine which routes remain available.
A single warning need not end the thesis. Falling HBM price may reflect healthy capacity while application demand continues to strengthen. A capex pause may improve free cash flow without collapsing utilization. The regime changes when several linked indicators turn in the same direction—for example, weaker application conversion, lower cloud utilization, canceled projects, shorter equipment lead times, and wider credit spreads.
Forecasts should remain conditional. Record the current constraint, the evidence that it is moving, the companies exposed to that change, and the assumptions already embedded in valuation. Revise the conclusion when several linked indicators change direction, rather than waiting for reported revenue to confirm what operating data already shows.
Sources
Linked evidence for this chapter's figures and load-bearing claims: 3 1 2
Footnotes
-
What the capex boom means for stock investors. Goldman Sachs, 2026-07-21; accessed 2026-07-25. ↩ ↩2
-
Energy and AI. International Energy Agency, 2025-04-10; accessed 2026-07-25. ↩ ↩2
-
The AI investment race. Bank for International Settlements, 2026-07-14; accessed 2026-07-25. ↩ ↩2