Interconnect & networking

How scale-up and scale-out networks turn purchased accelerators into effective compute—and who captures the content.

Chapter 2.6 — Interconnect & Networking

Ten thousand accelerators can deliver the useful work of far fewer if they spend their time waiting on one another. Networking is therefore not peripheral plumbing; it determines how much purchased compute becomes effective compute. The investment question is who captures the rising content without being designed out by open standards or integration.

Training a frontier model spreads the work across a cluster that can number more than a hundred thousand accelerators, and a cluster is only as fast as the slowest link between its chips. Interconnect is that link, and it comes in two layers. Scale-up is the very fast, short-range fabric that binds the dozens of GPUs inside a single rack into one giant processor; Nvidia's version is called NVLink. Scale-out is the broader network that connects thousands of racks across a data center; here the contest is between Nvidia's InfiniBand and ordinary Ethernet. As clusters grow, networking claims a rising share of the total bill, which is why this once-obscure layer is now a market that the biggest names in silicon are fighting over.

The easiest way to understand the problem is to imagine one accelerator finishing its calculation before the data it needs has arrived from another. It cannot move on, so it waits. Thousands of expensive chips can be powered, cooled, and depreciating while a communication bottleneck leaves part of their theoretical performance unused. The owner paid for installed compute but receives less effective compute.

Large model training creates this problem repeatedly. Each accelerator calculates on a portion of the model or data, then exchanges results so the group can update a common set of parameters. The larger the cluster, the more devices must coordinate and the more costly any delay becomes. In inference, especially for very large models or many simultaneous users, the system must also move weights, cached context, and intermediate results quickly enough to meet response-time targets.

Networking therefore creates value in two ways. It can shorten the time required to train a model, allowing the same cluster to complete more experiments. It can also raise cluster utilization, so a larger share of the capital already spent on accelerators produces billable work. This is why the right end metric is not switch ports or optical-module shipments. It is how much GPU waiting time the network removes.

A growing market contains two different network battles

Every AI cluster needs a scale-up fabric inside the rack and a scale-out network across racks. As clusters grow, both consume more switch stages, links, and optics; the addressable revenue grows even when accelerator unit growth slows.

2.6 tam

Nvidia's NVLink for scale-up and InfiniBand for scale-out are mature, high-performance technologies tied to its ecosystem. Industry coalitions have developed UALink for scale-up and Ultra Ethernet for scale-out so buyers can mix hardware from multiple vendors. Broadcom shipped its own Ethernet-based scale-up technology before UALink silicon reached the market, giving it an earlier commercial position. Arista is also backing Ethernet in scale-out. Production deployments, not standards membership alone, will determine adoption.

A growing market will not distribute profit evenly

The layer runs from the switch silicon through the connectivity components to the optics, and Broadcom appears at almost every level.

PlayerTickerRole
NvidiaNVDANVLink (scale-up) + InfiniBand / Spectrum-X (scale-out)
BroadcomAVGOTomahawk switches; Scale-Up Ethernet; co-packaged optics
Arista NetworksANETscale-out Ethernet switches
MarvellMRVLcustom interconnect, DSPs
Astera LabsALABretimers / connectivity fabric
CredoCRDOactive electrical cables
CoherentCOHRoptical components / lasers
LumentumLITElasers / optics
FabrinetFNoptical-module manufacturing

Each row captures a different part of the economics. The switch-chip designer controls how traffic is scheduled and can earn architecture-like margins. The system vendor combines chips, software, and support into a network the customer can operate. Retimers and active cables preserve signal quality as speed and distance rise. Optical-component suppliers convert electrical signals into light, while manufacturing partners assemble modules to the customer’s design.

Revenue can grow throughout the chain while profit pools move. If a switch vendor integrates functions that previously sat in a separate module, system content rises but an independent supplier can lose its socket. If an open standard expands the addressable market, it can also make parts more interchangeable and place pressure on price. The investor must identify not only which component count rises, but who controls the specification and qualification.

Scale-up and scale-out create different control points

The standards contest has two different markets. In scale-up, Nvidia's NVLink moves roughly 1.8 terabytes per second per GPU. UALink silicon arrived later, so hyperscalers seeking a 2026 product leaned on Broadcom's Scale-Up Ethernet through its Tomahawk switches. Broadcom therefore has the earlier commercial position in open scale-up systems, while UALink still has to prove volume adoption. In scale-out, Nvidia's InfiniBand competes with Ethernet. Arista is backing Ethernet and launched 1.6-terabit products in 2026; production socket share and pricing will show whether that momentum persists.

The distinction matters because the two networks solve different failure modes. Inside a rack, the customer wants dozens of accelerators to behave almost like one device. Latency must be extremely low, the topology is tightly controlled, and integration with the accelerator matters. Across racks, the customer needs to route traffic through a much larger facility, tolerate failures, and operate the network with familiar tools. A technology can be strong in one domain without winning the other.

It also changes the meaning of “open.” A published standard does not create a market by itself. Customers need interoperable chips, switches, cables, software, diagnostic tools, and support that ship on the same schedule. The incumbent’s proprietary product can remain rational if it delivers sooner and reduces integration risk, even when the customer would prefer more suppliers. Open standards become economically important only after they become complete, qualified systems.

This is why Broadcom is the structurally advantaged name in the whole layer. It wins on the custom ASICs of Chapter 2.5, on both scale-up and scale-out Ethernet switching here, and on the co-packaged optics below, which means it collects revenue almost regardless of which hyperscaler, which chip, or which networking standard prevails. Arista owns the scale-out Ethernet opportunity, and a cluster of connectivity specialists, Astera Labs and Credo among them, supply the retimers and cables that hold these enormous fabrics together as they scale.

Moving optics toward the package redraws the value chain

The physical limit the whole layer is running into is that pushing electrical signals fast enough over copper burns too much power and reaches only so far. The answer is optics, and 2026 is the year it moves from pluggable modules onto the switch package itself, so-called co-packaged optics. Nvidia's Spectrum-X Photonics switches, built on optical engines made with TSMC, and Broadcom's Tomahawk-based optical switch line are both arriving this year, and the same two companies that lead switching also lead optics. The transition is gradual, and moving optics onto the accelerator package rather than the switch is still years away, with the reliability of on-package lasers the technical crux, but it opens a durable new market for the laser and optical-component makers, Coherent, Lumentum, and Fabrinet.

For the general reader, the trade-off is similar to moving a utility closer to the machine that consumes it. A pluggable optical module is easier to replace when it fails. Integrating the optical engine next to the switch can save power and improve density, but a failure can affect a more valuable assembly and demands better manufacturing yield and service design. The transition therefore raises both content opportunity and qualification risk.

That is why a component announcement should not be treated as immediate volume revenue. The optical engine must meet thermal and reliability targets, qualify with the switch platform, enter production at acceptable yield, and remain in the architecture long enough to recover development cost. Supplier timing may lag the platform announcement by several quarters, and the highest technical content does not always carry the highest incremental margin.

More networking content does not guarantee better supplier returns

Interconnect revenue can grow faster than accelerator units because larger clusters require more switch stages, retimers, cables and optical links. That attractive content-per-system direction still does not make every supplier equally attractive. Switch silicon and network operating systems capture architecture value; optical modules and cables generally face more manufacturing competition, qualification churn and price erosion. Co-packaged optics can raise technical barriers, but it can also transfer value from a replaceable module into the switch vendor's integrated platform.

The layer also carries unusually high customer and program concentration. A merchant supplier may win a large hyperscaler design and report several years of rapid growth, yet the economics can reverse when that customer internalizes a component, changes topology or shifts the next platform to a rival. Investors should therefore track gross margin, customer concentration and content per deployed accelerator alongside headline networking revenue. The $25B forecast establishes the market opportunity, not the margin pool available to every participant.1

Work from one delayed training job to the supplier thesis

Consider a cluster owner whose accelerators are busy only part of the time because each synchronization step waits on congested links. Buying more GPUs may increase the installed performance figure without removing the delay. Replacing switches, adding optical capacity, improving traffic scheduling, or changing the fabric can be more valuable because it releases productive time from equipment the owner already bought.

That operational improvement creates several possible revenue pools. A switch-silicon vendor can sell a faster generation with more bandwidth. A network-system company can sell the switches, operating software, telemetry, and support needed to deploy it. Retimer, cable, and optics suppliers gain content as link count and speed rise. The cloud owner may also capture value if the upgraded cluster completes more training runs or serves more inference requests without adding the same amount of compute.

The gains are related, but they cannot all be credited with the full value of the recovered GPU time. An investor must locate the binding component and then ask how the customer buys it. If the hyperscaler designs its own switch around merchant silicon, the system vendor may not participate. If the platform owner bundles networking into a complete rack, independent components may gain physical volume while losing pricing power. If a transition from pluggable optics to co-packaged optics removes a replaceable module, value can move toward the switch vendor even as total optical content rises.

A practical thesis therefore needs four observations. Communication overhead must be material in the target workload. The proposed architecture must reduce it in production rather than in an isolated benchmark. The supplier must own a qualified design position that survives the next generation. And the expected content growth must be large relative to that supplier’s existing revenue without depending on one customer forever. This chain is harder to prove than “AI needs faster networks,” but it distinguishes a durable control point from a temporary component shortage.

Chip constraints make network efficiency more valuable in China

The interconnect layer is another American stronghold, owned by Nvidia, Broadcom, and Arista. China's answer is instructive and ties back to Chapter 2.5. Because its individual Ascend chips trail Nvidia's, Huawei compensates at the level of the network, building enormous clusters: its Atlas SuperPods bind thousands of Ascend chips, over eight thousand in one system, into a single machine using Huawei's own interconnect fabric. The strategy is to reach frontier-scale computing power through system-level density and networking rather than through per-chip superiority, which makes the interconnect the crucial enabler of China's whole compute effort. It is also years behind on the software and optics that make the Western fabrics efficient.

Open standards matter only when they become deployed systems

Watch whether open Ethernet genuinely displaces Nvidia's proprietary stack or whether Nvidia's vertical integration keeps it locked at the frontier, and in particular whether the UALink standard becomes real in deployments or whether Broadcom's head-start Ethernet keeps winning the sockets. Watch Ultra Ethernet's share against InfiniBand in newly-built training clusters, the clearest measure of the open coalition's progress. Watch the volume ramp of co-packaged optics in the second half of 2026, and the reliability of the on-package lasers. And watch whether Nvidia opens NVLink, a move that would concede the standards war while trying to keep the customers.

Who can retain the networking economics

Broadcom (AVGO) has the broadest public-market exposure in this layer because it captures value across custom silicon, switching, and optics under several standards outcomes. Arista (ANET) is a more concentrated expression of the shift toward Ethernet in scale-out networking. Nvidia (NVDA) still owns the integrated, proprietary stack and benefits as long as buying its systems means buying its network. The connectivity specialists, Astera Labs (ALAB) and Credo (CRDO), offer higher-beta exposure to cluster growth, while Coherent (COHR), Lumentum (LITE), and Fabrinet (FN) are exposed to the co-packaged-optics transition. These are different security profiles despite sharing the same demand driver; valuation and customer concentration determine whether strategic growth reaches shareholders.

GPU waiting time is the final test

The bullish case for the open-networking names weakens if Nvidia's proprietary stack proves stickier than expected and NVLink and InfiniBand hold their ground at the frontier, keeping the value inside Nvidia's system rather than letting it flow to merchant switching and Ethernet. Broadcom's position is the most robust in the layer, so the main risk to it is a broad AI-capex slowdown rather than a competitive loss. And the co-packaged-optics thesis slips if the reliability of on-package optics disappoints and the industry stays on pluggable modules longer than the 2026 timeline implies, delaying the opportunity for the optical-component suppliers.

For the platform-specific translation from switch and optical content into complete U.S. and Chinese systems, see the four supply-chain teardowns in §2.13.

The investment conclusion should be built from the workload outward. First ask whether clusters are becoming large enough, and communication-heavy enough, for networking content to rise. Next ask whether the relevant bottleneck sits inside the rack, across the data hall, or in the electrical-to-optical transition. Then identify the supplier that controls the design rather than merely assembling a replaceable part. Finally, test customer concentration and valuation against the expected duration of that socket.

The layer is attractive because networking can make already-purchased accelerators more productive. It is dangerous because architectures change quickly and a few customers control enormous programs. The cleanest evidence is not a market-size forecast. It is falling communication overhead in deployed clusters, sustained content per accelerator, stable gross margin, and repeated design wins across more than one customer generation. When those observations hold together, networking growth is becoming shareholder economics rather than simply another line in the AI bill of materials.


Sources

Linked evidence for this chapter's figures and load-bearing claims: 1

Footnotes

  1. Data Center AI Networking to Surge to Over $25B in 2028. 650 Group, 2024-01-24; accessed 2026-07-25. 2