- Broadcom’s switches play a key role in connecting data centers together
- The vendor said hyperscalers often want to own and control these important connections
- Inside AI data centers, AI NICs are key to GPU coordination
Telcos worldwide are eager to help hyperscalers and neoclouds create fast fiber connections between data centers, and a handful of deals are in place already. Telcos want to bring together fiber, optical line systems, switches and services needed to connect massive GPU clusters in disparate data centers. Broadcom, which makes the chips inside most high-capacity switches, sees demand for this type of connectivity growing steadily.
“In the last two years, it's been picking up because the cluster sizes are becoming bigger,” said Broadcom VP of Products Hasan Siraj. “Some of these hyperscalers are doing [it] themselves … they are owning that connectivity. But if they need the service provider in between, they absolutely use the service provider.”
The connectivity between AI data centers is often called “scale across,” distinct from “scale up” (connectivity within an AI server rack) and “scale out” (connectivity between racks). Siraj said Broadcom has solutions for all three use cases.
For scale across architectures, hardware-level security is mandatory, Siraj said, noting that Broadcom’s Jericho switch chips are deep buffered with line rate encryption and use telemetry to expose anomalies or denial of service attacks.
The network is the computer
Siraj, who has responsibility for Broadcom’s routing, switching, high-performance NIC and open networking software product portfolios, said that in AI infrastructure, the network is the computer. “When you're training these models or doing distributed inference, you can't fit these models on a single XPU,” he told Fierce. “You need, depending on what you're doing, tens, hundreds, hundreds of thousands, maybe half a billion. So you need something to connect all of these together. And that is the network.”
Even with the rising cost of compute, Siraj thinks hyperscalers are unlikely to reallocate significant budget away from networking. “Service providers and hyperscalers want the network to be in place three to six months before compute shows up,” he said. “They will build out the network. You will be sitting in racks, and when the compute comes, connectivity happens.”
A 10-megawatt data center could house roughly 7,000 XPUs/GPUs/TPUs, Siraj said. If a use case demands more, data centers need to be connected together. “These distances may be tens of kilometers to maybe 100 kilometers,” he said.
Analyst Earl Lum of EJL Wireless Research agreed that data center interconnects are needed to train AI models, but said he is less certain distributed inference will be a major driver for scale across connectivity. “You do scale across to pull data centers together to train these models and leverage all the compute, so you need to mesh them all together,” he said. “But if you no longer need to do that kind of training. you may no longer need those fiber optic interconnects.” Lum believes the portion of AI workloads dedicated to inference versus training is on the rise.
Whatever happens with scale across connectivity, AI is expected to generate skyrocketing demand for networking inside data centers. Siraj said that since AI architectures require tight communication between switches and network interface cards (NICs), Broadcom developed an AI NIC.
In the NIC of time
Broadcom’s move was well-timed since GPUs running AI workloads use NICs to communicate directly with one another. Demand for these cards has exploded, and the supply of AI network interface cards has tightened. Nvidia’s ConnectX cards, which dominate the market, can be hard to come by for all but the biggest buyers. Siraj said Broadcom’s Thor 2 and NetXtreme solutions are less constrained. “For us, it's a new market that we are entering, and yes, overall, the demand is extremely high, but for the customers that we are working with, we can fulfill that,” he said.
Unlike Nvidia’s NICs, which use the AI-specific Infiniband connectivity architecture, Broadcom’s NICs use Ethernet. “I think there is now unanimous consensus in the industry, almost, that for most of the use cases, Ethernet should be the de facto standard,” said Siraj, echoing a prediction made by other vendors who are not Nvidia.
Read more about telcos and data centers on Fierce Network
Telcos shoulder big risk building fiber for AI data centers
Some hyperscalers block fiber off-ramps, keeping rural America disconnected
