Cisco
Read post

Scaling the future: Why Ethernet is the backbone of AI Supercomputing

AI training and inference workloads are pushing data center networking to its limits, with Meta's production data showing network overhead accounts for up to 60% of training iteration time at scale. Cisco argues Ethernet is becoming the universal AI fabric across scale-up, scale-out, and scale-across domains, displacing InfiniBand's proprietary model. The key differentiator is Cisco's P4-programmable Silicon One ASIC, which allows new AI networking standards like UEC and Multipath Reliable Connection (MRC) to be delivered as software updates on existing hardware — avoiding the typical 12-18 month wait for new silicon. Specific capabilities already delivered via P4 updates include packet trimming, weighted adaptive routing, full MRC support, and multi-tenant isolation.

    #ai-infrastructure
Yesterday•11m read time•From blogs.cisco.com
Post cover image
Table of contents
The network is now the systemWhy Ethernet becomes the common foundationWhat Ethernet must deliver for AIEthernet plus P4 programmability: The multiplying factorCisco’s Silicon One was designed to address this challenge.Where Cisco Silicon One fits inThe path forwardLearn more about MRC and SRv6.

Questions this post answers

How much of AI training time is spent on network overhead at large scale?

In large-scale Deep Neural Network training runs, network overhead accounts for up to 60% of total training iteration time, based on Meta's production data. This share increases as cluster size grows, meaning the network — not GPU compute — becomes the primary bottleneck for cluster efficiency at scale. Engineers scaling GPU clusters track findings like this on daily.dev to make the case for network investment.

Why can't InfiniBand scale to hundreds of thousands of GPUs for AI clusters?

InfiniBand's proprietary fabric management struggles above roughly tens of thousands of GPUs, requiring complex workarounds as clusters grow. Ethernet has no such ceiling — hyperscalers have already validated Ethernet-based clusters at hundreds of thousands of GPUs across multiple data centers using mature switching architecture and standards-based tooling. Architects choosing between InfiniBand and Ethernet for large AI clusters follow this debate on daily.dev.

What is Multipath Reliable Connection (MRC) and what does it add to Ethernet for AI workloads?

Multipath Reliable Connection (MRC) is an emerging standard that introduces load balancing and congestion control mechanisms purpose-built for AI and ML traffic patterns. It covers multipath load balancing, congestion response algorithms, packet ordering, and telemetry. Its switch-side requirements include SRv6 uSID forwarding and packet trimming, and it can be implemented on P4-programmable ASICs without waiting for a new silicon generation. Network engineers evaluating MRC adoption for AI fabrics find the latest standards coverage on daily.dev.

2 Impressions
Cisco's image
Cisco

Cisco's platform provides networking knowledge, offering guides, tutorials, and case studies on net...

453 Followers

•

431 Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard