arXiv · 2610.07644
From ASIC to Fleet: Lessons from Building and Operating a Hyperscaler NIC
Abstract
We describe the operational infrastructure built to deploy and operate fbnic, a custom multi-host NIC, across hundreds of thousands of production hosts at Meta. Vendor multi-host NICs, designed by retrofitting single-host architectures, suffered from shared firmware and buffers that created cascading isolation failures over seven years. fbnic eliminates these through physical isolation, but shifting to in-house hardware shifts the entire operational burden to the hyperscaler. We present a hardware-in-the-loop CI pipeline testing firmware, driver, and kernel cross-products; a unified observability pipeline co-locating NIC and switch counters for cross-layer fault attribution; a driver-first architecture with fewer than ten firmware message types; a targeted firmware upgrade orchestrator at sub-sled granularity; and scoped repair automation confining blast radius to individual host slices. Over ten months, fbnic achieved a 12X reduction in unplanned unavailability, 37% lower mean time to repair, and 2.3X fewer hardware swaps compared to vendor NICs on the same platform.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Prankur Gupta, Alexander Duyck, Jakub Kicinski, Joseph Provine, Neal Peacock, Prabhakaran Ganesan, Rajiv Krishnamurthy, Chen Liu, Akshay Viswakumar, Timothy Vitkin, Jie Meng, Beatriz Padilla Hernandez, Michael Edwards, Andrei Kozlov, Viren Nathan, Tianyi Cui, Joy Chaoyue Xiong, Raul Hormazabal, Mohsin Bashir, Fred Feng, Nathan Walker, Lavin Khandelwal, Matt Maia, Milo Piazza, Mohanraj Thillainayagam, Mahmoud Mehr, Amithash Prasad, Lee Trager. 2026-10-06. From ASIC to Fleet: Lessons from Building and Operating a Hyperscaler NIC. https://arxiv.org/abs/2610.07644
Cite the original work for its findings. Save a collection to share your selection of sources.