Search arXiv⌕ Search

arXiv · 2609.31388

Adaptive Switching Between Leader-Based and Leaderless BFT Protocols

Abstract

Byzantine fault-tolerant (BFT) protocols are known for providing operational consistency and resilience in distributed systems. However, evolving network conditions, often driven by the network's inherent dynamism or adversarial influence, make it suboptimal to rely on a static protocol at all times. Existing BFT protocol adaptation solutions switch only among leader-based protocols and coordinate each switch through a separate consensus round, leaving them ineffective at handling severe asynchrony or situations in which an adaptive adversary targets the network's leader. We propose BFTide, a protocol adaptation architecture that enables a BFT system to intelligently and swiftly switch to a suitable protocol as network conditions shift. BFTide integrates a novel protocol switching layer that embeds protocol transition logic into the ongoing BFT operation, enabling safe and low-overhead transitions between partially synchronous leader-based protocols and asynchronous leaderless protocols. It further incorporates an offline-trained reinforcement learning policy that allows nodes to propose protocols at runtime based on observed system metrics. Experimental results show that BFTide reduces transaction latency under adverse network conditions compared with static BFT protocols and the state-of-the-art BFT protocol adaptation scheme BFTBrain (NSDI'25), while maintaining comparable throughput. The switching layer adds a modest 10-21% overhead to median latency when idle and requires no separate consensus round per switch.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sudip Bhujel, Yue Li, Ning Zhang, Y. Thomas Hou, Wenjing Lou, Yang Xiao. 2026-09-25. Adaptive Switching Between Leader-Based and Leaderless BFT Protocols. https://arxiv.org/abs/2609.31388

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple

General Matrix Multiplication (GEMM) is the cornerstone of HPC workloads and Deep Learning. State-of-the-art (SOTA) vendor libraries tune tensor layouts, parallelization schemes and cache blocking to minimize data movement across the memory hierarchy and maximize throughput. However, optimal settings for these parameters depend on the target platform and matrix shapes, making exhaustive tuning infeasible. In this work, we address this cumbersome scheduling search and tuning using space-filling curves (SFC). We partition the matrix multiplication using advancements in SFC, and obtain platform-oblivious and shape-oblivious matrix multiplication schemes with a high degree of data locality. We extend the SFC-based work partitioning to implement Communication-Avoiding (CA) algorithms with replication techniques in a seamless fashion. The resulting SFC-CA GEMM achieves provable asymptotic communication optimality for both square and rectangular matrix regimes. Across four x86 and Arm platforms, SFC-CA GEMM outperforms vendor libraries by up to 5.5$\times$ per shape and 1.8$\times$ in weighted harmonic mean (WHM) throughput. Last, we show the impact of our work on two real-world applications by leveraging our SFC-CA GEMM as a compute backend: i) prefill of LLM inference with speedups up to 1.85$\times$ over SOTA inference runtimes, and ii) distributed-memory matrix multiplication with speedups up to 2.3$\times$ over the SOTA distributed-memory GEMM framework with vendor-optimized compute backend.

cs.DC↗

MoEless: Efficient MoE LLM Serving with Serverless Experts

Large Language Models (LLMs) increasingly adopt Mixture-of-Experts (MoE) architectures to scale efficiently under stringent resource constraints. However, MoE's sparse activation causes severe expert load imbalance, where a few experts become stragglers while others remain underutilized, leading to inflated inference latency and cost. Existing solutions assume static, serverful model deployments, limiting expert elasticity and often incurring costly expert swapping or degraded output quality. We present MoEless, an efficient serverless MoE serving framework that mitigates expert load imbalance via elastic expert execution. MoEless leverages lightweight, layer-aware predictors to estimate incoming expert load distributions and proactively identify stragglers. We design optimized scaling and placement strategies to improve function locality, GPU utilization, and cross-expert load balance. MoEless is prototyped on top of Megatron-LM and deployed on an eight-GPU testbed. Experiments with open-source MoE models and real-world workloads show that MoEless reduces inference latency by 43% and inference cost by 84% compared to state-of-the-art solutions.

cs.DC↗

Cascadia: A Control-Plane-Free Alternative to Hyperconverged AI Infrastructure

We present Cascadia, a system for serving large language models on fleets of commodity Intel AIPCs using their CPU, integrated-GPU, and NPU resources. Every node embeds ingress, scheduling, and execution; inference requests require no dedicated routing control plane. Nodes join a libp2p QUIC mesh using CA-issued ed25519 admission certificates, gossip signed capabilities, exchange live load over direct peer streams, and route OpenAI-compatible requests to eligible peers. An operator-run certificate authority handles admission and fleet management outside the inference path. Three serving modes share one interface: whole-model execution on one node, load-balanced replicas, and pipeline-sharded chains using the compilation and speculative decoding mechanism of our companion paper. Optional KV-cache mobility reuses compatible conversation prefixes after a routing move, with cold recomputation on a miss. Signed response receipts and hash-chained logs support provenance and audit. A three-node Phi-3.5-mini NPU testbed delivered 3.10x the response throughput of its one-node configuration under ten concurrent requests; a separate four-node deployment recorded 4.06x the throughput of direct single-node serving. Paired latency observations, runtime measurements, and internal functional checks characterize the tested configurations. We compare Cascadia with IBM, Nutanix, VMware, and HPE platforms on deployment footprint, hardware requirements, scheduling, scaling, licensing, and trust, using vendor documentation. The paper repository provides benchmark scripts, curated measurements, and a claim-to-evidence map.

cs.DC↗