Search arXiv⌕ Search

arXiv · 2609.32197

SparSP: Exploiting Communication Sparsity for Sequence-Parallel Video DiTs

Abstract

Diffusion Transformers have become the dominant architecture for video generation. Their substantial computational cost motivates scaling inference across multi-GPU servers, yet efficient scaling remains challenging on commodity GPUs connected via PCIe, whose bandwidth is limited. Although sparse attention substantially reduces computation, its implications for communication remain underexplored. This paper argues that attention sparsity should be treated as a communication primitive. We present SparSP, an efficient sparse sequence parallel communication system that co-designs token distribution, communication routing, and asynchronous execution for sparse video diffusion models. First, Dependency-Aware Placement distributes sequence blocks according to diffusion models' sparse attention patterns. Second, Demand-Directed KV Routing transfers KV blocks directly to requesting GPUs without intermediate relays. Third, a Decoupled Transfer Runtime separates communication from GPU computation to reduce resource contention and maximize effective bandwidth. Our evaluation shows that SparSP improves attention performance by 1.38- 1.5$\times$, achieves an average 1.17$\times$ (up to 1.69$\times$) end-to-end speedup across three representative servers and three video diffusion models, and reduces communication volume by 12.54-23.05%. Moreover, we achieve an average 1.53-1.76$\times$ bandwidth improvement over NCCL.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Desen Sun, Xinrui Zhong, Yuke Wang, Sihang Liu. 2026-09-26. SparSP: Exploiting Communication Sparsity for Sequence-Parallel Video DiTs. https://arxiv.org/abs/2609.32197

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Recolorable Graph Exploration by an Oblivious Agent with Fewer Colors

Recently, Böckenhauer, Frei, Unger, and Wehner (SIROCCO 2023) introduced a novel variant of the graph exploration problem in which a single memoryless agent must visit all nodes of an unknown, undirected, and connected graph before returning to its starting node. Unlike the standard model for mobile agents, edges are not labeled with port numbers. Instead, the agent can color its current node and observe the color of each neighboring node. To move, it specifies a target color and then moves to an adversarially chosen neighbor of that color. Böckenhauer~et al.~analyzed the minimum number of colors required for successful exploration and proposed an elegant algorithm that enables the agent to explore an arbitrary graph using only eight colors. In this paper, we present a novel graph exploration algorithm that requires only six colors. Furthermore, we prove that five colors are sufficient if we consider only a restricted class of graphs, which we call the $φ$-free graphs, a class that includes every graph with maximum degree at most three and every cactus.

cs.DC↗

Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines

Agentic workflows carry out complex tasks by orchestrating multiple large language models (LLMs) and tools. Serving them at a target throughput with low latency is hard because they are written in arbitrary agentic frameworks and their execution times are unpredictable: execution branches, fans out, or recurs in data-dependent ways. Since their LLMs often outnumber the available GPUs, they also oversubscribe GPUs. We describe Scepsy, a serving system that schedules arbitrary multi-LLM agentic workflows onto a GPU cluster. Scepsy exploits the insight that, while the end-to-end latency of an agentic workflow is unpredictable, each LLM's fraction of execution time is comparatively stable across requests. Scepsy profiles each LLM under different parallelism degrees and combines the profiles with these fractions into an Aggregate LLM Pipeline, a lightweight throughput and latency predictor for allocations. To minimize latency at a target throughput, Scepsy uses the Aggregate LLM Pipeline to search over fractional GPU shares, tensor parallelism degrees, and replica counts. A hierarchical heuristic then places the chosen allocation onto the cluster, minimizing fragmentation and respecting network topology. On realistic agentic workflows, Scepsy achieves up to 2.5x higher throughput before saturation and 1.0-3.3x lower latency than systems that optimize LLMs independently or rely on user-specified allocations.

cs.DC↗

The World's Fastest Matching Engine Algorithm

We drove 247 matching engines through one C ABI harness on one identical workload: every open-source FIFO implementation we could find, deduplicated, and our own, on the same gate. The workload doubles as a byte-identical correctness oracle, 1,000,000,000+ order messages per engine, replayed against an independent-engine consensus. Only 47 are correct as shipped; we filed 181 GitHub issues upstream, 28 already fixed by their maintainers, none declined. Our engine leads the 160 that reproduce the consensus by ~95 M/s (12.6x the second best) on worst-case throughput. One core sustains 103.4 million order messages per second (122.09 million on AMD EPYC processors) worst-case, and in the engine's production configuration its wire pipeline keeps single-book P99 host-path latency, OUCH parsing and OUCH/ITCH encoding included, under a microsecond through 82 M msgs/s; a 96-core server (~$1,630/month, 3-year reserved) reaches ~1.3 billion/s across over 10,000 symbols, for scale, over 45x the CTA consolidated quote feed's provisioned capacity. The lead is structural: the 73 engines written inside the trading industry sit under the same 8.19 M/s ceiling as the rest of the field, and it is one of them that sets it.

cs.DC↗