Search arXiv⌕ Search

arXiv · 2609.34162

SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs

Abstract

Host-resident layer streaming enables full-parameter LLM fine-tuning beyond GPU memory, but data-parallel ranks compete for shared host resources. Replicated transfers amplify traffic, while strong scaling can expose host work as computation windows shrink. We present SlideDP, a synchronous data-parallel runtime for shared-host multi-GPU systems. It maintains one authoritative host state, decouples communication routes from state layout, and pipelines parameter delivery, gradient aggregation, and CPU updates across ranks and chunks. An analytical step-time model characterizes resource bottlenecks and pipeline exposure; runtime measurements guide communication, chunking, and activation policies under a GPU memory budget. In matched-batch sweeps, SlideDP achieves geometric-mean throughput ratios of 1.46-2.64$\times$ over SlideFormer, MegaTrain, and ZeRO-Offload. On four H100s, SlideDP approaches GPU-resident FSDP2 throughput for Qwen3-14B at a smaller batch size. With a larger batch, it processes over 1M tokens per step and exceeds FSDP2's measured peak throughput by 11.2%. Separately, it supports 256K-token sequences for the same model and fine-tunes Qwen2.5-72B on four RTX 4090 GPUs. Project page: https://github.com/RegiaYoung/SlideDP.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ruijia Yang, Shiyuan Lin, Yulong Ao, Zhiyu Li, Yingli Zhao, Xianduo Li, Yonghua Lin, Zeyi Wen. 2026-09-28. SlideDP: Scaling Host-Resident LLM Fine-Tuning Across Multiple GPUs. https://arxiv.org/abs/2609.34162

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Recolorable Graph Exploration by an Oblivious Agent with Fewer Colors

Recently, Böckenhauer, Frei, Unger, and Wehner (SIROCCO 2023) introduced a novel variant of the graph exploration problem in which a single memoryless agent must visit all nodes of an unknown, undirected, and connected graph before returning to its starting node. Unlike the standard model for mobile agents, edges are not labeled with port numbers. Instead, the agent can color its current node and observe the color of each neighboring node. To move, it specifies a target color and then moves to an adversarially chosen neighbor of that color. Böckenhauer~et al.~analyzed the minimum number of colors required for successful exploration and proposed an elegant algorithm that enables the agent to explore an arbitrary graph using only eight colors. In this paper, we present a novel graph exploration algorithm that requires only six colors. Furthermore, we prove that five colors are sufficient if we consider only a restricted class of graphs, which we call the $φ$-free graphs, a class that includes every graph with maximum degree at most three and every cactus.

cs.DC↗

Hierarchical Secure Distributed Linearly Separable Computation with Arbitrary Heterogeneous Data Assignment

This paper studies secure distributed linearly separable computation over a three-layer hierarchical network, where clustered users communicate with a central server through relays. The server aims to recover Kc linear combinations of K intermediate outcomes, where each intermediate outcome is a separable function of one dataset. We consider a more general setting with arbitrary heterogeneous data assignment across users, where ''arbitrary'' means that the data assignment is given in advance (which can be in any form) and ''heterogeneous'' means that the users may hold different numbers of datasets. Under this assignment, each user computes the intermediate outcomes of its assigned datasets and sends masked messages to its associated relay. The relays subsequently process and forward the received messages to the server. We impose two security constraints: (i) security against server, requiring the server to learn only the desired task function without gaining any additional information about users' inputs; and (ii) security against relays, ensuring each relay learns nothing about users' inputs. Moreover, the server or any relay may collude with a subset of users. For Kc=1, the underlying computation reduces to distributed gradient coding. We propose a secure scheme tolerating user dropouts and user collusion, achieving the optimal two-layer communication rates in one regime and order-optimal communication rates within a factor of 2 in the other regime. For Kc>1, we extend the proposed construction to multi-dimensional linearly separable tasks under the no-dropout setting.

cs.DC↗

ParaAnya: Accelerating Parallel Diffusion Sampling with Plug-and-Play Output Caching

Diffusion models have achieved remarkable success in generative tasks, but their inherently sequential sampling process introduces a severe computational bottleneck. Recent Parallel-in-Time (PinT) solvers attempt to mitigate this by parallelizing generation across a sliding window of timesteps, advancing the window only when step-wise changes stabilize. However, this overlapping window mechanism forces the network to repeatedly evaluate the same timesteps. When the input variations between iterations are minimal, these redundant evaluations lead to significant computational waste. To address this inefficiency, we propose ParaAnya, an output cache mechanism agnostic to the parallel sampling algorithm that can reduce the number of function evaluations (NFE). ParaAnya caches input-output pairs of diffusion models and reuses the cached output at overlapping timesteps. By dispatching only cache-miss timesteps to GPU workers, our approach eliminates redundant computation while preserving the structure of the underlying algorithms' update rules. We integrate ParaAnya into four representative parallel sampling algorithms and evaluate its performance on Stable Diffusion v1.5. Across four parallel samplers evaluated with DDIM on eight GPUs, ParaAnya provides $1.30$--$2.43\times$ speedups over their uncached counterparts and reduces NFE by up to 70.1\%, reaching up to a $5.62\times$ speedup over single-GPU serial sampling while maintaining comparable CLIP scores.

cs.DC↗