Search arXiv⌕ Search

arXiv subjects

Sifat Munim

Publications and source records attributed to Sifat Munim.

2 recordsLinked to original sources

Communication-Efficient Distributed Training via Ring-Based Coded Approximate All-Reduce

Ring All-Reduce is widely used within large-scale distributed training for exact computation of the aggregate gradient. For a system with $N$ workers, its normalized per-worker communication is $2(N-1)/N$. In this work we present a communication-efficient Ring All-Reduce (CERAR) protocol for ``approximate'' gradient aggregation. CERAR partitions each local gradient into $c$ components and crucially relies on linear encoding and decoding operations. It performs $L=c+N-2$ communication rounds over an $N$-worker ring, yielding normalized communication rate $1+(N-2)/c$ and normalized storage $1+N/c$. We present an explicit Vandermonde-based construction, whose approximation error can be made arbitrarily close to zero with communication and storage rates approaching one with increasing $c$. However, this limit is achieved through ill-conditioned encoding matrices. Accordingly, we give a multiplicative perturbation construction whose error is $O(ε)$, while the relevant condition numbers are $O(ε^{-r_\star})$, where $r_\star=\lceil c/N\rceil-1$. This naturally motivates a condition-number-constrained optimization formulation for trading off the competing objectives and obtaining numerically stable practical designs. We present experiments on multi-GPU clusters with low-bandwidth and high-bandwidth interconnects. Our results demonstrate clear benefits in the low-bandwidth setting, even for moderate parameter length. We expect corresponding improvements even in the high-bandwidth setting for experiments with much higher parameter lengths.

cs.IT↗

Communication-Efficient Approximate Gradient Coding

Large-scale distributed learning aims at minimizing a loss function $L$ that depends on a training dataset with respect to a $d$-length parameter vector. The distributed cluster typically consists of a parameter server (PS) and multiple workers. Gradient coding is a technique that makes the learning process resilient to straggling workers. It introduces redundancy within the assignment of data points to the workers and uses coding theoretic ideas so that the PS can recover $\nabla L$ exactly or approximately, even in the presence of stragglers. Communication-efficient gradient coding allows the workers to communicate vectors of length smaller than $d$ to the PS, thus reducing the communication time. While there have been schemes that address the exact recovery of $\nabla L$ within communication-efficient gradient coding, to the best of our knowledge the approximate variant has not been considered in a systematic manner. In this work we present constructions of communication-efficient approximate gradient coding schemes. Our schemes use structured matrices that arise from bipartite graphs, combinatorial designs and strongly regular graphs, along with randomization and algebraic constraints. We derive analytical upper bounds on the approximation error of our schemes that are tight in certain cases. Moreover, we derive a corresponding worst-case lower bound on the approximation error of any scheme. For a large class of our methods, under reasonable probabilistic worker failure models, we show that the expected value of the computed gradient equals the true gradient. This in turn allows us to prove that the learning algorithm converges to a stationary point over the iterations. Numerical experiments corroborate our theoretical findings.

cs.IT↗