Distributed Linear Solvers and Data Heterogeneity
We consider the problem of solving a large-scale system of linear equations in a distributed/federated setting. The taskmaster solves the system with the help of a set of machines, each of which possesses a subset of the equations. While various solutions for this problem exist, a fundamental understanding and rigorous comparison between the convergence rates of the main algorithmic classes - the projection-based methods and the optimization-based ones - is missing. We provide the first comprehensive analysis and comparison of these two classes of algorithms, with a particular focus on the fastest representative method from each class, i.e., the Accelerated Projection-Based Consensus (APC) and the Distributed Heavy-Ball Method. We introduce a novel notion of data heterogeneity called angular heterogeneity, discussing its significance. Using this notion, we characterize and compare the optimal convergence rates of the algorithms of interest and capture the effects of the number of machines, the number of equations, and cross-machine and local data heterogeneity on these rates. Our analysis sheds light on the previously observed superior performance of APC in realistic scenarios, where there is often large data heterogeneity, and provides several insights into the effect of angular heterogeneity on the efficiencies of different algorithms. Additionally, we provide distributed algorithms for efficiently computing the angular heterogeneity metrics. Lastly, as a by-product of this investigation, we obtain a tight bound on the condition number of an arbitrary matrix with full column rank in terms of the Euclidean norms of its columns and the angles between them. Numerical analyses validate our theoretical results, supporting the predicted advantage of APC in the high-heterogeneity regime and providing a deeper understanding of the effects of angular heterogeneity on convergence rates.