Search arXiv⌕ Search

arXiv · 2009.03527

Approximate Multiplication of Sparse Matrices with Limited Space

Abstract

Approximate matrix multiplication with limited space has received ever-increasing attention due to the emergence of large-scale applications. Recently, based on a popular matrix sketching algorithm -- frequent directions, previous work has introduced co-occuring directions (COD) to reduce the approximation error for this problem. Although it enjoys the space complexity of $O((m_x+m_y)\ell)$ for two input matrices $X\in\mathbb{R}^{m_x\times n}$ and $Y\in\mathbb{R}^{m_y\times n}$ where $\ell$ is the sketch size, its time complexity is $O\left(n(m_x+m_y+\ell)\ell\right)$, which is still very high for large input matrices. In this paper, we propose to reduce the time complexity by exploiting the sparsity of the input matrices. The key idea is to employ an approximate singular value decomposition (SVD) method which can utilize the sparsity, to reduce the number of QR decompositions required by COD. In this way, we develop sparse co-occuring directions, which reduces the time complexity to $\widetilde{O}\left((\nnz(X)+\nnz(Y))\ell+n\ell^2\right)$ in expectation while keeps the same space complexity as $O((m_x+m_y)\ell)$, where $\nnz(X)$ denotes the number of non-zero entries in $X$ and the $\widetilde{O}$ notation hides constant factors as well as polylogarithmic factors. Theoretical analysis reveals that the approximation error of our algorithm is almost the same as that of COD. Furthermore, we empirically verify the efficiency and effectiveness of our algorithm.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yuanyu Wan, Lijun Zhang. 2024-06-23. Approximate Multiplication of Sparse Matrices with Limited Space. https://arxiv.org/abs/2009.03527

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Efficient Constrained Graph Search for Post-hoc Error Correction in Binary Classifiers

We introduce a model-agnostic framework for constrained post-hoc error correction in binary classifiers. Given a frozen base classifier, the method searches for an interpretable conjunction of feature--threshold rules that corrects residual false-positive or false-negative errors while explicitly constraining newly introduced errors. The approach combines graph-based search over candidate rule paths, depth-dependent dynamic constraints, and a reduced-histogram procedure for efficient threshold evaluation. Unlike retraining or modifying the base classifier, the learned correction path operates on its predictions and can therefore be applied to arbitrary binary classifiers with suitable input features. Experiments on a large binary-classification problem demonstrate that the method can identify compact correction rules efficiently; for example, one configuration removes 90\% of false positives while sacrificing 5\% of true positives.

cs.LG↗

Stochastic Bilevel Optimization with Heavy-Tailed Noise

This paper considers the smooth bilevel optimization in which the lower-level problem is strongly convex and the upper-level problem is possibly nonconvex. We focus on the stochastic setting where the algorithm can access the unbiased stochastic gradient evaluation with heavy-tailed noise, which is prevalent in many machine learning applications, such as training large language models and reinforcement learning. We propose a nested-loop normalized stochastic bilevel approximation (N$^2$SBA) for finding an $ε$-stationary point with the stochastic first-order oracle (SFO) complexity of $\tilde{\mathcal{O}}\big(κ^{\frac{7p-3}{p-1}} σ^{\frac{p}{p-1}} ε^{-\frac{4 p - 2}{p-1}}\big)$, where $κ$ is the condition number, $p\in(1,2]$ is the order of central moment for the noise, and $σ$ is the noise level. Furthermore, we specialize our idea to solve the nonconvex-strongly-concave minimax optimization problem, achieving an $ε$-stationary point with the SFO complexity of~$\tilde{\mathcal O}\big(κ^{\frac{2p-1}{p-1}} σ^{\frac{p}{p-1}} ε^{-\frac{3p-2}{p-1}}\big)$. All the above upper bounds match the best-known results under the special case of the bounded variance setting, i.e., $p=2$. We also conduct the numerical experiments to show the empirical superiority of the proposed methods.

cs.LG↗

FLAME: Flow Enhanced Legendre Memory Models for General Time Series Forecasting

In this work, we introduce FLAME, a family of extremely lightweight and capable Time Series Foundation Models, which support versatile forecasting tasks via generative probabilistic modeling, while ensuring both efficiency and robustness. FLAME utilizes the Legendre Memory for strong generalization capabilities. By adapting variants of Legendre Memory, i.e., translated Legendre (LegT) and scaled Legendre (LegS), in the Encoding and Decoding phases, FLAME can effectively capture the inherent inductive bias within data and make efficient long-range inferences. To enhance the accuracy of probabilistic forecasting while remaining efficient, FLAME adopts a normalizing-flow-based forecasting head, which can model complex distributions over the forecasting horizon in a generative manner. Comprehensive experiments on three well-recognized benchmarks, including TSFM-Bench, ProbTS, and TFB, demonstrate that FLAME is a strong out-of-the-box tool for decision intelligence.

cs.LG↗