Low-Latency Coded Tensor--Matrix Multiplication for Distributed Signal Processing Systems
Large-scale signal processing systems, including aeronautical and aerospace platforms, increasingly rely on tensor operations, whose distributed execution is constrained by communication, memory, and latency bottlenecks. In this paper, we propose a tensor-aware coded computation framework for distributed mode-1 tensor--matrix multiplication that operates directly on tensor subtensors, avoiding explicit unfolding and preserving multi-dimensional structure. For a fixed partitioning configuration, the proposed tensor-PolyDot scheme achieves the same recovery threshold, communication cost, worker-side computation, and memory requirements as conventional matrix-based PolyDot schemes. However, by leveraging multivariate encoding, it replaces high-degree univariate interpolation at the fusion node with structured low-dimensional decoding. This leads to a substantial reduction in decoding complexity and latency, without affecting any other system metric. Additionally, the tensor-based approach improves memory locality and enables parallel decoding across tensor modes, making it well-suited to modern hardware architectures. Numerical results show decoding speedups proportional to the tensor partitioning factor, reaching up to an order-of-magnitude improvement in practical settings. These advantages make the proposed framework particularly relevant to latency-critical aeronautical and aerospace distributed signal processing systems.