Search arXivSearch

arXiv · 2609.22087

When Is Availability-Aware Training Worth It? A Benchmark and Empirical Study of Interruption-Resilient Optimization Under Predictable Compute Schedules

Abstract

Training under non-stationary but predictable compute availability (satellites under eclipse, duty-cycled edge devices, power-capped datacenters) is often framed as needing specialized, availability-aware optimizers. We test that premise. We release OrbitTrace, a benchmark of 50 physics-grounded availability traces from SGP4 propagation of live two-line element sets across three orbital regimes, and ask a falsifiable question: when an availability gap interrupts training, is a specialized resumption strategy worth it, or is competent checkpoint-and-resume enough? Our central finding: when optimizer state can be preserved across a gap, the gap is essentially free. A strong checkpoint baseline that restores full optimizer state and indexes its learning-rate schedule in effective (active) time matches uninterrupted training to within data-ordering noise on CIFAR-10/ResNet-18, and exactly on a GPT-2/AdamW task. Advantages previously reported for availability-aware methods, including our own three-pillar method AAT, arise almost entirely from comparison against a weak baseline that indexes its schedule on wall-clock time. Against the strong baseline, reactive adaptation provides no advantage across stationary, optimizer-state-loss, and distribution-drift regimes. We isolate one narrow regime where it helps: for large models whose optimizer state cannot be persisted across gaps and that are interrupted by frequent, short pauses, reconstructing a decayed optimizer moment recovers only ~21% of the state-loss penalty on average (and not robustly across seeds); this vanishes for eclipse-scale gaps, where the decayed moment is indistinguishable from zero. Our contributions are a benchmark, a strong reproducible baseline protocol, and a clear characterization of when interruption-resilient optimization is worth its complexity, and when it is not.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Subhadip Mitra. 2026-06-22. When Is Availability-Aware Training Worth It? A Benchmark and Empirical Study of Interruption-Resilient Optimization Under Predictable Compute Schedules. https://arxiv.org/abs/2609.22087

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Circulant ADMM-Net for Fast High-resolution DoA Estimation

This paper introduces CADMM-Net and CHADMM-Net, two deep neural networks for direction of arrival estimation within the least-absolute shrinkage and selection operator (LASSO) framework. These two networks are based on a structured deep unfolding of the alternating direction method of multipliers (ADMM) algorithm through the use of circulant as well as Hermitian-circulant matrices. Along with a computational complexity of $\mathcal{O}(N\log(N))$ per layer for the inference, where $N$ is the length of the dictionary $\mathbf{A}$, they additionally exhibit a memory footprint of $N$ and approximately half of $N$ for CADMMNet and CHADMM-Net, respectively, compared with $N^{2}$ for ADMM-Net. Furthermore, these structured networks exhibit a competitive performance against ADMM-Net, LISTA, TLISTA, and THLISTA with respect to the detection rate, the angular root-mean square error, and the normalized mean squared error.

eess.SP

Adaptive Probabilistic Constellation Shaping based on Enumerative Sphere Shaping for FSO Channel with Turbulence and Pointing Errors

Free-space optical (FSO) transmission enables fast, secure, and efficient next-generation communications with abundant spectrum resources. However, atmospheric turbulence, pointing errors, path loss, and atmospheric loss induce random attenuation, challenging link reliability. Adaptive coded modulation technology enhances spectrum utilization and reliability. We propose an adaptive probabilistic constellation shaping (A-PCS) coherent system utilizing enumerative sphere shaping (ESS) for distribution matcher (DM). With PCS-64QAM, the system achieves continuous rate control from conventional QPSK-equivalent to 64QAM spectral efficiency, providing quasi-continuous control with granularities of approximately $0.05$~bits/4D for spectral efficiency and $0.1$~dB for the post-FEC SNR threshold, and a maximum control depth of $12.5$~dB. Leveraging ESS for efficient sequence utilization, it offers higher spectral efficiency and finer control granularity than constant composition distribution matcher (CCDM)-based A-PCS systems. We further model and analyze the FSO channel, presenting calculations and comparisons of outage probability and ergodic capacity under varying turbulence intensities and pointing errors. Results demonstrate 99.999~\% reliability at maximum $σ_\mathrm{R}^2 = 1.02$ and $σ_\mathrm{s} = 0.51~\mathrm{m}$, meeting requirements under severe turbulence and large pointing errors. {Furthermore, under non-ideal delayed channel state information (CSI) feedback conditions, the system adapts to varying turbulence coherence times and feedback delays, with results showing that finer modulation granularity (provided by A-PCS-ESS) enhances immunity to feedback delay, maintaining a performance advantage over conventional adaptive schemes across a range of channel environments and delay values.

eess.SP

Rate-Splitting--Inspired Bistatic OFDM-ISAC

Achieving effective uplink bistatic ISAC over an OFDM waveform gives rise to challenging interference structures. These are mostly due to unequal direct- and echo-path contributions and Doppler-induced ICI, rendering orthogonal resource separation and fixed SIC strategies inadequate. To address this problem, we propose a RS-inspired framework where the transmitter splits each communication message into a robust and a supplementary stream, which are jointly superposed over a sensing signal. Furthermore, we present the design of a staged sensing-communication receiver. Based on this framework, we derive tractable per-subcarrier SINR expressions and establish the relation between sensing accuracy and communication reliability based on the Fisher information. Building on these, we formulate a joint power-allocation problem for SE maximization under sensing-performance and power constraints. The resulting non-convex formulation is solved using convex surrogates and fractional programming. Numerical results demonstrate that, compared to NOMA-inspired baselines, the proposed framework provides more effective IFI management and improved robustness to Doppler-induced ICI.

eess.SP