Search arXiv⌕ Search

arXiv subjects

Yongqiang Wang

Publications and source records attributed to Yongqiang Wang.

At least 19 recordsLinked to original sources

Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share. We train only the student encoder to regress the teacher's per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher-student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces we support, and applies both to a tokenizer pretrained alone and to one jointly trained with a language model. At 2.8x compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher-student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.

cs.SD↗

Apple Intelligence Foundation Language Models

We present foundation language models developed to power Apple Intelligence features, including a ~3 billion parameter model designed to run efficiently on devices and a large server-based language model designed for Private Cloud Compute. These models are designed to perform a wide range of tasks efficiently, accurately, and responsibly. This report describes the model architecture, the data used to train the model, the training process, how the models are optimized for inference, and the evaluation results. We highlight our focus on Responsible AI and how the principles are applied throughout the model development.

cs.AI↗

Local Updates in Distributed Optimization: Provable Acceleration and Topology Effects

Inspired by the success of performing multiple local optimization steps between communication rounds in federated learning, incorporating such local updates into distributed optimization has recently attracted growing interest. However, unlike federated learning, where local updates can accelerate training by reducing gradient estimation error under minibatch settings, it remains unclear whether similar benefits persist when exact gradients are available. Moreover, existing theoretical results typically require reducing the step size when multiple local updates are employed, which can entirely offset any potential benefit of these additional local updates. In this paper, we focus on the classic DIGing algorithm and leverage the tight performance bounds provided by Performance Estimation Problems (PEP) to show that incorporating local updates can indeed accelerate distributed optimization. To the best of our knowledge, this is the first rigorous demonstration of such acceleration for a broad class of objective functions. Our analysis further reveals that, under an appropriate step size, performing only two local updates is sufficient to achieve the maximal possible improvement, and that additional local updates provide no further gains. Because more updates increase computational cost, these findings offer practical guidance for efficient implementation. We also show that these speed gains depend critically on the network structure, with sparser or less connected graphs, characterized by the spectral properties of the mixing matrix, yielding smaller improvements. Extensive experiments on both synthetic and real-world datasets corroborate the theoretical findings.

eess.SY↗

Accelerating Optimization and Machine Learning through Decentralization

Decentralized optimization enables multiple devices to learn a global machine learning model while each individual device only has access to its local dataset. By avoiding the need for training data to leave individual users' devices, it enhances privacy and scalability compared to conventional centralized learning, where all data has to be aggregated to a central server. However, decentralized optimization has traditionally been viewed as a necessary compromise, used only when centralized processing is impractical due to communication constraints or data privacy concerns. In this study, we show that decentralization can paradoxically accelerate convergence, outperforming centralized methods in the number of iterations needed to reach optimal solutions. Through examples in logistic regression and neural network training, we demonstrate that distributing data and computation across multiple agents can lead to faster learning than centralized approaches, even when each iteration is assumed to take the same amount of time, whether performed centrally on the full dataset or decentrally on local subsets. This finding challenges longstanding assumptions and reveals decentralization as a strategic advantage, offering new opportunities for more efficient optimization and machine learning.

cs.LG↗

Gradient Manipulation in Distributed Stochastic Gradient Descent with Strategic Agents: Truthful Incentives with Convergence Guarantees

Distributed learning has gained significant attention due to its advantages in scalability, privacy, and fault tolerance.In this paradigm, multiple agents collaboratively train a global model by exchanging parameters only with their neighbors. However, a key vulnerability of existing distributed learning approaches is their implicit assumption that all agents behave honestly during gradient updates. In real-world scenarios, this assumption often breaks down, as selfish or strategic agents may be incentivized to manipulate gradients for personal gain, ultimately compromising the final learning outcome. In this work, we propose a fully distributed payment mechanism that, for the first time, guarantees both truthful behaviors and accurate convergence in distributed stochastic gradient descent. This represents a significant advancement, as it overcomes two major limitations of existing truthfulness mechanisms for collaborative learning:(1) reliance on a centralized server for payment collection, and (2) sacrificing convergence accuracy to guarantee truthfulness. In addition to characterizing the convergence rate under general convex and strongly convex conditions, we also prove that our approach guarantees the cumulative gain that an agent can obtain through strategic behavior remains finite, even as the number of iterations approaches infinity--a property unattainable by most existing truthfulness mechanisms. Our experimental results on standard machine learning tasks, evaluated on benchmark datasets, confirm the effectiveness of the proposed approach.

cs.LG↗

Local Differential Privacy for Distributed Stochastic Aggregative Optimization with Guaranteed Optimality

Distributed aggregative optimization underpins many cooperative optimization and multi-agent control systems, where each agent's objective function depends both on its local optimization variable and an aggregate of all agents' optimization variables. Existing distributed aggregative optimization approaches typically require access to accurate gradients of the objective functions, which, however, are often hard to obtain in real-world applications. For example, in machine learning, gradients are commonly contaminated by two main sources of noise: the randomness inherent in sampled data, and the additional variability introduced by mini-batch computations. In addition to the issue of relying on accurate gradients, existing distributed aggregative optimization approaches require agents to share explicit information, which could breach the privacy of participating agents. We propose an algorithm that can solve both problems with existing distributed aggregative optimization approaches: not only can the proposed algorithm guarantee mean-square convergence to an exact optimal solution when the gradients are subject to noise, it also simultaneously ensures rigorous differential privacy, with the cumulative privacy budget guaranteed to be finite even when the number of iterations tends to infinity. To the best of our knowledge, this is the first algorithm able to guarantee both accurate convergence and rigorous differential privacy in distributed aggregative optimization. Besides characterizing the convergence rates under nonconvex/convex/strongly convex conditions, we also rigorously quantify the cost of differential privacy in terms of convergence rates. Experimental results on personalized machine learning using benchmark datasets confirm the efficacy of the proposed algorithm.

eess.SY↗

Quantum Nanophotonic Interface for Tin-Vacancy Centers in Thin-Film Diamond

The negatively charged tin-vacancy center in diamond (SnV$^-$) is an excellent solid state qubit with optically-addressable transitions and a long electron spin coherence time at elevated ($\sim1.7$ K). However, implementing scalable quantum nodes with high-fidelity optical readout of the electron spin state requires efficient photon emission and collection from the system. In this manuscript, we report a quantum photonic interface for SnV$^-$ centers based on one-dimensional photonic crystal cavities fabricated in diamond thin films. Furthermore, we provide a rigorous description of the spontaneous emission dynamics of our system, taking into account individual contributions from both the C and D transitions of the emitter. This allows for determination of Purcell factors per transition and, by extension, the C/D branching ratio SnV$^{-}$ zero phonon line. We observe quality factors up to $\sim$6000 across this sample, and measure up to a 12-fold lifetime reduction, which translates into a Purcell factor of $F_C=26.2\pm1.5$ for a targeted C transition. By considering the cavity mode polarization alignment with the C and D transition dipole moments, we validate the C/D branching ratio to be $η_{\text{BR}}=0.75\pm0.01$, in line with previous theoretical and experimental findings.

quant-ph↗

Irradiation-induced amplification of electric fields at oxide interfaces as revealed by correlative DPC-STEM and DFT

Heterointerfaces are ubiquitous in modern devices, found in technologies ranging from microelectronics to structural components for energy applications. Many of these emerging technologies are found in applications such as satellites, batteries, and next generation nuclear reactors, that are subject to harsh environments. In some scenarios, multiple extreme conditions, such as irradiation and corrosion, act on the material simultaneously. Extending the lifetime of these technologies is dependent on a detailed understanding of how their component materials platforms and interfaces respond in extreme environments, where irradiation and corrosion may couple in unique ways, distinct from corrosion under ambient conditions. Oxides, which form readily over metal underlayers, can act as protective coatings; enhancing the robustness of oxide overlayers to protect underlying metal alloys is a potential avenue towards corrosion mitigation. Here we study the impact of irradiation-induced non-equilibrium defects on charge segregation and electric fields at and near multi-phase oxide heterointerfaces. We perform a detailed study of irradiated Fe2O3-Cr2O3 thin film heterostructures using first-principles DFT electronic structure modeling paired with 4D-STEM DPC and EELS techniques to measure nanoscale changes in electric fields. Our results show clear evidence that irradiation drives substantial modulation of interfacial electric fields that can be tailored by controlling the atomistic chemical structure of the oxide interface. We show that irradiation can selectively induce built-in electric fields, thereby altering their direction; this suggests a pathway to engineering protective oxide heterostructure overlayers that can electrically control the spatial distribution of defects, with significant implications for the design of corrosion-resistant materials for extreme environments.

cond-mat.mtrl-sci↗

AXLearn: Modular, Hardware-Agnostic Large Model Training

AXLearn is a production system which facilitates scalable and high-performance training of large deep learning models. Compared to other state-of-art deep learning systems, AXLearn has a unique focus on modularity and support for hardware-agnostic training. AXLearn's internal interfaces between software components follow strict encapsulation, allowing different components to be assembled to facilitate rapid model development and experimentation on different hardware infrastructure. AXLearn maintains constant complexity as we scale the components in the system, compared to linear or quadratic complexity in state-of-the-art training systems. This allows integrating features such as Rotary Position Embeddings (RoPE) into AXLearn across hundred of modules with just 10 lines of code, compared to hundreds as required in other systems. At the same time, AXLearn maintains equivalent performance compared to state-of-the-art training systems. Finally, we share our experience in the development and operation of AXLearn at Apple.

cs.LG↗

Provably Convergent Decentralized Optimization over Directed Graphs under Generalized Smoothness

Decentralized optimization has become a fundamental tool for large-scale learning systems; however, most existing methods rely on the classical Lipschitz smoothness assumption, which is often violated in problems with rapidly varying gradients. Motivated by this limitation, we study decentralized optimization under the generalized $(L_0, L_1)$-smoothness framework, in which the Hessian norm is allowed to grow linearly with the gradient norm, thereby accommodating rapidly varying gradients beyond classical Lipschitz smoothness. We integrate gradient-tracking techniques with gradient clipping and carefully design the clipping threshold to ensure accurate convergence over directed communication graphs under generalized smoothness. In contrast to existing distributed optimization results under generalized smoothness that require a bounded gradient dissimilarity assumption, our results remain valid even when the gradient dissimilarity is unbounded, making the proposed framework more applicable to realistic heterogeneous data environments. We validate our approach via numerical experiments on standard benchmark datasets, including LIBSVM and CIFAR-10, using regularized logistic regression and convolutional neural networks, demonstrating superior stability and faster convergence over existing methods.

math.OC↗

Differentially-Private Distributed Model Predictive Control of Linear Discrete-Time Systems with Global Constraints

Distributed model predictive control (DMPC) has attracted extensive attention as it can explicitly handle system constraints and achieve optimal control in a decentralized manner. However, the deployment of DMPC strategies generally requires the sharing of sensitive data among subsystems, which may violate the privacy of participating systems. In this paper, we propose a differentially-private DMPC algorithm for linear discrete-time systems subject to coupled global constraints. Specifically, we first show that a conventional distributed dual gradient algorithm can be used to address the considered DMPC problem but cannot provide strong privacy preservation. Then, to protect privacy against the eavesdropper, we incorporate a differential-privacy noise injection mechanism into the DMPC framework and prove that the resulting distributed optimization algorithm can ensure both provable convergence to a global optimal solution and rigorous $ε$-differential privacy. In addition, an implementation strategy of the DMPC is designed such that the recursive feasibility and stability of the closed-loop system are guaranteed. Simulation results are provided to demonstrate the effectiveness of the developed approach.

eess.SY↗

Guaranteeing Both Consensus and Optimality in Decentralized Nonconvex Optimization with Multiple Local Updates

Scalable decentralized optimization in large-scale systems hinges on efficient communication. A common way to reduce communication overhead is to perform multiple local updates between two communication rounds, as in federated learning. However, extending this strategy to fully decentralized settings poses fundamental challenges. Existing decentralized algorithms with multiple local updates guarantee accurate convergence only under strong convexity, limiting applicability to the nonconvex problems prevalent in machine learning. Moreover, many methods require exchanging and storing auxiliary variables, such as gradient-tracking vectors or correction terms, to ensure convergence under data heterogeneity, incurring high communication and memory costs. In this paper, we propose MILE, a fully decentralized algorithm that guarantees both consensus and optimality under multiple local updates in general nonconvex settings. This is achieved through a novel periodic-system-based formulation and a lifting-based analysis, which together yield a closed-form expression for the state evolution across local updates, a theoretical advance not achieved previously. This closed-form characterization allows us to establish, for the first time, guaranteed consensus and optimality in decentralized nonconvex optimization under multiple local updates, in contrast to prior results that only ensure optimality of the average state. We prove that MILE achieves an $O(1/T)$ convergence rate under both exact and stochastic gradients, while requiring only a single variable exchange per interacting agent pair, minimizing communication and memory costs. Numerical experiments on benchmark datasets confirm its effectiveness.

math.OC↗

Disorder-broadened topological Hall phase and anomalous Hall scaling in FeGe

Magnetic skyrmions are topologically protected spin textures that are promising candidates for low-power spintronic memory and logic devices. Realizing skyrmion-based devices requires an understanding of how structural disorder affects their stability and transport properties. This study uses Ne$^{+}$ ion irradiation at fluences from $10^{11}$ to $10^{14}$ ions-cm$^{-2}$ to systematically vary defect densities in 80 nm epitaxial FeGe films and quantify the resulting modifications to magnetic phase boundaries and electronic scattering. Temperature- and field-dependent Hall measurements reveal that increasing disorder progressively extends the topological Hall signal from a narrow window near 200K in pristine films down to 4K at the highest fluence, with peak amplitude more than doubling. Simultaneously, the anomalous Hall effect transitions from quadratic Berry curvature scaling to linear skew scattering behavior, with the skew coefficient increasing threefold. These results establish quantitative correlations between defect concentration, skyrmion phase space, and transport mechanisms in a chiral magnet. It demonstrates that ion-beam modification provides systematic control over both topological texture stability and electrical detectability.

cond-mat.mtrl-sci↗

Critical Disconnect Between Structural and Electronic Recovery in Amorphous GaAs during Recrystallization

Understanding the evolution of structure and functionality through amorphous to crystalline phase transitions is critical for predicting and designing devices for application in extreme conditions. Here, we consider both aspects of recrystallization of irradiated GaAs. We find that structural evolution occurs in two stages, a low temperature regime characterized by slow, epitaxial front propagation and a high-temperature regime above dominated by rapid growth and formation of dense nanotwin networks. We link aspects of this structural evolution to local ordering, or paracrystallinity, within the amorphous phase. Critically, the electronic recovery of the materials is not commensurate with this structural evolution. The electronic properties of the recrystallized material deviate further from the pristine material than do those of the amorphous phase, highlighting the incongruence between structural and electronic recovery and the contrasting impact of loss of long range order versus localized defects on the functionality of semiconducting materials.

cond-mat.mtrl-sci↗

Data-Centric Lessons To Improve Speech-Language Pretraining

Spoken Question-Answering (SQA) is a core capability for useful and interactive artificial intelligence systems. Recently, several speech-language models (SpeechLMs) have been released with a specific focus on improving their SQA performance. However, a lack of controlled ablations of pretraining data processing and curation makes it challenging to understand what factors account for performance, despite substantial gains from similar studies in other data modalities. In this work, we address this gap by conducting a data-centric exploration for pretraining SpeechLMs. We focus on three research questions fundamental to speech-language pretraining data: (1) how to process raw web-crawled audio content for speech-text pretraining, (2) how to construct synthetic pretraining datasets to augment web-crawled data and (3) how to interleave (text, audio) segments into training sequences. We apply the insights from our controlled data-centric ablations to pretrain a 3.8B-parameter SpeechLM, called SpeLangy, that outperforms models that are up to 3x larger by 10.2% absolute performance. We hope our findings highlight the impact of effective data curation for speech-language pretraining and guide future data-centric exploration in SpeechLMs.

eess.AS↗

Skeleton-based Robust Registration Framework for Corrupted 3D Point Clouds

Point cloud registration is fundamental in 3D vision applications, including autonomous driving, robotics, and medical imaging, where precise alignment of multiple point clouds is essential for accurate environment reconstruction. However, real-world point clouds are often affected by sensor limitations, environmental noise, and preprocessing errors, making registration challenging due to density distortions, noise contamination, and geometric deformations. Existing registration methods rely on direct point matching or surface feature extraction, which are highly susceptible to these corruptions and lead to reduced alignment accuracy. To address these challenges, a skeleton-based robust registration framework is presented, which introduces a corruption-resilient skeletal representation to improve registration robustness and accuracy. The framework integrates skeletal structures into the registration process and combines the transformations obtained from both the corrupted point cloud alignment and its skeleton alignment to achieve optimal registration. In addition, a distribution distance loss function is designed to enforce the consistency between the source and target skeletons, which significantly improves the registration performance. This framework ensures that the alignment considers both the original local geometric features and the global stability of the skeleton structure, resulting in robust and accurate registration results. Experimental evaluations on diverse corrupted datasets demonstrate that SRRF consistently outperforms state-of-the-art registration methods across various corruption scenarios, including density distortions, noise contamination, and geometric deformations. The results confirm the robustness of SRRF in handling corrupted point clouds, making it a potential approach for 3D perception tasks in real-world scenarios.

cs.CV↗

Robust Partial 3D Point Cloud Registration via Confidence Estimation under Global Context

Partial point cloud registration is essential for autonomous perception and 3D scene understanding, yet it remains challenging owing to structural ambiguity, partial visibility, and noise. We address these issues by proposing Confidence Estimation under Global Context (CEGC), a unified, confidence-driven framework for robust partial 3D registration. CEGC enables accurate alignment in complex scenes by jointly modeling overlap confidence and correspondence reliability within a shared global context. Specifically, the hybrid overlap confidence estimation module integrates semantic descriptors and geometric similarity to detect overlapping regions and suppress outliers early. The context-aware matching strategy smitigates ambiguity by employing global attention to assign soft confidence scores to correspondences, improving robustness. These scores guide a differentiable weighted singular value decomposition solver to compute precise transformations. This tightly coupled pipeline adaptively down-weights uncertain regions and emphasizes contextually reliable matches. Experiments on ModelNet40, ScanObjectNN, and 7Scenes 3D vision datasets demonstrate that CEGC outperforms state-of-the-art methods in accuracy, robustness, and generalization. Overall, CEGC offers an interpretable and scalable solution to partial point cloud registration under challenging conditions.

cs.CV↗

Desensitizing for Improving Corruption Robustness in Point Cloud Classification through Adversarial Training

Due to scene complexity, sensor inaccuracies, and processing imprecision, point cloud corruption is inevitable. Over-reliance on input features is the root cause of DNN vulnerabilities. It remains unclear whether this issue exists in 3D tasks involving point clouds and whether reducing dependence on these features can enhance the model's robustness to corrupted point clouds. This study attempts to answer these questions. Specifically, we quantified the sensitivity of the DNN to point cloud features using Shapley values and found that models trained using traditional methods exhibited high sensitivity values for certain features. Furthermore, under an equal pruning ratio, prioritizing the pruning of highly sensitive features causes more severe damage to model performance than random pruning. We propose `Desensitized Adversarial Training' (DesenAT), generating adversarial samples using feature desensitization and conducting training within a self-distillation framework, which aims to alleviate DNN's over-reliance on point clouds features by smoothing sensitivity. First, data points with high contribution components are eliminated, and spatial transformation is used to simulate corruption scenes, generate adversarial samples, and conduct adversarial training on the model. Next, to compensate for information loss in adversarial samples, we use the self-distillation method to transfer knowledge from clean samples to adversarial samples, and perform adversarial training in a distillation manner.Extensive experiments on ModelNet-C and PointCloud-C demonstrate show that the propose method can effectively improve the robustness of the model without reducing the performance of clean data sets. This code is publicly available at \href{https://github.com/JerkyT/DesenAT/tree/master}{https://github.com/JerkyT/DesenAT}.

cs.CV↗