Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 181 records · Page 10Linked to original sources

Finite Index and Do Carmo Question for Constant Mean Curvature Hypersurfaces

We prove that any finite $δ$-index hypersurface $M$ in ${\mathbb R}^{n+1}$ with constant mean curvature must be minimal, provided either of the following conditions holds: - the volume growth of $M$ is sub-exponential; - the Ricci curvature of $M$ satisfies $\operatorname{Ric}_M\geq -\frac{C(1-δ)}{n-1}|A|^2g,$ where $A$ is the second fundamental form, $g$ is the metric on $M$ and $C$ is a positive constant smaller than 4. We emphasize that no restriction on the dimension is imposed. Moreover, the statement in the second case is new even for finite index hypersurfaces ($δ=0$).

math.DG↗

Survival is the Only Reward: Sustainable Self-Training Through Environment-Mediated Selection

Self-training systems often degenerate due to the lack of an external criterion for judging data quality, leading to reward hacking and semantic drift. This paper provides a proof-of-concept system architecture for stable self-training under sparse external feedback and bounded memory, and empirically characterises its learning dynamics and failure modes. We introduce a self-training architecture in which learning is mediated exclusively by environmental viability, rather than by reward, objective functions, or externally defined fitness criteria. Candidate behaviours are executed under real resource constraints, and only those whose environmental effects both persist and preserve the possibility of future interaction are propagated. The environment does not provide semantic feedback, dense rewards, or task-specific supervision; selection operates solely through differential survival of behaviours as world-altering events, making proxy optimisation impossible and rendering reward-hacking evolutionarily unstable. Analysis of semantic dynamics shows that improvement arises primarily through the persistence of effective and repeatable strategies under a regime of consolidation and pruning, a paradigm we refer to as negative-space learning (NSL), and that models develop meta-learning strategies (such as deliberate experimental failure in order to elicit informative error messages) without explicit instruction. This work establishes that environment-grounded selection enables sustainable open-ended self-improvement, offering a viable path toward more robust and generalisable autonomous systems without reliance on human-curated data or complex reward shaping.

cs.AI↗

Space Filling Curves is All You Need: Communication-Avoiding Matrix Multiplication Made Simple

General Matrix Multiplication (GEMM) is the cornerstone of HPC workloads and Deep Learning. State-of-the-art (SOTA) vendor libraries tune tensor layouts, parallelization schemes and cache blocking to minimize data movement across the memory hierarchy and maximize throughput. However, optimal settings for these parameters depend on the target platform and matrix shapes, making exhaustive tuning infeasible. In this work, we address this cumbersome scheduling search and tuning using space-filling curves (SFC). We partition the matrix multiplication using advancements in SFC, and obtain platform-oblivious and shape-oblivious matrix multiplication schemes with a high degree of data locality. We extend the SFC-based work partitioning to implement Communication-Avoiding (CA) algorithms with replication techniques in a seamless fashion. The resulting SFC-CA GEMM achieves provable asymptotic communication optimality for both square and rectangular matrix regimes. Across four x86 and Arm platforms, SFC-CA GEMM outperforms vendor libraries by up to 5.5$\times$ per shape and 1.8$\times$ in weighted harmonic mean (WHM) throughput. Last, we show the impact of our work on two real-world applications by leveraging our SFC-CA GEMM as a compute backend: i) prefill of LLM inference with speedups up to 1.85$\times$ over SOTA inference runtimes, and ii) distributed-memory matrix multiplication with speedups up to 2.3$\times$ over the SOTA distributed-memory GEMM framework with vendor-optimized compute backend.

cs.DC↗

An AI-ready fine-tuning framework for accurate machine-learning interatomic potentials in solid-solid battery interfaces

Atomistic modeling of solid-solid battery interfaces is essential for understanding electro-chemo-mechanical coupling, but the complex interfacial chemistry and heterogeneous environments pose major challenges for quantum-accurate, data-efficient modeling. Herein, we propose an approach of fine-tuning with integrated replay and efficiency (FIRE), a general framework for universal machine-learning interatomic potentials by combining efficient configurational sampling with a replay-argumented continual strategy, achieving quantum-level accuracy at moderate cost. Across six solid-solid battery interface systems, FIRE consistently achieves root-mean-square errors in energy below 1 meV/atom and in force near 20 meV/angstrom, marking an order-of-magnitude improvement over existing models while requiring only 10% of the original datasets. In addition, the fine-tuned model successfully reproduces key mechanical and electrochemical properties of the materials, in close agreement with experimental data. The FIRE offers a generalizable and data-efficient approach for developing accurate interatomic potentials across diverse materials, enabling predictive simulations beyond the reach of first-principles methods.

cond-mat.mtrl-sci↗

ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule

We consider time discretization for score-based diffusion models to generate samples from a learned reverse-time dynamic on a finite grid. Uniform and hand-crafted grids can be suboptimal given a budget on the number of time steps. We introduce Adaptive Reparameterized Time (ART), which controls the clock speed of a reparameterized time variable to redistribute computation along the sampling trajectory while preserving the terminal time, with the objective of minimizing the aggregate Euler discretization error. We derive a randomized companion ART-RL that recasts ART as a continuous-time reinforcement learning problem with Gaussian policies, and prove a two-directional bridge between the two: the deterministic ART optimum lifts to an optimal Gaussian policy, and conversely any optimal Gaussian policy must recover the ART control through its mean. This bridge turns continuous-time actor--critic learning into a principled, rather than heuristic, route to the deterministic timestep optimum. Within the official EDM pipeline, ART-RL improves FID on CIFAR--10 across a wide range of budgets; after one-time offline training, the distilled deterministic schedule transfers without retraining to AFHQv2, FFHQ, and ImageNet at no extra inference cost.

cs.LG↗

Efficient Application of Tensor Network Operators to Tensor Network States Through Successive Deterministic Compression

We introduce an algorithm to apply tree tensor network operators to tree tensor network states followed by compression. Especially when applying a matrix product operator to matrix product states, this serves as a key subroutine in tensor network algorithms. We name the method successive deterministic compression (SDC). We compare SDC to existing operator application and compression algorithms in benchmarks involving random tensor network operators and states, as well as circuit simulation with standard entangling gates. We find that successive deterministic compression is consistently among the best-performing techniques when the tensor network state (after the operator application) is not highly compressible. Our results also highlight the flexibility of general tree tensor network structures for simulating circuits with long-range entanglement.

quant-ph↗

Lowest Span Confidence: Zero-Shot Hallucination Detection from a Single LLM Response

Hallucinations in Large Language Models (LLMs), i.e., plausible but non-factual generations, pose a significant challenge to reliable deployment in high-stakes environments. However, many existing hallucination detectors require expensive repeated sampling for consistency checks or access to model-internal states unavailable in common API-based scenarios. To this end, we propose an efficient zero-shot metric called Lowest Span Confidence (LSC) for hallucination detection under minimal resource assumptions. Concretely, LSC evaluates the local confidence of adjacent complete-word spans. By selecting the lowest aggregated confidence across neighboring words whose token widths can vary, LSC captures localized uncertainty associated with factual inconsistency. This boundary-aligned smoothing reduces the global dilution of perplexity and the sensitivity of minimum token probability to isolated noise. Our main evaluation spans four model families {Llama-2, Qwen2.5, Gemma-2, Mistral} and seven benchmarks {NQ, TriviaQA, SQuAD, CoQA, HotpotQA, RAGTruth, FELM}. Additional analyses examine word reconstruction, span width, and the role of adjacency in preserving local confidence. Across these settings, LSC is competitive with methods that use multiple responses or model-internal information while requiring only one response and its output token probabilities, without training a separate detector or using an auxiliary model.

cs.CL↗

OGPO: One-Step Generative Policy Optimization for Real-Time Robot Control

Real-time robot control demands fast action generation. Diffusion and flow matching policies for robot control require multi-step sampling, limiting their deployment in real-time scenarios. Natively reducing the sampling steps to one sacrifices representation quality and task performance, creating a trilemma among speed, fidelity, and performance. We present One-Step Generative Policy Optimization (OGPO), a systematic framework to resolve this trilemma. OGPO first pairs a lightweight architecture with the interval velocity principle for distillation-free one-step inference, while representation spreading prevents representation quality degradation. It then performs on-policy reinforcement learning (RL) fine-tuning on this fast, stable policy to break the imitation learning ceiling. Experiments on RoboMimic and OpenAI Gym benchmarks show that OGPO matches or exceeds multi-step baselines while achieving a 5-20 times inference speedup and over 120Hz control frequency. Physical deployment on a Franka-Emika-Panda robot validates real-world applicability. Project page: https://ogpo-project.github.io/

cs.RO↗

Tunneling probe-based characterisation of the sp${}^3$ dangling bond on the H-C(100):$2\times1$ surface

The sp${}^3$ dangling bond on the diamond surface plays a critical role in the performance and fabrication of diamond quantum technologies. For the former, the magnetic and electric properties of this defect can impede the performance of quantum sensors and computers. For the latter, the chemical properties of the dangling bond are integral to proposed methods for bottom-up fabrication of scalable diamond quantum devices. In pursuit of high-performance, scalable diamond quantum technology, tunnelling probe-based techniques offer the ability to create and modify the sp${}^3$ dangling bond with atomic-scale precision. However, these capabilities cannot be realised either deterministically or at scale without a means of identifying the sp${}^3$ dangling bond amidst the myriad of other defects on the diamond surface. Consequently, in this work we provide a comprehensive experimental and theoretical framework for STS-based characterisation of the sp${}^3$ defect on the H-terminated (100) diamond surface. This capability provides the foundation for future tunnelling probe studies in the modification of dangling bonds.

cond-mat.mes-hall↗

Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

The generalization capabilities of robotic manipulation policies are heavily influenced by the choice of visual representations. Existing approaches typically rely on representations extracted from pre-trained encoders, using two dominant types of features: global features, which summarize an entire image via a single pooled vector, and dense features, which preserve a patch-wise embedding from the final encoder layer. While widely used, both feature types mix task-relevant and irrelevant information, leading to poor generalization under distribution shifts, such as changes in lighting, textures, or the presence of distractors. In this work, we explore an intermediate structured alternative: Slot-Based Object-Centric Representations (SBOCR), which group dense features into a finite set of object-like entities. This representation permits to naturally reduce the noise provided to the robotic manipulation policy while keeping enough information to efficiently perform the task. We benchmark a range of global and dense representations against intermediate slot-based representations, across a suite of simulated and real-world manipulation tasks ranging from simple to complex. We evaluate their generalization under diverse visual conditions, including changes in lighting, texture, and the presence of distractors. Our findings reveal that SBOCR-based policies outperform dense and global representation-based policies in generalization settings, even without task-specific pretraining. These insights suggest that SBOCR is a promising direction for designing visual systems that generalize effectively in dynamic, real-world robotic environments.

cs.RO↗

Dynamics-Inspired Diffusion for Foreground-Preserving Document Background Editing

We revisit diffusion-based generation for structured visual content and identify a fundamental limitation of existing approaches: foreground preservation and background stylization are typically enforced through external interventions, such as hard masking or corrective post-processing, rather than arising from the generative process itself. Here, we define background as the generative content outside designated foreground regions (e.g., text and layout elements), while preserving the structural integrity of the foreground. We propose a dynamical systems perspective on diffusion, in which controllable generation is formulated as trajectory shaping in latent space. Under this view, we introduce Auxiliary Context Diffusion (ACD), a state-space control framework that integrates heterogeneous signals (layout-derived foreground indicators, document summaries, and style representations) directly into the diffusion dynamics. This formulation induces time-scale separation in the generative process, where foreground regions become dynamically stabilized while background regions remain expressive. To address stylistic drift across multi-page documents, we further introduce style directions as persistent latent constraints that guide diffusion trajectories within a shared stylistic subspace. Unlike prior approaches that entangle style with prompt conditioning, our formulation enables reusable and consistent style control across pages. We validate the proposed perspective through controlled experiments on synthetic document benchmarks, demonstrating that trajectory-level control provides a unified and extensible mechanism for structured generation without retraining, hard masking, or corrective post-processing. These results suggest a new direction for controllable diffusion in document-centric and multimodal applications.

cs.CV↗

Contrastive Representation Shaping for LLM Unlearning

Most LLM unlearning methods aim to approximate retrain-from-scratch behaviors with minimal distribution shift, often via alignment-style objectives defined in the prediction space. While effective at reducing forgotten content generation, such approaches may act as suppression: forgotten concepts can persist in representations and remain entangled with retained knowledge. We introduce CLReg, a contrastive representation regularizer that identifies forget features while pushing them away from retain features, reducing forget--retain interference while empirically preserving the scale and shape of retain features. As light motivation for the mechanism, we provide a one-step analysis showing that CLReg decreases a simple entanglement proxy in the embedding space. Across unlearning benchmarks and LLMs of different sizes, CLReg decreases forget-retain representation entanglement to enhance mainstream unlearning methods without extra privacy risks, inspiring future unlearning work to remove forget concepts via representation shaping. Code is available at https://github.com/HaoranTang/CLReg.

cs.LG↗

Fairness-Aware Performance Evaluation for Multi-Party Multi-Objective Optimization

In multiparty multiobjective optimization problems, solution sets are usually evaluated using classical performance metrics, aggregated across DMs. However, such mean-based evaluations may be unfair by favoring certain parties, as they assume identical geometric approximation quality to each party's PF carries comparable evaluative significance. Moreover, prevailing notions of MPMOP optimal solutions are restricted to strictly common Pareto optimal solutions, representing a narrow form of cooperation in multiparty decision making scenarios. These limitations obscure whether a solution set reflects balanced relative gains or meaningful consensus among heterogeneous DMs. To address these issues, this paper develops a fairness-aware performance evaluation framework grounded in a generalized notion of consensus solutions. From a cooperative game-theoretic perspective, we formalize four axioms that a fairness-aware evaluation function for MPMOPs should satisfy. By introducing a concession rate vector to quantify acceptable compromises by individual DMs, we generalize the classical definition of MPMOP optimal solutions and embed classical performance metrics into a Nash-product-based evaluation framework, which is theoretically shown to satisfy all axioms. To support empirical validation, we further construct benchmark problems that extend existing MPMOP suites by incorporating consensus-deficient negotiation structures. Experimental results demonstrate that the proposed evaluation framework is able to distinguish algorithmic performance in a manner consistent with consensus-aware fairness considerations. Specifically, algorithms converging toward strictly common solutions are assigned higher evaluation scores when such solutions exist, whereas in the absence of strictly common solutions, algorithms that effectively cover the commonly acceptable region are more favorably evaluated.

cs.NE↗

Manifold-Aware Perturbations for Constrained Generative Modeling

Generative models have enjoyed widespread success in a variety of applications. However, they encounter inherent mathematical limitations in modeling distributions where samples are constrained by equalities, as is frequently the setting in scientific domains. In this work, we develop a computationally cheap, mathematically justified, and highly flexible distributional modification for combating known pitfalls in equality-constrained generative models. We propose perturbing the data distribution in a constraint-aware way such that the new distribution has support matching the ambient space dimension while still implicitly incorporating underlying manifold geometry. Through theoretical analyses and empirical evidence on several representative tasks, we illustrate that our approach consistently enables data distribution recovery and stable sampling with both diffusion models and normalizing flows.

cs.LG↗

Heavy-particle production during inflation and its gravitational-wave signal

We study a mechanism for producing heavy particles during inflation through a transient quadratic $U(1)$-breaking interaction of a complex scalar coupled derivatively to the inflaton. The rolling inflaton induces an effective chemical potential, while sufficiently strong symmetry breaking creates a tachyonic instability that amplifies the scalar fluctuations. The finite duration of the instability limits their growth, allowing efficient production of particles heavier than the Hubble scale in regimes where homogeneous backreaction remains small. The amplified fluctuations source a peaked gravitational-wave spectrum through their anisotropic stress. We calculate this spectrum by following the full time evolution of the source. For suitable parameters and production times, our idealized forecasts suggest that the signal could be observable with LISA, subject to the assumed switching profile and reheating history.

astro-ph.CO↗

Functional Subspace, where language models can use vector algebra to solve problems

Large language models (LLMs) were invented for natural language tasks such as translation, but they have proved that they can perform highly complex functions across domains. Additionally, they have been thought to develop new skills without being trained on them. These learning capabilities lead to LLMs adoption in a wide range of domains. Thus, it is imperative that we understand their operating mechanisms and limitations for proper diagnostics and repair. The earlier studies proposed that high level concepts are encoded as linear directions in LLMs activation space and that the geometry of embeddings have semantic meanings. Inspired by these studies, we hypothesize that LLMs may use subspaces and vector algebra in subspaces to perform tasks. To address this hypothesis, we analyze LLMs' functional modules and residual streams collected from LLMs engaging in in-context learning (ICL), one of the emergent abilities. Our analyses suggest that 1) LLMs can create subspaces, where evidence can be accumulated and 2) ICL tasks can be solved via simple algebraic operations in subspaces.

cs.CL↗

MatGPTQ: Efficient and Accurate Inference over Nested Quantized Models

Matryoshka Quantization (MatQuant), Any-Precision-LLM (AP) and AnyBCQ (AB) are recent quantization approaches showing that a single integer-quantized model can be served across multiple precisions. In this paradigm, lower-precision models are extracted from a higher-precision model by simply reading fewer bits of the weights. This enables a single checkpoint to cover a wide range of memory and latency budgets, but makes both quantization and efficient execution substantially harder. Existing methods rely on expensive quantization-aware training (QAT) or gradient-based post-training quantization (PTQ) rather than fast one-shot PTQ, and offer limited system support: dedicated kernels are either missing or restricted to single- or small-batch decoding. We address these limitations with Post-Training Matryoshka Quantization (MatGPTQ), an end-to-end pipeline for nested-model quantization and inference. MatGPTQ casts Matryoshka quantization as multi-precision error compensation, producing a single "sliceable" parent model jointly optimized for multiple target precisions in one pass over a small calibration set. We further refine the MatQuant representation so that an $r$-bit model reads exactly $r$ bits, and introduce the first dedicated inference kernels for this format, supporting batch sizes beyond one and integrated into vLLM. Across standard LLMs and benchmarks, MatGPTQ outperforms MatQuant while remaining competitive with AP and AB at the smallest checkpoint size, and our kernels achieve end-to-end speedups of up to 3.5$\times$ over BF16 at the low-bit regime. Overall, MatGPTQ makes nested quantized models practical to serve from a single, compact checkpoint. Code is available at https://github.com/IST-DASLab/MatGPTQ.

cs.LG↗

Multiplicative Subgroups of $\mathbb{Z}_p^*$ that are Generalized Arithmetic Progressions

We prove that a multiplicative subgroup $A_k$ of $\mathbb{Z}_p^*$ is a generalized arithmetic progression if and only if $|A_k| = 2,\ 4,$ or $p-1$. Much of the argument builds upon recent work studying additive decompositions of subgroups, and we generalize a result of Hanson and Petridis to show that any additive $n$-decomposition of a subgroup must be a direct sum. We also show how this classification quickly follows from Kalmynin's recent work resolving Sárközy's conjecture for quadratic residues.

math.NT↗