Search arXiv⌕ Search

SEARCH · Search arXiv

Search Search arXiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Exploring Forum Post Retrieval with Generative Modeling

Generative recommendation (GR) has emerged as an alternative to embedding-based retrieval, building on the success of generative models in language and vision. We are exploring GR on Facebook Forum, a standalone application for medium-to-heavy users of Facebook Groups. Because Forum is a new surface, its own interaction data are too sparse to train a GR model from scratch. We address this with transfer along two axes: we train on a broader corpus of Facebook Groups engagements rather than Forum sessions alone, and we reuse hierarchical, prefix-based semantic IDs (SIDs) learned from cross-platform Facebook Feed data instead of fitting a Forum-specific tokenizer. A 3B-parameter instruction-tuned language model is then supervised-fine-tuned to generate SIDs directly from user context. We systematically ablate the design choices that matter most in practice, including SID construction, the composition and length of user history, and the inclusion of user-profile features. Our results show that cross-platform SIDs transfer to a new recommendation surface, and offer practical guidance for teams deploying GR on real-world social platforms.

cs.IR↗

Bandits with Multiple Optimal Arms: Minimax Regret and Non-Adaptivity

We study multi-armed bandits (MAB) with multiple optimal arms, motivated by the fact that many practical decision making problems admit multiple correct answers. For $K$-armed bandits with $A$ optimal arms, we first provide a sharper analysis of previous sub-sampling algorithms (De Heide et al., 2021; Zhu and Nowak, 2020), establishing a $\tilde{O}\Big(\frac{K-A}{\sqrt{KA}}\sqrt{T} \Big)$ minimax regret, where $T$ is the total number of interactions and $\tilde O(\cdot)$ drops all constant and logarithmic factors, improving the previous $\tilde{O}(\sqrt{KT/A})$ regret. We then provide a matching lower bound up to logarithmic factors, indicating that our established rate is nearly minimax-optimal. We further show that the knowledge of $A$ up to $\tilde{O}(1)$ factors is necessary to achieve near-optimal regret, as near-optimal algorithms for one number of optimal arms must incur substantially larger regret than optimal regret for a smaller number. Overall, our results provide a comprehensive minimax characterization of $K$-armed bandits with $A$ over the entire range of $1 \leq A \leq K-1$.

stat.ML↗

Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation

Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt to scene complexity or production context. We propose SMART, a Self-evolving Multi-Agent system for long-foRm subtitle Translation. During test-time training, SMART builds persistent series-level memory and translates a subset of sentences through a dynamic router and Mixture-of-Agents layer with tools for terminology verification, subtitle constraint validation, and contextual retrieval. A judge-refiner loop scores candidates and uses textual critiques to update agent prompts and routing policies without retraining the underlying LLMs. During test-time inference, the evolved configuration translates the remaining series. We also introduce Subtitle Arena, covering 14 genres, 2--198 episodes per series, production years 1959--2023, and 15 target locales, together with SubMQM, a subtitle-adapted MQM framework with seven dimensions and 19 error categories. SMART achieves the best overall MQM score in all 15 Subtitle Arena directions, reducing average penalty by 6.9% over the strongest competing agent system. On a public benchmark, MuSC, SMART obtains the best model result across all 4 language pairs. SMART also achieves the best result in human evaluation with an overall score of 4.50/5.

cs.CL↗

dattri-LLM: A Unified and Efficient Library for Training Data Attribution at LLM Scale

Training data attribution (TDA) estimates the contribution of individual training examples to model outputs. Most scalable TDA methods rely on per-example gradients, whose computation and use at LLM scale pose challenges in efficiency, compatibility, and extensibility. We introduce dattri-LLM, a TDA library that makes gradient-based attribution more practical at scale. For efficiency, dattri-LLM uses compact gradient representations and dynamically routes gradient operations based on a cost model. For compatibility, its capture mechanism collects per-example gradients from existing training loops that call backward(), without requiring changes to the loop or its configuration. This includes distributed training with DDP and FSDP and pipelines built with HuggingFace Transformers, TRL, and OLMo. For extensibility, dattri-LLM exposes reusable gradient operations and training-time callbacks for implementing attribution methods and applications. These interfaces support a variety of attribution methods, including gradient similarity, curvature-based influence, and trajectory-based methods, as well as applications that act on gradients during training, such as online data selection. On the same hardware and workload, dattri-LLM achieves 3.2x the throughput of the fastest competing library on average, scales multiple attribution methods to 110B-parameter models across four H200 GPUs, and offers superior attribution fidelity-cost trade-offs across a range of models with different model families and scales. The source code of dattri-LLM is available at https://github.com/TRAIS-Lab/dattri-llm.

cs.LG↗

Entangling Atomic Quantum Memories Using High-Order Modulated Coherent States

High-order coherent-state modulation and collective midpoint measurements enable multi-ebit entanglement distribution between cavity-coupled atomic memories. An SRM-inspired 16-QAM receiver achieves $1.75$ ebits per network mode use at $0.5$~dB end-to-end loss, exceeding single-qubit-per-mode benchmarks by $5.8$~dB and lying $3.8$~dB below the half-link capacity bound. The advantage persists for per-interface loss below $0.2$~dB and calibrated phase errors below $0.2π$. For 4-QAM, an optimized POVM with twice as many outcomes raises the achievable rate by $4.0\%$, revealing measurement-design headroom.

quant-ph↗

Blackboard Intelligence Can Surpass Autoregressive on Globally Constrained Problems

Next-token prediction has driven remarkable progress in large language models, yet a growing body of evidence suggests that they can struggle on problems governed by complex global constraints. In this work, we focus on this regime and ask whether some of these limitations arise from the inference interface induced by next-token prediction itself. We study this question through blackboard intelligence: an inference-time perspective in which a model works on a fixed, revisable canvas and searches over candidate solution states rather than committing to a causal, left-to-right trajectory. We instantiate this idea with diffusion language models, whose any-order prediction interface naturally exposes predictions over partially filled solution states. Our key observation is that mean confidence, a simple model-internal quantity available from the standard masked diffusion objective, provides a useful proxy for global coherence and can guide inference-time search and revision. Empirically, across ZebraLogic, Nurse Rostering, and Job-Shop Scheduling, Blackboard consistently improves inference while holding the fine-tuned LLaDA-8B-Instruct checkpoint fixed and substantially outperforms same-scale autoregressive baselines, reaching 90.4% accuracy on ZebraLogic-Hard, 76.4% exact feasibility on Nurse Rostering, and 80.2% optimality on JSSP. Stronger autoregressive search and refinement also fail to close the gap on ZebraLogic-Hard, while Blackboard surpasses tested frontier LLMs there and on JSSP despite their substantially greater scale and strong test-time reasoning. We open-source our codebase at https://github.com/jwoosang1/blackboard-intelligence.

cs.LG↗

Divisors and harmonic morphisms on metric graphs of pseudocompact type

A metric graph is of pseudocompact type if identifying parallel edges produces a tree. We give a constructive proof that, for such graphs, divisorial $d$-gonality is equivalent to the existence of a degree $d$ harmonic morphism to a tree. This mirrors the algebraic correspondence, for curves of compact type, between limit linear series of dimension one and admissible covers. We also deduce lifting results for positive-rank divisors, with a genus-preserving refinement when identifying parallel edges produces a path. Finally, we study the Brill-Noether theory in the path case.

math.AG↗

scTrilemma: Balancing Identity, Invariance, and Fidelity in Single-Cell Representation Learning

Single-cell RNA-seq representation learning is fundamentally label-free: cell identities, states, and contexts are not fixed training targets, so what constitutes signal or nuisance is analysis-dependent. A single representation must therefore preserve biological identity and state, remain robust to nuisance context, and retain the gene-level variation needed for expression analysis, three demands we call the representation trilemma. To tackle this problem, we introduce scTrilemma, a latent-bottleneck VAE that routes expression-derived variation to the embedding, the decoder, or the prior rather than forcing all of it through one embedding. It gates gene tokens by expression, routes the cell representation through the decoder, and conditions the prior on unlabeled pseudo-bulk context, under a single reconstruction objective and without target annotations or auxiliary representation losses. In release-based zero-shot evaluation on successive CZ CELLxGENE Census releases, scTrilemma leads all three demands at once and preserves biological-state, differential-expression, and pathway structure across multiple disease settings. Latent interventions further show that context can be removed at almost no cost to the other demands, leaving identity against fidelity as the remaining tension. Code is publicly available at https://github.com/yunhak0/scTrilemma.

cs.LG↗

EmbodiRSI: Recursive Self-Improvement for Data-Efficient Robot Adaptation

Adapting robot manipulation policies to new tasks and environments remains highly data-intensive, while the data needed for further improvement depends on the policy's current capabilities and failure modes. We introduce EmbodiRSI, an agentic system for recursive self-improvement (RSI) in a real-to-sim-to-real setting, where task-specific simulations are constructed from target deployment scenarios and used as low-cost environments for iterative policy improvement before transfer back to the physical world. EmbodiRSI uses policy execution feedback to guide subsequent experience acquisition and policy updates. Two complementary mechanisms close this loop: Collaborative Error Correction generates agent-assisted corrective trajectories from policy-reached states, while Adaptive Data Collection directs expert demonstration generation toward the current policy's weaknesses. The task-specific simulation serves as a reusable workspace for policy warm-up, repeatable evaluation, failure diagnosis, and targeted data generation across successive RSI rounds. Across three tabletop environments and 14 subtasks, EmbodiRSI increases scene-balanced autonomous simulation success from 50.4% to 83.5% over two RSI updates. With 400 adaptive simulated trajectories and only ten real-world refinement trajectories per subtask, EmbodiRSI achieves 83.1% scene-balanced autonomous real-world success, compared with 75.0% for adaptation using 200 real-world demonstrations per subtask. These results demonstrate that feedback-driven recursive improvement in deployment-specific simulations can enable data-efficient adaptation of embodied policies to physical environments.

cs.RO↗

Boundary-rank obstructions and measurement-assisted recovery in dissipative flat-band preparation

Number-conserving cooling can fail to prepare an interacting flat-band target when Pauli exclusion blocks its transfer destinations. We study this failure on vertex-edge decorated graphs at one particle per flat orbital. The rank of a cut block of the flat-orbital Gram matrix bounds the number of residual source directions that can remain flat. For sublinear-range cooling with no Hamiltonian term, separated spin domains on periodic decorated hypercubic lattices support exponentially many stationary states with a nonzero bright-particle density. For a chain with fixed finite cooling range, we construct a physical Fock initial state whose overlap with a wrong dark state is independent of system size. This gives fidelity and bright-density bounds valid at every time. In higher dimensions, the bare overlap decays with boundary area, while local boundary rotations prepare states with a finite failure weight in constant circuit depth. Dephasing the occupations of all compact bright modes restores global attraction to the ferromagnetic target when the graph and cooling destinations satisfy the stated conditions. The result allows Hamiltonians that preserve the target. The compressed dynamics in the strong-dephasing limit also remains attractive and has a positive gap at each fixed size. Gram calculations, finite-system dynamics, and Liouvillian spectra test the analytic results. Weak-dephasing and particle-transport bounds constrain the preparation time despite eventual convergence.

quant-ph↗

Tight Post-Quantum Parallel Repetition for Private-Coin Arguments

We show that assuming the existence of homomorphic encryption, parallel repetition of all interactive arguments (after being run under homomorphic encryption) reduces the soundness error at a tight exponential rate even in the post-quantum setting. Moreover, we generalize this result to hold for threshold verifiers, where the parallel repeated verifier accepts if and only if at least $t$ of the executions are accepted (for some threshold $t$). Prior to this work, these results were known only when the cheating prover was assumed to be classical, and it was not known how to achieve tight bounds. As a corollary, we construct the first constant-round succinct argument for $\mathsf{QMA}$ with negligible completeness and soundness errors assuming only the existence of quantum homomorphic encryption.

quant-ph↗

Mitigating Object Hallucination in Large Vision-Language Models via False Discovery Controlled Visual Data Splitting

Multiple object hallucination, where large vision-language models (LVLMs) generate objects not supported by the visual input, is a persistent challenge caused by visual uncertainty during decoding. Existing methods reduce hallucinations using contrastive signals, but they rely on heuristics and lack principled control of false positives at the image level. To address this, we propose False Discovery Rate-COntRol of HALlucination (CORAL), a training-free framework that models visual uncertainty using an uncertainty-aware visual data splitting strategy and leverages mirror statistics to quantify visual contrast during decoding. By computing mirror statistics from paired, symmetrically perturbed visual inputs, CORAL estimates spurious object predictions and sets a data-driven threshold to control the expected fraction of false discoveries per image, suppressing hallucinations while retaining high power for truly grounded objects. The framework is flexible, supports multiple LVLMs, and mitigates hallucinations without retraining or supervision. Extensive experiments on multiple benchmarks with several evaluation metrics demonstrate that CORAL consistently outperforms state-of-the-art methods, providing more reliable and robust hallucination control. Code is available at: https://changliu1993-cl.github.io/CORAL/

cs.CV↗

From Verification Failures to Reusable Guidance for Coding Agents

Coding agents need to establish that a program satisfies a specification and that the specification captures the requested behavior. We study how expert diagnosis of verification failures can become reusable guidance for this work. Our approach combines executable language definitions in the K framework with a kit of procedures for constructing specifications, repairing proofs, and auditing their adequacy. A human-guided development campaign on HumanEval, a benchmark of 164 Python programming tasks, achieves a 164/164 success rate with the semantics and the kit, measured by final AI audit Pass verdicts after two targeted repairs. To examine whether auditing detects problems that successful proofs leave unresolved, we construct 12 author-reviewed pairs of clean and defective packages. Every package passes its K proofs, and completed audits identify all defects and accept all clean packages. We then use KleverBench to test specification and proof construction for 31 programs with changed operator meanings. Comparisons with complete acceptance rules and equally long generic advice yield mixed results across two model and budget settings, motivating further work on selecting useful guidance within resource limits. Human-reviewed Optimism proofs establish expected pause reverts for six operations within declared input bounds under London semantics with unbounded gas. We report progress, difficulties, and lessons toward agents that deliver programs with checkable correctness arguments.

cs.SE↗

Metachecks in Bivariate Bicycle Codes: Syndrome Distance, Measurement Faults, and Repair Limits

Faulty syndrome measurements can corrupt an otherwise correct quantum-error correction step. Bivariate bicycle (BB) codes contain dependent stabilizer checks, so every valid syndrome obeys additional parity constraints, or metachecks. We study how far this built-in redundancy can identify measurement faults and when the remaining ambiguity is unavoidable, while separately checking the code's logical structure. A logical decomposition is used as a preliminary safety check: it identifies a $k/2$-dimensional annihilator subspace and a $k/2$-dimensional colon quotient, and the minimum-weight logical need not be visible from the annihilator side alone. On the measurement side, translation symmetry partitions syndrome locations into classes that carry identical metacheck information. This gives an exact characterization of the leading single-fault ambiguity, a bound on how many fault locations can be distinguished, and a family-level repair limit when the encoded dimension stays bounded while the block length grows. Under a static-data assumption, the same calculation gives the minimum number of checks that must be remeasured to remove every single-fault ambiguity. Exact finite-code calculations illustrate both regimes: all single measurement faults are distinguishable in a 72-qubit BB code, whereas the 144-qubit Gross code merges its 72 syndrome locations into 36 indistinguishable pairs. In a 108-qubit example, one logical component first appears at weight 12 while the other contains a weight-10 logical. Sustained phenomenological experiments show that joint data--measurement decoding is more robust than a separated repair stage on the more ambiguous codes. The resulting tests apply to general two-block BB codes, including non-coprime periods and repeated-root cases.

quant-ph↗

Sharp Stationary Gaussian Approximation for Constant-Stepsize SGD

We prove a sharp Gaussian approximation for the invariant law of constant-stepsize SGD with bounded additive noise generated by an exogenous uniformly ergodic Markov chain. For a smooth, strongly convex objective with a Lipschitz Hessian and nondegenerate long-run noise covariance, the centered iterate normalized by the square root of the stepsize is $O(\sqrtα)$-close in 1-Wasserstein distance to its limiting Gaussian. The proof combines blockwise Gaussian comparison with long-run contraction. A four-state example gives a matching lower bound although the one-time noise marginal is symmetric and every nonzero-lag autocovariance vanishes. In this example, an adjacent third-order mixed moment produces the leading correction.

cs.LG↗

Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds

Persistent spatial memory enables embodied agents to navigate familiar environments across repeated visits. However, targets may move while unobserved, including during navigation, making remembered locations unreliable by the time an agent arrives. Despite advances in memory retrieval and state prediction, accounting for continued hidden world evolution and revising beliefs under limited visibility remain challenging. We study Evolving-World Navigation, where agents infer target locations from intermittent observations, predict their states at inspection time, and revise beliefs using visual evidence. We propose EvolvingNav, which constructs a time-indexed belief from timestamped 3D object histories through a structured persistence-relocation model. The belief distinguishes persistence at the last observed location from relocation to alternative locations and retains probability mass outside the known candidate set. An event-driven filter propagates the current belief as time elapses, forecasts target occupancy at candidate inspection times, and incorporates new RGB-D evidence. Negative observations downweight location hypotheses according to calibrated, visibility-conditioned detection probabilities, while evidence tracking prevents repeated use of the same observations. A frozen, zero-shot vision-language controller uses the updated belief to choose actions and replan. We further introduce EvoWorld-Bench, a benchmark grounded in human activity traces, comprising 54 scenes and 803,680 tasks with controlled changes before and during navigation. In simulation and real-robot experiments, EvolvingNav improves navigation success and search efficiency over the evaluated baselines. Paired experiments show the clearest gains under learnable temporal patterns, while ablations demonstrate the value of preserving uncertainty and incorporating visibility-aware evidence.

cs.AI↗

Fast and Sample Efficient Safety Verification via Extreme Learning Machine

Deep learning methods like neural networks have greatly simplified the computation of safety certificates for complex nonlinear systems with unknown dynamics. However, due to the data-driven nature of these certificates and the complex architecture of neural networks, computation time as well as robustness guarantees across unseen data remain a challenge. This work aims to formally verify safety properties of discrete-time unknown systems by synthesizing extreme learning machine (ELM)-based barrier certificates. Compared to neural network counterparts, this approach greatly improves convergence guarantees and computational time due to its architectural simplicity and the convex nature of the underlying optimization problem. By minimizing the Lipschitz constant of the candidate barrier, we present a grid-based sampling technique to formally verify its validity using the minimum number of samples required. We demonstrate through numerical examples the effectiveness of our approach, and compare with traditional deep-learning based certificate synthesis to highlight its benefits.

eess.SY↗

QATFactory: A Versatile, Deployment-Aligned Framework for Quantization-aware Training and Distillation of LLMs

Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive post-training quantization (PTQ) can degrade model quality. We present QATFactory, an open-source framework for deployment-aligned quantization-aware distillation (QAD) and reinforcement learning (QARL). QATFactory simulates deployment-time quantization while performing matrix multiplications in BF16, allowing models to adapt to quantization noise without requiring training hardware that natively supports the target format; for example, it supports NVFP4 training on H100 GPUs, which lack FP4 Tensor Cores. The framework supports NVFP4, MXFP4, and $\text{llama}.\text{cpp}$'s Q4_K format; dense and mixture-of-experts models; and both full-parameter and LoRA-based training. It exports checkpoints directly to vLLM and $\text{llama}.\text{cpp}$ without an additional lossy conversion step or added inference overhead. With QATFactory, we conduct extensive experiments on models ranging from 8B to 230B parameters and evaluate exported checkpoints in production inference engines. Across models and formats, QAD consistently improves deployed-model quality over strong PTQ baselines. On Qwen3.5-9B, QAD achieves average benchmark accuracies of 68.9% under NVFP4 and 66.0% under MXFP4, outperforming the best PTQ results of 65.4% and 56.4%, respectively. Through our experiments, we found that although both FP4 formats quantize weights and activations at deployment, the best training strategy is format-dependent: NVFP4 generally performs better when only weights are quantized during training, whereas MXFP4 benefits from quantizing both weights and activations. At a fixed training token budget, training on fewer 32K sequences improves average accuracy by 1.9 points over training on more 4K sequences.

cs.LG↗