Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 217 records · Page 12Linked to original sources

Revisiting transportation problems under Monge costs with applications to the Discrete Ordered Median Problem

We investigate the classical Hitchcock-Koopmans Transportation Problem under the assumption that the cost matrix satisfies the Monge property. In this setting, we derive compact formulas for optimal dual solutions based on the northwest-corner rule. As an application illustrating how these formulas yield structural insight while enhancing computational performance, we consider a broad class of facility location problems. In particular, the expressions are used within a Benders decomposition framework to derive novel mixed-integer linear programming formulations for the Discrete Ordered Median Problem with non-increasing weights. Numerical experiments validate that the resulting formulations achieve state-of-the-art performance and exhibit strong robustness across a wide range of instances.

math.OC↗

Dynamics of perturbed elliptical billiard tables

Dynamical billiards consist of a particle on a two-dimensional table, bouncing elastically off the boundary. The state of the system is given by two numbers: one describing the location along the curve where the bounce occurs, and another describing the incoming angle of the trajectory before the bounce. Tracking these numbers over successive bounces defines a two-dimensional area preserving map, and iterating this map gives a dynamical system first studied by Birkhoff. Although there are powerful theoretical results showing that generic (strictly convex) billiards exhibit chaotic dynamics, it is nevertheless difficult (if not impossible) to decide when a given billiard table is generic and one will often resort to numerical computations. In this paper, we employ the parameterization method to compute high order Taylor expansions of the local stable/unstable manifolds attached to periodic orbits of billiard maps. Globalizing appropriate fundamental domains locates transverse intersections, providing insight into the existence and location of chaotic invariant sets. A key step in implementing the parameterization method is computing the composition of a given polynomial with the (implicitly defined) billiard map, and we show that this can be done efficiently using the discrete Fourier transform (DFT). The DFT requires evaluating the composition on a disk in the complex plane, so that we must first extend the billiard table/map to a complex domain. This problem is addressed from a computational perspective.

math.DS↗

Efficient Generative Modeling beyond Memoryless Diffusion via Adjoint Schrödinger Bridge Matching

Diffusion models often yield highly curved trajectories and noisy score targets due to an uninformative, memoryless forward process that induces independent data-noise coupling. We propose Adjoint Schrödinger Bridge Matching (ASBM), a generative modeling framework that recovers optimal trajectories in high dimensions via two stages. First, we view the Schrödinger Bridge (SB) forward dynamic as a coupling construction problem and learn it through a data-to-energy sampling perspective that transports data to an energy-defined prior. Then, we learn the backward generative dynamic with a simple matching loss supervised by the induced optimal coupling. By operating in a non-memoryless regime, ASBM produces significantly straighter and more efficient sampling paths. Compared to prior works, ASBM scales to high-dimensional data with notably improved stability and efficiency. Extensive experiments on image generation show that ASBM improves fidelity with fewer sampling steps. We further showcase the effectiveness of our optimal trajectory via distillation to a one-step generator.

cs.CV↗

A discrete gradient scheme for preserving QSR-dissipativity

The notion of dissipative dynamical systems provides a formal description of processes that cannot generate energy internally. For these systems, changes in energy can only occur due to an external energy supply or dissipation effects. Unfortunately, dissipative properties tend to deteriorate in numerical computations, especially in nonlinear systems. Discrete gradient methods can help mitigate this problem. In this paper, we present a class of structure-preserving time discretization schemes based on discrete gradients for a special class of systems that are dissipative with respect to a quadratic supply rate.

math.NA↗

AI as Coordination-Compressing Capital: Task Reallocation, Organizational Redesign, and the Regime Fork

Task-based models of AI hold organizational structure fixed. We model AI as agent capital that compresses managers' per-link coordination costs toward residual floors for consequential decisions. Positive floors bound spans of control; zero floors do not. For a fixed manager pool, proportional team sizes, equal mean team quality and fixed within-firm pay shares, and when every floor is positive, the floors determine long-run managerial wage dispersion, independent of execution-compression parameters. With a common positive floor, distinct skills and skill-biased execution savings, managerial inequality is purely transitional: zero at zero capital and in the long run, positive at every positive finite capital level. If every lower-floor manager approaches capacity faster than every higher-floor one, inequality exceeds its long-run level at sufficiently large finite capital and converges from above, a non-monotone path. Numerical illustrations, not estimates, include a simulation with skill-sorted workers, to which the wage and Gini results do not apply.

econ.GN↗

Computing Equilibria in Games with Stochastic Action Sets

The study of learning in games typically assumes that each player always has access to all of their actions. However, in many practical scenarios, players' available actions might be restricted due to exogenous stochasticity. To model this setting, for a game $\mathcal{G}_{\mathrm{orig}}$ with action set $A_i$ for each player $i$, we introduce the corresponding Game with Stochastic Action Sets (GSAS) which is parametrized by a probability distribution over the players' set of possible action subsets $\mathcal{S}_i\subseteq 2^{A_i}\setminus\{\varnothing\}$. In a GSAS, players' strategies and Nash equilibria (NE) admit prohibitively large representations, and existing algorithms for NE computation scale poorly. Under the assumption that action availabilities are independent between players, we show that NE in two-player zero-sum (2p0s) GSAS can be approximately represented by a compact vector of size $\vert A_i\vert$, overcoming the naïve exponential-sized representation. Computationally, we introduce an algorithm that minimizes ranking regret, converging to NE with high probability in 2p0s-GSAS with rate $O(\sqrt{\vert A_i\vert\log\vert A_i\vert/T})$ for time horizon $T$. Finally, using the iterates of our algorithm, we develop a stochastic approximation procedure to recover compactly represented NE.

cs.GT↗

SeedFlood: A Step Toward Scalable Decentralized Fine-Tuning of LLMs

This work presents SeedFlood, a new approach to decentralized LLM fine-tuning designed to scale across large models, large collaborations, and complex network topologies while achieving global consensus with negligible communication overhead. Traditional methods suffer from high communication costs that grow with model size, while information decay over network hops renders global consensus inefficient. SeedFlood takes a significant departure from these practices by exploiting the seed-reconstructible structure of zeroth-order gradients and effectively making the messages to transmit near-zero in size, allowing them to be flooded to every client in the network, and thereby enhancing scalability of decentralized training. Consequently, SeedFlood enables training in regimes previously considered impractical, such as billion-parameter scale models or distributed across hundred of clients. Our experiments on decentralized LLM fine-tuning demonstrate that SeedFlood consistently outperforms the standard zeroth-order baselines in both communication efficiency and generalization performance, and even achieves results comparable to first-order gossip-based methods in large-scale settings, while requiring orders-of-magnitude less communication cost. We also provide theoretical analysis to formalize that SeedFlood avoids topology-dependent consensus terms in the convergence bound while retaining the acceleration enabled by increased client participation.

cs.LG↗

Capabilities Ain't All You Need: Measuring Propensities in AI

AI evaluation has primarily focused on measuring capabilities, with formal approaches inspired from Item Response Theory (IRT) being increasingly applied. Yet propensities - the tendencies of models to exhibit particular behaviours - play a central role in determining both performance and safety outcomes. However, traditional IRT describes a model's success on a task as a monotonic function of model capabilities and task demands, an approach unsuited to propensities, where both excess and deficiency can be problematic. Here, we introduce the first formal framework for measuring AI propensities by using a bilogistic formulation for model success, which attributes high success probability when the model's propensity is within an "ideal band". Further, we estimate the limits of the ideal band using LLMs equipped with newly developed task-agnostic rubrics. Applying our framework to six families of LLM models whose propensities are incited in either direction, we find that we can measure how much the propensity is shifted and what effect this has on the tasks. Critically, propensities estimated using one benchmark successfully predict behaviour on held-out tasks. Moreover, we obtain stronger predictive power when combining propensities and capabilities than either separately. More broadly, our framework showcases how rigorous propensity measurements can be conducted and how it yields gains over solely using capability evaluations to predict AI behaviour.

cs.LG↗

JEPA-Bisim: Learning Robust Visual Representations for Planning with Joint-Embedding Predictive World Models

World models learned from high-dimensional visual observations allow agents to make decisions and plan directly in latent space, avoiding pixel-level reconstruction. However, recent latent predictive architectures (JEPAs), including the DINO world model (DINO-WM), display a degradation in test time robustness due to their sensitivity to ``slow features". These include visual variations such as background changes and distractors that are irrelevant to the task being solved. We address this limitation by augmenting the predictive objective with a bisimulation encoder that enforces control-relevant state equivalence, mapping states with similar transition dynamics to nearby latent states while limiting contributions from slow features. We evaluate our model on a navigation task (PointMaze) and on a manipulation task (PushT) under different test-time background changes and visual distractors. Across all benchmarks, our model consistently improves robustness to slow features while operating in a reduced latent space, up to $10\times$ smaller than that of DINO-WM. Moreover, our model is agnostic to the choice of pre-trained visual encoder and maintains robustness when paired with DINOv2, SimDINOv2, and iBOT features.

cs.LG↗

Quantum Hamiltonian-Based Generative Modeling of Single-Cell Transcriptomics for Gene Regulatory Network Inference

We introduce a novel quantum Hamiltonian-based gene expression model (QHGM), a generative framework for modeling pseudotime-ordered single-cell gene expression data. In QHGM, gene interactions are encoded via a parameterized Hamiltonian, and the outcomes of quantum measurements provide a discrete representation of the gene expression profile. To learn the Hamiltonian parameters and infer gene regulatory networks (GRNs), we develop a scalable variational quantum algorithm for network inference (VQ-Net) based on empirical risk minimization. We derive finite-sample recovery guarantees for accurate parameter estimation, demonstrating polynomial scaling with the number of genes. Experiments on synthetic data demonstrate accurate GRN recovery, with VQ-NET achieving over 25% improvement in edge recovery and over 50% improvement in parameter-sign recovery compared to state-of-the-art classical methods. We further apply the framework to glioblastoma scRNA-seq data, where it identifies biologically plausible regulatory interactions associated with cancer progression, highlighting the potential of quantum-like modeling beyond classical probabilistic frameworks.

quant-ph↗

InterviewSim: A Scalable Framework for Interview-Grounded Personality Simulation

Simulating real personalities with large language models requires grounding generation in authentic personal data. Existing evaluation approaches rely on demographic surveys, personality questionnaires, or short AI-led interviews as proxies, but lack direct assessment against what individuals actually said. We address this gap with an interview-grounded evaluation framework for personality simulation at a large scale. We extract over 671,000 question-answer pairs from 23,000 verified interview transcripts across 1,000 public personalities, each with an average of 11.5 hours of interview content. We propose a multi-dimensional evaluation framework with four complementary metrics measuring content similarity, factual consistency, personality alignment, and factual knowledge retention. Through systematic comparison, we find that interview grounding yields consistent gains in content alignment and exact-match factual recall over biographical profiles and parametric prompting. We further find complementary strengths: retrieval-augmented methods tend to preserve personality alignment, while larger chronological contexts generally reduce contradictions and improve factual recall. Our evaluation framework enables principled method selection based on application requirements, and our empirical findings provide actionable insights for advancing personality simulation research.

cs.CL↗

Characterization-free classification and identification of the environment between two quantum players

Identifying the causal structure of quantum channels is essential for verifying quantum networks and certifying quantum resources. We introduce a characterization-free protocol enabling two isolated players, Alice and Bob, to identify the definite-order strategy adopted by an unknown environment mediating their channels. Without assuming knowledge of their devices or the environment, the players infer the causal order solely from input-output statistics by testing Markovian conditions that we prove are necessary and sufficient for each strategy class. Remarkably, we prove that, under an explicit generic-sampling condition, a randomly selected binary measure-and-prepare setting retains exact-distribution identifiability with probability one. In the optical experiment, we use a reduced-randomness construction in which several preparation states are kept fixed. Nevertheless, the Markov-condition-based procedure yields the expected causal-order and memory-presence classification for every tested process realization and setting. This observation suggests that the randomization assumptions of the general theorem may be relaxed. Our results provide an operational framework for causal inference in quantum networks.

quant-ph↗

AuditBench: Evaluating Alignment Auditing Techniques on Models with Hidden Behaviors

We introduce AuditBench, an alignment auditing benchmark. AuditBench consists of 56 language models with implanted hidden behaviors. Each model has one of 14 concerning behaviors--such as sycophantic deference, opposition to AI regulation, or secret geopolitical loyalties--which it does not confess to when directly asked. AuditBench models are highly diverse--some are subtle, while others are overt, and we use varying training techniques both for implanting behaviors and training models not to confess. To demonstrate AuditBench's utility, we develop an investigator agent that autonomously employs a configurable set of auditing tools. By measuring investigator agent success using different tools, we can evaluate their efficacy. Notably, we observe a tool-to-agent gap, where tools that perform well in standalone non-agentic evaluations fail to translate into improved performance when used with our investigator agent. We find that our most effective tools involve scaffolded calls to auxiliary models that generate diverse prompts for the target. White-box interpretability tools can be helpful, but the agent performs best with black-box tools. We also find that audit success varies greatly across training techniques: models trained on synthetic documents are easier to audit than models trained on demonstrations, with better adversarial training further increasing auditing difficulty. We release our models, agent, and evaluation framework to support future quantitative, iterative science on alignment auditing.

cs.CL↗

Starbursts hiding in the main sequence: a pathway toward quenching?

Star-forming galaxies spend most of their lifetimes on the star-forming main sequence, which establishes a tight empirical and statistical relation between stellar mass and star-formation rate. Occasional episodes of rapid star formation can push them temporarily above this sequence, turning them into starbursts. Yet some galaxies display starburst-like traits -- rapid, dense, and compact star formation -- while still remaining within the scatter of the main sequence. These "starbursts in the main sequence" (SBMSs) reveal the complexity and diversity of star formation modes, making them crucial for understanding how galaxies evolve and transition between different regimes. In this paper, we identify SBMSs in the cosmological simulation NewHorizon and follow their evolution across time to uncover their physical origins and the role of this special regime in shaping galaxy evolution. We explain the existence of SBMSs by a comparatively earlier assembly of their stellar mass, driven in particular by more frequent and repeated mergers as the other galaxies, as well as exceptionally productive starburst events triggered by these interactions. As a result, this regime appears preferentially -- though not exclusively -- in the most massive galaxies. The SBMS behavior is not continuous within individual galaxies but instead arises intermittently as a short-lived (~ 30 Myr) evolutionary mode. Nevertheless, such SBMS episodes exist throughout cosmic time across the galaxy population... [abridged]

astro-ph.GA↗

ProtoDCS: Towards Robust and Efficient Open-Set Test-Time Adaptation for Vision-Language Models

Large-scale Vision-Language Models (VLMs) exhibit strong zero-shot recognition, yet their real-world deployment is challenged by distribution shifts. While Test-Time Adaptation (TTA) can mitigate this, existing VLM-based TTA methods operate under a closed-set assumption, failing in open-set scenarios where test streams contain both covariate-shifted in-distribution (csID) and out-of-distribution (csOOD) data. This leads to a critical difficulty: the model must discriminate unknown csOOD samples to avoid interference while simultaneously adapting to known csID classes for accuracy. Current open-set TTA (OSTTA) methods rely on hard thresholds for separation and entropy minimization for adaptation. These strategies are brittle, often misclassifying ambiguous csOOD samples and inducing overconfident predictions, and their parameter-update mechanism is computationally prohibitive for VLMs. To address these limitations, we propose Prototype-based Double-Check Separation (ProtoDCS), a robust framework for OSTTA that effectively separates csID and csOOD samples, enabling safe and efficient adaptation of VLMs to csID data. Our main contributions are: (1) a novel double-check separation mechanism employing probabilistic Gaussian Mixture Model (GMM) verification to replace brittle thresholding; and (2) an evidence-driven adaptation strategy utilizing uncertainty-aware loss and efficient prototype-level updates, mitigating overconfidence and reducing computational overhead. Extensive experiments on CIFAR-10/100-C and Tiny-ImageNet-C demonstrate that ProtoDCS achieves state-of-the-art performance, significantly boosting both known-class accuracy and OOD detection metrics. Code will be available at https://github.com/O-YangF/ProtoDCS.

cs.CV↗

Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data

How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article presents a comparison between a human-respondent survey of 420 Silicon Valley coders and developers and synthetic survey data designed to simulate real survey takers generated by five leading Generative AI Large Language Models: ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro plus Claude CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, and DeepSeek 3.2. Our findings reveal that while AI agents produced technically plausible results that lean more towards replicability and harmonization than assumed, none were able to capture the counterintuitive insights that made the human survey valuable. Moreover, deviations grouped together for all models, leaving the real data as the outlier. Our key finding is that while leading LLMs are increasingly being used to scale, replicate and replace human survey responses in research, these advances only show an increased capacity to parrot conventional wisdom in harmony with each other rather than revealing novel findings. If synthetic respondents are used in future research, we need more replicable validation protocols and reporting standards for when and where synthetic survey data can be used responsibly, a gap that this paper fills. Our results suggest that synthetic survey responses cannot meaningfully model real human social beliefs within organizations, particularly in contexts lacking previously documented evidence. We conclude that synthetic survey-based research should be cast not as a substitute for rigorous survey methods, but as an increasingly reliable pre- or post-fieldwork instrument for identifying societal assumptions, conventional wisdoms, and other expectations about research populations.

cs.CY↗

MergeDJD: A Fast Constructive Algorithm with Piece Merging for the Two-Dimensional Irregular Bin Packing Problem

The two-dimensional irregular bin packing problem (2DIBPP) aims to pack a given set of irregular polygons, referred to as pieces, into fixed-size rectangular bins without overlap, while maximizing bin utilization. Although numerous metaheuristic algorithms have been proposed for the 2DIBPP, many industrial applications favor simpler constructive heuristics due to their deterministic behavior and low computational overhead. Among such methods, the DJD algorithm proposed by L'opez-Camacho et al. is one of the most competitive constructive heuristics for the 2DIBPP. However, DJD is less effective for cutting instances, in which many pieces can be seamlessly combined into larger polygons. To address the issue, we propose MergeDJD, a novel constructive algorithm that integrates and extends the DJD framework. MergeDJD first preprocesses the instance by iteratively identifying groups of pieces that can be combined into larger and more regular piece. It then employs an improved version of DJD, in which the placement strategy is enhanced to better handle non-convex and combined shapes, to pack all resulting pieces into bins. Computational experiments on 1,089 well-known benchmark instances show that MergeDJD consistently outperforms DJD on 1,083 instances while maintaining short runtimes. Notably, MergeDJD attains new best known values on 515 instances. Ablation studies further confirm the effectiveness of the proposed components. To facilitate reproducibility and future research, we have open-sourced the complete implementation and provided interfaces for visualizing packing results.

cs.CG↗

From Literature to Hypotheses: An AI Co-Scientist System for Biomarker-Guided Drug Combination Hypothesis Generation

The rapid growth of biomedical evidence makes it difficult to translate biomarker mechanisms into actionable drug combination hypotheses. We present CoDHy, an interactive AI co-scientist for biomarker-guided hypothesis generation in oncology. CoDHy constructs task-specific knowledge graphs from curated databases and biomedical literature, then combines graph embeddings with agent-based reasoning to generate, validate, and rank evidence-grounded drug combinations. Through a web interface, researchers specify the biomarker, cancer context, and literature scope; inspect supporting evidence and intermediate results; and iteratively refine the generated hypotheses. The demonstration presents CoDHy's end-to-end workflow and shows how researchers can interactively explore and compare mechanistically supported drug combinations while remaining in control of hypothesis prioritization.

cs.CL↗