Search arXiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 253 records · Page 14Linked to original sources

Man and machine: artificial intelligence and judicial decision making

The integration of artificial intelligence (AI) into judicial decision making -- particularly in pretrial, sentencing, and parole contexts -- has generated a substantial and rapidly growing literature. Across computer science, economics, law, criminology, and psychology, researchers have examined the reliability, fairness, and real-world effects of AI-assisted decision making. Yet this literature remains fragmented, and differences in assumptions, concepts, and research priorities make it difficult to assess what is actually known. Using criminal justice risk assessment as a focal case, this article makes two contributions. First, we develop a conceptual framework that distinguishes and relates three central questions: (1) the predictive validity of automated risk assessment tools; (2) how algorithmic risk assessments compare with human predictions (AI-versus-Human); and (3) how algorithmic recommendations affect judges' decisions (AI-plus-Human). Second, we use this framework to synthesize the empirical evidence addressing each of these questions. Our review identifies important limitations in existing research on predictive validity, as well as substantial gaps in understanding how judges respond to AI advice and how those responses vary across individuals and decision-making environments. The available evidence suggests that AI decision aids have, so far, had at most modest effects on pretrial and sentencing decisions. We conclude that further research is needed to understand how judges make decisions in noisy informational environments and under what conditions AI tools can produce meaningful improvements in judicial decision making.

cs.AI↗

Neither Here Nor There: Cross-Lingual Representation Dynamics of Code-Mixed Text in Multilingual Encoders

Multilingual encoder-based language models are widely used for code-mixed analysis, yet their internal representations of code-mixed inputs -- and their relationship to the constituent languages -- remain poorly understood. Using Hindi-English as a case study, we construct a unified trilingual corpus of parallel English, Hindi (Devanagari), and Romanized code-mixed sentences. We then probe cross-lingual representation alignment in standard multilingual encoders and their code-mix-adapted variants using CKA, token-level saliency, and entropy-based uncertainty analysis. We find that while standard models align English and Hindi well, code-mixed inputs remain loosely connected to either language -- and that continued pre-training on code-mixed data improves English-code-mixed alignment at the cost of English-Hindi alignment. Interpretability analyses further reveal a clear asymmetry: models process code-mixed text through an English-dominant semantic subspace, while native-script Hindi provides complementary signals that reduce representational uncertainty. Motivated by these findings, we introduce a trilingual post-training alignment objective that brings code-mixed representations closer to both constituent languages simultaneously, yielding more balanced cross-lingual alignment and downstream gains on sentiment analysis and hate speech detection -- showing that grounding code-mixed representations in their constituent languages meaningfully helps cross-lingual understanding. Code is available at: https://github.com/debajyotimaz/tri_align_EMNLP_2026.

cs.CL↗

EgoForge: Goal-Directed Egocentric World Simulator

Generative world models have shown promise for simulating dynamic environments, yet egocentric video remains challenging due to rapid viewpoint changes, frequent hand-object interactions, and goal-directed procedures whose evolution depends on latent human intent. Existing approaches either focus on hand-centric instructional synthesis with limited scene evolution, perform static view translation without modeling action dynamics, or rely on dense supervision, such as camera trajectories, long video prefixes, and synchronized multi-camera capture. In this work, we introduce EgoForge, an egocentric goal-directed world simulator that generates coherent, first-person video rollouts from minimal static inputs: a single egocentric image, a high-level instruction, and an optional auxiliary exocentric view. To improve intent alignment and temporal coherence, we introduce GRAFT, a trajectory-level diffusion refinement method that uses positive and negative rollout distributions, derived from goal, temporal, scene-consistency, and perceptual rewards, to steer the diffusion velocity field toward coherent, goal-complete egocentric simulations. Extensive experiments show EgoForge achieves consistent gains in semantic alignment, geometric stability, and motion fidelity over strong baselines, and performs robustly in real-world smart-glasses experiments.

cs.CV↗

On the full set of unitarizable supermodules over $\mathfrak{sl}(m\vert n)$

We classify all simple unitarizable supermodules over special linear Lie superalgebras using the algebraic Dirac operator introduced by Huang and Pandžić and the associated Dirac inequalities. The same argument treats finite-dimensional and infinite-dimensional supermodules without requiring explicit realizations or complete branching rules.

math.RT↗

Virtual absorption modes of Schwarzschild-de Sitter spacetimes in semi-open systems

We present a study of virtual absorption modes (VAMs) in Schwarzschild-de Sitter (SdS) spacetime under semi-open boundary conditions, where the VAMs correspond to total transmission modes (TTMs) with the reflection amplitude being vanishing. Our numerical analysis reveals that as the reflectivity $|\mathcal{K}|$ decreases, the imaginary parts of VAM spectra systematically decrease, with each overtone exhibiting a critical reflectivity at which $\text{Im}(ω_{\text{VAM}})=0$. Using simulations based on spectral collocation methods, we demonstrate that excitation precisely at a VAM spectrum leads to a virtual absorption process. These results establish VAMs as the spectral signatures of virtual absorption processes for exotic compact objects (ECOs).

gr-qc↗

The NCS-Model: A seismic foundation model trained on the Norwegian repository of public data

We present the NCS-models, a family of seismic foundation models pretrained on a large curated share of full-stack seismic cubes from the Norwegian Continental Shelf (NCS) available through the public DISKOS database. The model weights are open-sourced for the wider geoscience community. Foundation models trained with large-scale self-supervision are emerging as a promising basis for automatic seismic interpretation. However, most existing seismic models rely on limited or proprietary datasets, and it remains unclear how well natural-image foundation models transfer to seismic data. Our goals are to develop basin-scale seismic foundation models, provide practical recipes for scalable 3D training, and compare the downstream performance of practical configurations trained under similar pretraining budgets. Using masked autoencoders with Vision Transformer backbones, we pretrain models on a DISKOS-derived corpus of 3D time- and depth-migrated seismic volumes. The NCS-model variants use 2D, 2.5D multi-view, and 3D tokenization and each variant is pretrained using approximately 250 GPU-hours on the same hardware. We evaluate the models using frozen backbones together with k-nearest neighbors and linear probing on NCS interpretation benchmarks and one out-of-basin benchmark. Baselines include an ImageNet-pretrained MAE, a frontier vision foundation model, and a globally pretrained seismic foundation model. Following large-scale pretraining on the curated DISKOS corpus, NCS-2.5D achieves the highest average performance among all evaluated models. The resulting embeddings additionally support similarity search for interactive interpretation.

physics.geo-ph↗

Casewise and Cellwise Robust Tensor-on-Tensor Regression

Tensor-on-tensor regression is an important tool for the analysis of tensor data, aiming to predict a set of response tensors from a corresponding set of predictor tensors. However, standard tensor-on-tensor regression is sensitive to outliers, which may be present in both the response and the predictor. It can be affected by casewise outliers, which are observations that deviate from the bulk of the data, as well as by cellwise outliers, which are individual anomalous cells within the tensors. The latter are particularly common due to the typically large number of cells in tensor data. This paper introduces a novel robust tensor-on-tensor regression method, named ROTOT, that can handle both types of outliers simultaneously, and can cope with missing values as well. This method uses a single loss function to reduce the influence of both casewise and cellwise outliers in the response. The outliers in the predictor are handled using a robust Multilinear Principal Component Analysis method. Graphical diagnostic tools are also proposed to visualize the different types of detected outliers. The performance of ROTOT is evaluated through extensive simulations and further illustrated using the Labeled Faces in the Wild dataset, where ROTOT is applied to predict facial attributes.

stat.ME↗

Image-Domain Poisson-Perturbation Robustness of NCCT Slice Classification

Image-domain Poisson perturbation may alter normalized NCCT appearance and downstream models. We evaluated classification of ischemic core/penumbra-bearing non-contrast CT (NCCT) slices under five simulated settings. The source cohort was CPAISD (112 hyperacute ischemic stroke patients); the test partition contained 10 patients and 809 slices. First, a fixed-checkpoint audit compared direct ResNet-18 classification (P1) with residual U-Net denoising followed by the classifier (P2). P1 average precision (AP) ranged from 0.694 to 0.901, whereas P2 AP ranged from 0.509 to 0.797 (substantially lower at settings 10-40). The fixed 0.5 threshold had 0-12.8% sensitivity for P1 and 0% for P2. Second, a prospectively locked de novo experiment compared direct noisy classification (DNC), joint denoising-classification (JDC-0), and the same joint model with privileged training-only lesion-boundary supervision (JDC-B). Across-setting mean AP was 0.861 +/- 0.028 for DNC, 0.840 +/- 0.038 for JDC-0, and 0.856 +/- 0.031 for JDC-B. Hierarchical paired-bootstrap differences were -0.021 (95% CI -0.063 to 0.019) for JDC-0 minus DNC, 0.016 (-0.019 to 0.054) for JDC-B minus JDC-0, and -0.005 (-0.041 to 0.026) for JDC-B minus DNC; none excluded zero. A frozen stress test on the 52-patient AISD partition also showed limited transportability. Thus, ordinary joint training did not demonstrate a classification benefit, and the boundary term recovered part of its point-estimate loss without a statistically supported advantage. Image fidelity, ranking, calibration, and clinical utility must be evaluated separately. This study does not validate acquired low-dose, portable, or cone-beam CT, nor patient-level stroke diagnosis.

cs.CV↗

Real-time Appearance-based Gaze Estimation for Open Domains

Appearance-based gaze estimation (AGE) has achieved remarkable performance in constrained settings, yet we reveal a significant generalization gap where existing AGE models often fail in practical, unconstrained scenarios, particularly those involving facial wearables and poor lighting conditions. We attribute this failure to two core factors: limited image diversity and inconsistent label fidelity across different datasets, especially along the pitch axis. To address these, we propose a robust AGE framework that enhances generalization without requiring additional human-annotated data. First, we expand the image manifold via an ensemble of augmentation techniques, including synthesis of eyeglasses, masks, and varied lighting. Second, to mitigate the impact of anisotropic inter-dataset label deviation, we reformulate gaze regression as a multi-task learning problem, incorporating multi-view supervised contrastive (SupCon) learning, discretized label classification, and eye-region segmentation as auxiliary objectives. To rigorously validate our approach, we curate new benchmark datasets designed to evaluate gaze robustness under challenging conditions, a dimension largely overlooked by existing evaluation protocols. Our MobileNet-based lightweight model achieves generalization performance competitive with the state-of-the-art (SOTA) UniGaze-H, while utilizing less than 1\% of its parameters, enabling high-fidelity, real-time gaze tracking on mobile devices.

cs.CV↗

Operator Norm Bounds for Multi-leg Matrix Tensors and Applications to Random Matrix Theory

We study the extremal values of multi-leg traces of matrix tensors under operator norm constraints. A graphical representation gives upper and lower bounds expressed as powers of the common tensor-factor dimension. The lower bounds are attained by matrices that permute tensor factors and are exact within this family. We prove the bounds first for two tensor factors and then for an arbitrary number. Leaving some indices uncontracted produces matrices for which we obtain both moment and operator norm bounds; a single choice of coefficient matrices attains the lower bounds for every positive integer moment. We also obtain exact scalar and operator norm maxima in special cases with sufficiently many tensor factors sharing the same cyclic contraction. As an application, the operator norm bounds yield a uniform comparison between Ginibre products and their free circular counterparts with growing matrix coefficients.

math.OA↗

Who Wrote the Book? Detecting and Attributing LLM Ghostwriters

In this paper, we introduce GhostWriteBench, a dataset for LLM authorship attribution. It comprises long-form texts (50K+ words per book) generated by frontier LLMs, and is designed to test generalisation across multiple out-of-distribution (OOD) dimensions, including domain and unseen LLM author. We also propose TRACE -- a novel fingerprinting method that is interpretable and lightweight -- that works for both open- and closed-source models. TRACE creates the fingerprint by capturing token-level transition patterns (e.g., word rank) estimated by another lightweight language model. Experiments on GhostWriteBench demonstrate that TRACE achieves state-of-the-art performance, remains robust in OOD settings, and works well in limited training data scenarios.

cs.CL↗

Central Limit Theorems for Outcome Records in Disordered Quantum Trajectories

We prove annealed functional central limit theorems for finite pattern counts in the measurement record of discrete-time quantum trajectories, with the instrument applied at each step determined by an invertible, probability-preserving base dynamical system. When the base is ergodic, under summable strong-mixing coefficients of the instrument process and a summable uniform annealed trace-norm forgetting rate for the associated non-selective channel cocycle, we establish a joint functional CLT for bounded vector-valued functions of finite outcome blocks under the annealed law determined by the dynamically stationary state. We then extend this limit to every measurable random initial state, yielding a universal functional CLT with unchanged stationary centering and asymptotic covariance. We also provide practical sufficient criteria ensuring the existence and uniqueness of the dynamically stationary state and the required annealed trace-norm forgetting. We illustrate the results through a broad family of examples, including disordered walk-type models generated by finite group actions, measurement followed by preparation, and instruments with reset components. The results apply to general disordered quantum instruments and are not restricted to the perfect-measurement regime; they complement the law of large numbers established by Ekblad, Moreno-Nadales, and Pathirana (2026) for the same disordered setting and provide a disordered counterpart of the homogeneous CLT of Attal, Guillotin-Plantard, and Sabot (2014).

math-ph↗

Beyond Idealized Patients: Evaluating LLMs under Challenging Patient Behaviors in Medical Consultations

Large language models (LLMs) are increasingly used for medical consultation and health information support, where safety depends not only on medical knowledge but also on robust responses to unclear, inconsistent, or misleading patient input. However, most existing medical LLM evaluations assume idealized and well-posed patient questions, limiting their realism. We study challenging patient behaviors that commonly arise in real medical consultations and complicate safe clinical reasoning. We define four clinically grounded categories of such behaviors: information contradiction, factual inaccuracy, self-diagnosis, and care resistance. For each behavior, we specify concrete failure criteria that capture unsafe responses. Building on four existing medical dialogue datasets, we introduce CPB-Bench (Challenging Patient Behaviors Benchmark), a bilingual (English and Chinese) benchmark of multi-turn dialogues annotated for these behaviors. We find that although models perform well overall, they exhibit consistent behavior-specific failures, especially when handling contradictory or medically implausible patient information. We further evaluate four intervention strategies and find inconsistent improvements, with some interventions introducing unnecessary corrections.

cs.CL↗

Revisiting Strategic Protectionism in Network Markets

This paper studies the welfare effects of an industrial policy targeting a sector with network externalities in a two-country model with strategic trade and R&D investment. When externalities are weak or the goods are close substitutes, the business-stealing effect dissipates more surplus than the externality creates. Under sufficiently strong externalities and weak substitutability or complementarity of the goods, industrial policy competition can make both countries simultaneously better off compared to the laissez-faire outcome because of the mutual business-enhancement effect. The case is stronger for product innovation than for process innovation, as the former directly affects the demand and triggers a stronger network effect than the latter which operates indirectly through the industry supply. Thus, the network externalities create an opportunity for win-win industrial policies, but its realisation depends on the market structure and the nature of innovation.

econ.TH↗

A Taxonomy of Programming Languages for Code Generation

The world's 7,000+ languages vary widely in the availability of resources for NLP, motivating efforts to systematically categorize them by their degree of resourcefulness (Joshi et al., 2020). A similar disparity exists among programming languages (PLs); however, no resource-tier taxonomy has been established for code. As large language models (LLMs) grow increasingly capable of generating code, such a taxonomy becomes essential. To fill this gap, we present the first reproducible PL resource classification, grouping 646 languages into four tiers. We show that only 1.9% of languages (Tier 3, High) account for 74.6% of all tokens in seven major corpora, while 71.7% of languages (Tier 0, Scarce) contribute just 1.0%. Statistical analyses of within-tier inequality, dispersion, and distributional skew confirm that this imbalance is both extreme and systematic. Our results provide a principled framework for dataset curation and tier-aware evaluation of multilingual LLMs.

cs.CL↗

Hamiltonicity of inhomogeneous random graphs

We provide a complete characterization of those graphons $W$ for which the inhomogeneous random graph \(\G(n,W)\) is asymptotically almost surely Hamiltonian. The characterization involves three conditions. Two of them constitute the characterization of $\G(n,W)$ being a.a.s.\ connected, as was shown recently by Hladký and Viswanathan. The third condition captures a geometric obstacle which prevents $\G(n,W)$ from having perfect fractional matchings. A prominent feature of the positive direction of our proof is the use of weak* limits to translate the absence of the above geometric obstacle into favourable matching properties of a regularization of $\G(n,W)$.

math.CO↗

Detecting Multi-Agent Collusion Through Multi-Agent Interpretability

As LLM agents are increasingly deployed in multi-agent systems, they introduce risks of covert coordination that may evade standard forms of human oversight. While linear probes on model activations have shown promise for detecting deception in single-agent settings, collusion is inherently a multi-agent phenomenon, and the use of internal representations for detecting collusion between agents remains unexplored. We introduce NARCBench, a benchmark for evaluating collusion detection under environment distribution shift, and propose five probing techniques that aggregate per-agent deception scores to classify scenarios at the group level, evaluated across four open-weight models (Qwen3-32B, Llama-3.1-70B, DeepSeek-R1 32B, GPT-OSS-20B) and six probe architectures. We frame this as a distributed anomaly detection problem, identifying three collusion signatures that map onto distinct anomaly types and detection paradigms. Every model reaches 1.00 AUROC in-distribution; on our strongest model (Llama-3.1-70B), our five probing techniques achieve 0.73 to 0.93 AUROC when transferred zero-shot to structurally different multi-agent scenarios and 0.99 to 1.00 on a steganographic blackjack card-counting task, with detection performance scaling with model capability. We find that no single probing technique dominates across all collusion types, consistent with the framework's prediction that different anomaly types require different detection paradigms. This work takes a step toward multi-agent interpretability: extending white-box inspection from single models to multi-agent contexts, where detection requires aggregating signals across agents. These results suggest that model internals provide a complementary signal to text-level monitoring for detecting multi-agent collusion. Code and data available at https://github.com/aaronrose227/narcbench.

cs.AI↗

The Hitchhikers Guide to Rubric Quality Understanding and Enrichment

Rubrics distill notions of expert quality and measure agent performance. However, the quality of rubrics themselves have not been systematically measured and are often left to downstream performance.We import apparatuses from measurement theory built for exactly this: quantitative signals based on the rubric's content, and introduce the RubrIc-Failure Taxonomy (RIFT), of nine possible ways a rubric fails, organized under reliability and content validity. Every mode leaves a distinct signature. To show the signals track failure causally, we seed 720 corruptions, injecting each RIFT mode into clean rubrics at known severity levels. A linear probe over the signals identifies which mode was injected at $75.0\%$ accuracy, beating $56.7\%$ for a frontier model asked to name the failure directly. Surprisingly across GDPval and Terminal-Bench, 10 of 48 expert-authored rubrics weight their criteria backwards, putting more of the score on requirements an expert panel judged less essential. This means a response can fail what matters most and still be graded well. This paper serves as a comprehensive guide on how to understand failure modes in rubrics and create better versions using quality signals, causal experiments, and provides a taxonomy with its rules and examples.

cs.AI↗