Search arXivSearch

SEARCH · Search arXiv

Results for “physics.bio-ph”

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

925 records · Page 9Linked to original sources

Measurement Validity in LLM Cultural Alignment

Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments like the Inglehart-Welzel Cultural Map and drawing conclusions about which cultures a model resembles. While a model's answer to a value-laden questions may be interpreted as a cultural signal, it also carries sampling noise and, can be quite sensitive to question framing. In this paper, we separate survey responses, sampling noise and question framing for multiple LLMs. We decompose response variance from these models into variation across random seeds, prompt rewordings. We employ noise-to-signal ratio (NSR) to test whether a model's apparent cultural position is distinguishable from noise. When applied across a dozen models from four geographic origins, calibrated against 88 Integrated Values Survey countries, the answer is often no. NSR exceeds 1.0 on 49 of 117 valid model-question pairs (42%), reaching 5.56 in the worst case. Two models even refuse to answer sufficient number of survey questions outright. Our results corroborate previous findings that LLMs cluster toward Western, English-speaking cultural positions. However, what does not hold up in this study is the precision with which anyone can currently interpret a specific model's coordinates: prompt tone alone can shift a model by 2.4 map units, comparable to the distance between actual countries in the Inglehart-Welzel Cultural Map. These findings suggest that cultural attribution from LLM survey responses requires establishing the reliability of the underlying measurements before interpreting model coordinates as evidence of cultural representation.

physics.soc-ph

Generative Nested Sampling of Atomistic Thermodynamic Landscapes

Nested sampling (NS) resolves the thermodynamics of an atomistic system from a single simulation, but its practical reach is limited by the Markov-chain updates needed to decorrelate walkers within each likelihood-constrained ensemble. Flow-based NS has removed this bottleneck for gravitational-wave (GW) inference, yet its transfer to atomistic systems is not merely a change of application. Comparing a GW150914-like binary-black-hole likelihood with an eight-particle two-dimensional Lennard-Jones (LJ) system of comparable dimensionality, we show that the two landscapes differ fundamentally: atomistic multimodality is discrete and combinatorial, generated by particle permutations separated by hard collision walls, and its coordinate coupling is dense and collective, whereas the GW posterior exhibits smooth degeneracies and localized parameter coupling. Guided by this diagnosis, we introduce NS-Flows: a single conditional normalizing flow, conditioned on the NS energy bound and trained on a sliding window of recent live sets, that replaces MCMC by direct parallel draws corrected by importance-weighted rejection resampling. Live sets supply data self-consistently, allowing flow training without structured priors or a pre-existing dataset. For LJ disks in PBC, the algorithm reduces energy evaluations by over two orders of magnitude and wall-clock time by roughly one third, an advantage that becomes increasingly favorable as the cost of the potential grows. The flow's generation efficiency further acts as a physical diagnostic: it varies non-monotonically along the annealing trajectory, is lowest in the dense disordered regime, and is quantitatively captured by the constrained ensemble's internal mode complexity together with target drift across the training window, identifying liquid-like ensembles, rather than prior-target separation, as the hard case for current flow architectures.

cond-mat.stat-mech

La Agente Óptima: Towards Agentic Self-Driving Laboratories

Self-driving laboratories (SDLs) combine automated experimentation with adaptive decision-making to accelerate scientific discovery. Their operation nevertheless often depends on human specialists who translate scientific objectives into executable closed-loop campaigns. Specialists adjust them as data and operating conditions change. Here, we present La Agente Óptima, an agentic framework that constructs and supervises Bayesian optimization campaigns across computational and experimental systems while maintaining a persistent optimization state. By separating large language model (LLM) reasoning from executed campaigns, Óptima runs repetitive optimization loops consistently, returns control to the agent only when progress requires interpretation or campaign revision, and keeps every decision auditable. We evaluate Óptima across ablation studies, five digital discovery tasks, and two physical platforms. Throughout, Óptima maintained executable campaigns as both the scientific problem and execution environment evolved. In a closed-loop contact angle optimization campaign, Óptima identified and corrected a mid-run measurement failure, bringing the contact angle from 71.4 to 67.8 degrees, just above the 64-66 degree range. From this result, Óptima correctly inferred that the target was likely unattainable with the available reagents and recommended changing the formulation. In a five-day multi-objective flow-chemistry campaign, Óptima increased the yield from 30% to 59% over 23 experiments. Despite substantial inference costs, it cost less and used substantially less starting material than a human-directed campaign, while selecting a more mass-efficient operating point. These results show that LLM-based agents can make rigorous, long-running optimization campaigns accessible to domain scientists without specialist setup, expanding the scope of SDLs.

cs.AI

GadIR: A Spatial-Topology Preserving Compiler for Quantum Many-Body Systems Simulation

Simulating quantum many-body systems has been one of the most important applications of quantum computation. For simulation, the Hamiltonian of a physical system is compiled into quantum programs with native instructions for quantum hardware. In previous works, the Hamiltonian is represented as Pauli strings, then compiled and optimized based on the quantum circuit model. Such representation paradigm neglects the spatial topology of original physical models, which is vital information to reducing the overhead of compiling many-body systems Hamiltonians. To address such neglect, we introduce a spatial-topology preserving compiler for quantum many-body simulation. Using Pauli gadgets as the representations of the Hamiltonian, we introduce our intermediate representation -- GadIR, to preserve the spatial-topology information of original physical models. Our compiler frontend performs the group reduction algorithm based on Pauli gadget model, which is a hardware-independent optimization. Our compiler backend performs trotterization and scheduling on Pauli gadgets, then synthesizes the Pauli gadgets into hardware-native quantum programs. We evaluate our compiler on all the canonical quantum many-body system models, while achieving a significant reduction on compilation overhead regarding four major quantum architectures. Overall, our spatial-topology preserving IR exploits the compilation optimization space for quantum many-body systems Hamiltonian.

quant-ph

WolfSociety: Understanding Collective Risk from Harmful-Agent Scaling in Financial Agent Societies

Safety evaluations typically focus on individual agents, but interacting agents can spread harmful information and influence the environment in which later decisions are made. We study how collective failure changes with harmful-agent fraction and society size in a controlled financial agent society, where agents communicate over a social network and trade in a shared market. In the primary financial scenario, collective failure requires broad harmful diffusion together with severe price dislocation or liquidity stress. Across all tested society sizes, failure remains rare at low harmful fractions but rises sharply over a narrow range. As society size grows from N=100 to N=2000, the harmful fraction associated with a 50% failure probability decreases from 4.7% to 2.2%, while the corresponding number of harmful agents increases from approximately 5 to 44. In contrast, when the number of harmful agents is held fixed, their impact becomes weaker as the society grows. Controlled interventions further show that broader network reach shifts the collapse boundary toward lower harmful fractions, whereas stronger conformity alone has little effect. To characterize these effects, we introduce Agent Society Dynamics, a finite-size framework for relating harmful-agent fraction, society size, and interaction structure to collective failure. Overall, our results reveal a nonlinear, size-dependent collapse transition in financial agent societies, showing that collective failure depends not only on the prevalence of harmful agents but also on the size and interaction structure of the surrounding society. Code is available at https://github.com/SAIL-Research-Lab/WolfSociety.

physics.soc-ph

Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling

Global weather reanalyses and forecasts resolve the evolving atmospheric state on coarse grids, but site-specific applications require predictions at arbitrary locations where near-surface conditions also depend on unresolved terrain and land-surface properties. Existing probabilistic downscalers address this gap using hand-crafted topographic and surface descriptors. We ask instead whether Earth observation foundation models can provide transferable subgrid surface representations for probabilistic weather downscaling. We augment a convolutional conditional neural process (ConvCNP) that downscales coarse ERA5 reanalysis fields at ~25 km resolution with a learned local surface descriptor, obtained by compressing a patch of TESSERA embeddings at 10 m resolution. Although these embeddings summarize annual surface conditions, they improve downscaling by encoding persistent surface properties that capture a location's departure from the coarse-grid atmospheric state. Across five climatically diverse regions, the embedding improves point and probabilistic skill at stations held out in both space and time, overall improving CRPS skill by 11.5% for 2 m temperature and 6.2% for 10 m wind speed relative to a topography-only ConvCNP baseline. A hand-crafted descriptor incorporating richer surface information than topography alone captures comparable persistent subgrid signal but yields far smaller predictive gains than the learned embedding representation. These improvements persist when forecasts from the Aurora AI model replace ERA5 reanalysis fields and when predicting at newly deployed weather station networks. To our knowledge, this is the first evidence that long-timescale Earth observation embeddings can support short-timescale weather downscaling where subgrid departures are systematically structured by persistent surface properties.

cs.LG

Modelling infodemics on a global scale: A 30 countries study using epidemiological and social listening data

Infodemics represent a significant threat to public health, arising from complex interactions between online and offline phenomena. The continuous feedback loops between digital information ecosystems and real-world contingencies make infodemics particularly challenging to define operationally, measure, and eventually model in quantitative terms. This study aims to evaluate the effect of various epidemic-related variables on the dynamics of the COVID-19 infodemic, using a regression modeling framework applied to data from 30 countries across diverse income groups. We use World Health Organization (WHO) COVID-19 surveillance data on new cases and deaths, vaccination data from the Oxford COVID-19 Government Response Tracker, infodemic data (volume of public conversations and social media content) from the WHO EARS platform, and Google Trends data to represent information demand. Our findings show that new deaths are the strongest predictor of document production, and that the epidemic burden in neighboring countries exerts a greater influence on document production than domestic epidemic conditions. Building on these results, we propose a data-driven classification of country-level response that highlights country-specific discrepancies between the evolution of the infodemic and the epidemic. Further, an analysis of the temporal evolution of the relationship between the two phenomena quantifies the extent to which discussions surrounding vaccine rollouts may have shaped the development of the infodemic. Beyond underscoring the value of a holistic approach that integrates both online and offline dimensions, our results demonstrate that the evolution of infodemics and their relationship with epidemic variables can be closely monitored, even over short time windows.

cs.SI

Modeling of Mobility and Energy Policies in an Agent-Based Framework: Case Studies for Chicago Region in 2050

Metropolitan regions are simultaneously pursuing several interventions to improve mobility, accessibility, and energy efficiency, necessitating integrated tools to understand how these policies interact to affect travel behavior, energy use, and infrastructure needs. This paper evaluates the combined impacts of electrification, freight demand management, road pricing, parking reform, and transit expansion on the Chicago metropolitan transportation system in 2050, using a business-as-usual (BAU) scenario as the baseline. We employ POLARIS, a large-scale agent-based modeling framework calibrated to 2019 conditions, to simulate nine policy scenarios for the seven-county northeastern Illinois region. The framework co-simulates activity-based passenger demand, endogenous freight generation, multimodal traffic assignment, and transit operations, with charging infrastructure and freight operations optimized for each case. Our findings reveal that under the high electrification scenario, total fuel mass declines by 68% while total charging energy increases by approximately 4-8x from BAU, resulting in a peak power demand near 4 GW concentrated in the urban core. Furthermore, freight management policies reduce freight VMT by increasing trip frequency but shortening distances, smart road pricing most effectively reduces auto VMT, and transit expansion boosts ridership by 18% relative to BAU. By presenting the first integrated, agent-based scenario framework for Chicago that jointly evaluates these interventions, this study provides actionable insights for regional transportation planning, grid infrastructure investment, and emissions reduction, highlighting the value of targeted charger upgrades and coordinated policy bundles.

physics.soc-ph

AI-Augmented Inquiry and Regulation in Hybrid Systems: A Control Allocation Architecture for Preserving Epistemic Agency in Hybrid Human-AI Cognition

Generative artificial intelligence (genAI) systems are increasingly integral to epistemic processes such as hypothesis generation, explanation construction, and decision-making. Although they reliably enhance performance, emerging evidence reveals a metacognitive dilemma: as external generative capacity increases, internal monitoring, calibration, and cognitive engagement may decline. This reflects a redistribution of cognitive control within distributed human-AI systems that cannot be explained by automation bias or reliance on algorithms alone. We propose the AIRIS (AI-Augmented Inquiry and Regulation in Hybrid Systems) framework to analyze this dilemma and specify where regulatory intervention can counteract it. AIRIS is a multi-level control allocation architecture specifying the conditions under which epistemic agency can be preserved in hybrid generative systems. Drawing on distributed cognition, cognitive load theory, multimedia learning, and self-regulated learning, it identifies seven interacting mechanisms through which hybrid cognition may become destabilized, from delegation and calibration drift to motivational-affective drift. Five regulatory operators (Anticipate, Interrogate, Reflect, Integrate, and Synthesize) target internal generative engagement at points of emerging instability. The architecture does not itself improve learning; it specifies what must remain in place for genAI-supported work to sustain understanding, whether through instructional design, teacher guidance, or learners' own regulation. We derive testable propositions concerning the seven mechanisms and the five operators, reframing AI augmentation as a problem of control allocation in distributed generative systems. Beyond theory, AIRIS offers a research agenda, a design framework for genAI-integrated learning environments, and a conceptual toolkit for the governance of hybrid human-AI cognition.

physics.ed-ph

A Reproducible Method for Mapping Electricity Transmission Infrastructure for Space Weather Risk Assessment

Space weather risk assessment is constrained by the lack of available asset information needed to model geomagnetically induced currents (GICs) in electricity transmission infrastructure. We propose a systematic method that enables risk analysts to collect their own open-source substation data. Using a web browser platform for annotation, we convert OpenStreetMap (OSM) substation locations into high-resolution, component-level mappings of electricity transmission assets. We convert an initial 1,313 high-voltage (>=230 kV) substations to 52,273 components using low-altitude, satellite, and Street View imagery accessed through Google Earth, identifying 7,949 transformers. Compared to the OSM baseline, this approach provides detailed insights on voltage levels and substation configurations. We then construct a geospatial GIC network for the Tennessee Valley Authority (TVA) region, comparing May 2024 results with the University of Illinois Urbana-Champaign 150-bus (UIUC150) synthetic network and with measured ground GICs at 13 monitoring devices. The transformer types at unannotated substations and the grounding resistances are unknown, so we sample both across a Monte Carlo ensemble. This gives a median TVA 95th-percentile peak ground GIC of 29.1 A, with a 90 percent confidence interval of 23.9-36.1 A. The UIUC150 network yields a 95th-percentile peak ground GIC of 35.8 A under the same forcing, falling within this interval, and the modeled time series broadly capture the temporal morphology of the geomagnetic storm at the monitoring sites. This method shows promise for spatially explicit, screening-level GIC assessment without requiring access to operator data.

physics.geo-ph

CoDiNG -- Naming Game with Continuous Latent Opinions of Individual Agents

Understanding the mechanisms behind opinion formation is crucial for gaining insight into the processes that shape the spread of political beliefs, cultural attitudes, consumer choices, and social movements in society. This work introduces a realistic model of opinion dynamics that captures the intricacies of real-world opinion dynamics by synthesizing principles from cognitive science. The proposed model is a hybrid continuous-discrete extension of the well-known Naming Game opinion model. The continuous layer captures the strength of each opinion through reinforcement and forgetting in the human brain, akin to memory imprints. The discrete layer allows for converting intrinsic continuous opinion into a discrete form, which often occurs when we publicly verbalize our opinions. We evaluated our model on longitudinal data combining real communication events with repeated surveys of the same individuals, comparing it against the Naming Game, the hybrid SJBO model, and four simple baselines at the population and the individual level. Unlike rigid baselines and the classic Naming Game, which inherently capture only a single aspect, hybrid models can be tuned to model either individual-level opinions or aggregate opinion dynamics. However, this flexibility comes with a strict trade-off, as they cannot accurately reproduce both simultaneously. Out of the six analysed topics, our model exceeds or matches SJBO, showing that reinforcement and forgetting grounded in cognition contribute to explaining opinion dynamics. Additionally, in our empirical data, individuals change their opinions while the aggregate distribution stays almost stationary, so a model reproducing no dynamics can still score well. This observation indicates that evaluating opinion models must be multidimensional and rely on more than one metric.

cs.SI

CARDIO-Affect: A Hamiltonian-Variability Framework for Spatio-Temporal Emotional Pattern Recognition with Manifold-Based Individual and Group Profiling

We present CARDIO-Affect, a complex-systems theoretical framework for long-term emotional dynamics in bounded social groups, with explicit uncertainty quantification at every layer. Long-period naturalistic emotion in stable small groups exhibits hallmarks of complex systems -- multi-stable attractors, weak chaos, long-range memory, and sparse heterogeneous coupling -- invisible to conventional short-clip facial-emotion analysis. CARDIO-Affect treats individual emotion as a multi-stable nonlinear stochastic dynamical system and group emotion as a sparsely-coupled network with emergent macrostates, formalised through six propositions and four pillars: (i) statistical mechanics with neural-parameterised Hamiltonian SDE over asymmetric potentials; (ii) information geometry on a 45-dimensional Fisher-Rao manifold; (iii) topological data analysis for invariant trajectory signatures; (iv) HRV-inspired Emotional Variability Analytics (EVA) decomposing each person-day into multi-scale time/frequency/nonlinear measures. We validate on the first 30.1-month longitudinal in-the-wild facial-emotion corpus (companion: arXiv:2510.15221) by discovering three falsifiable paradoxes: Sparse-Contagion (R_0=0.36, density 2.7%, 8 BH-FDR edges), Asymmetric-Persistence (negative dwell 5.85x positive, 1.77D potential gap), and Crisis-Inversion (Shanghai 2022 lockdown naive d=-0.40 collapses to permutation-p=0.94 under BSTS + synthetic-control). On synthetic benchmarks, CARDIO-EBM v2 matches asymptotically optimal Granger on linear VAR data (Class A AUROC 0.984+/-0.012 vs Granger 0.997+/-0.001, 5 seeds) but fails on tanh-coupled nonlinear data (Class B AUROC 0.490 vs Granger 0.796), a documented limitation of the linear mask-self estimator. We release framework code and the full reproduction pipeline.

physics.soc-ph

BEAM3R: Beam's-eye-view architecture with Mamba-3 for implicit dose reconstruction

To enable accurate and rapid photon control point and proton beamlet dose calculation in the DoseRAD2026 challenge, we present BEAM3R, a dose estimation framework operating in beam's-eye-view (BEV). Our core innovation combines a Mamba-3 state-space depth-sequence core with physics-based transport conditioning to model long-range depth transport without expensive 3D convolutions. BEAM3R shares a 2D CNN encoder-decoder architecture for photon and proton dose tasks, processing per-plane BEV slices. Proton beamlets are conditioned on water equivalent thickness and remaining range, encoding the parameters determining Bragg peak position. Photon models use a bidirectional Mamba-3 core to capture dose contributions from materials downstream of the calculation point, while the proton model uses a forward core with learned energy-prefix tokens and a Bragg-peak refinement module. To reduce interpolation artifacts and support high spatial resolution, we introduce axial grid alignment of BEV lattices with CT slices and an implicit super-resolution representation via sub-pixel phase packing, evaluated by a differentiable Triton-accelerated resampler that reconstructs packed cubic B-spline coefficients directly in CT space. For MRI-based tasks, synthetic CTs (sCT) are generated by a patch-based conditional GAN with a SwinUNETR backbone. On the preliminary DoseRAD2026 test set, CT-to-photon and CT-to-proton models achieved 1%/1 mm local gamma pass rates of 96.8% and 96.0%, with stratified plan-level MAEs of 0.0041 and 0.0079. Substituting sCT reduced gamma pass rates to 89.7% for photon and 75.4% proton plan level doses, with stratified plan-level MAEs of 0.0093 and 0.0336. Standardised runtimes were 23.4 s and 18.4 s for CT-to-photon and CT-to-proton prediction, increasing to 39.7 s and 42.8 s for the corresponding MRI-based pipelines.

physics.med-ph

Can AI-Assisted Inquiry Enhance Students' Decision-Making Skills in Socio-Scientific Issues? A Three-Group Experimental Study on Climate Change

Climate change is a socio-scientific issue: it rests on science but cannot be settled by science, because any serious response forces people to weigh costs, values, and competing interests under uncertainty. Helping students make such decisions well is a central aim of science education, and the arrival of generative artificial intelligence raises a sharp question: does a conversational AI partner deepen students' reasoning, or simply do the thinking for them? This study tested whether AI-assisted inquiry improves secondary students' decision-making about climate change. Using a pretest-posttest design with three groups (AI-assisted inquiry, inquiry without AI, and traditional instruction; 270 students, 90 per group), reasoning was assessed across seven decision-making steps, from defining the problem to monitoring with adaptive management, using a four-level analytic rubric scored through content analysis with high inter-coder agreement. All three groups began at comparable, mostly low levels and all improved, but the gains differed sharply. The AI-assisted group improved most, ahead of inquiry-only and of traditional instruction. Between-group effect sizes on gains were large for AI-assisted versus traditional instruction and moderate-to-large for AI-assisted versus inquiry-only, with the clearest advantages on stakeholder engagement, alternatives, implementation, and monitoring. Within the AI group, the number of times students checked the AI's claims against the sources predicted their gains, and no student was flagged for over-reliance. The findings suggest that AI helps most when it is designed to question rather than to answer, and that the inquiry it is embedded in carries much of the benefit.

cs.CY

Cognitive Cells: A Compositional Framework for Populations of Small Language Models

Recent work on large language models and agentic systems raises a basic question that current practice leaves open: how should artificial cognition be decomposed, measured, and composed? We propose studying multi-agent systems from a fixed unit we call a cognitive cell: a small, frozen language model with bounded memory and a message interface. The methodological commitment, the fixed-cell principle, is to hold this unit constant and vary only the population size, the communication topology, the message bandwidth, and the coordination protocol, so that collective behavior becomes a measurable property of a known device rather than an artifact of per-study engineering. We characterize a single cell by a compact datasheet of measurable parameters, and we ask when replicating and connecting cells improves performance: first we measure how one cell behaves alone, then we replicate it and test when voting, communication, and topology help. Instantiating the framework with small frozen models (1.5 and 3 billion parameters), we report a first round of measurements. Adding cells helps only when their errors are not too correlated. A simple correct/incorrect voting model is a useful but conservative null: real open-ended voting can exceed it, because errors are dispersed across many wrong answers rather than concentrated on one. Popular interactive protocols, namely debate, a shared blackboard, and chain revision, do not beat a matched-cost voting baseline in our setting. Finally, a cell's ability to relay several facts, itself a datasheet quantity, predicts whether a population can solve tasks whose evidence exceeds any single cell's memory. We present these as initial measurements within a broader program on scalable artificial cognition, in which multi-agent architectures appear as the special case of cells autonomous enough to be treated as agents.

physics.soc-ph

A priori Assessment of Tensor-Network Encoding for Isotropic Turbulent Flows

Tensor networks (TNs), originally developed for simulating many-body quantum systems, provide a systematic framework for approximating high-dimensional fields. This is achieved by factorizing the field into interconnected tensors with small bond dimensions, thereby restricting the correlations captured across field bipartitions. Belonging to the family of TNs, the matrix product state (MPS) ansatz is utilized here as a reduced-order modeling framework to construct truncated representations of isotropic turbulent flow data. Two direct numerical simulation (DNS) datasets are considered: the hydrodynamic field of an incompressible three-dimensional flow, and a conserved Fickian scalar in a similar flow. Each field is encoded as an MPS through a sequence of singular value decompositions (SVDs) in which small singular values are discarded. The truncated representation is contracted back to the full grid, and the resulting reconstructed field is compared against DNS. An interleaved ordering of the spatial tensor indices of the transport variables is applied prior to decomposition in order to localize the dominant inter-tensor correlations. Velocity reconstructions achieve $99.8\%$ fidelity using only $5\%$ of the original DNS memory, while the scalar field reaches the same fidelity at $15\%$ memory usage. A wide range of lower- and higher-order statistics, including velocity gradients, dissipation, and structure functions, are systematically examined. At these compression levels, the total kinetic energy and the scalar energy are both recovered within $0.2\%$ relative error, while the mean dissipation and mean scalar dissipation remain within approximately $10\%$ of the DNS generated values. These findings support the suitability of MPS for scalable reduced-order analysis of complex turbulent datasets and motivate further exploration of TN-based methods in computational turbulence.

physics.flu-dyn

How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making

Production deployments of large language model (LLM) agents remain unreliable on long, multi-step workflows even as benchmark success rates climb steadily. We argue this gap is largely an artifact of task horizon: benchmarks are dominated by short-to-medium horizons where success remains high, while production workloads demand an order of magnitude more dependent steps. We measure the effect directly, characterizing the shape of agent degradation and disentangling its cause across a large controlled study spanning nine models, six open models from 1.2B to 671B parameters, and three deployed proprietary systems; four task families, including a genuinely agentic tool-use loop; five horizons; and three context regimes. Task success follows a geometric law governed by a single per-step reliability parameter, which rises with model scale but saturates well below 1 even for the strongest models, guaranteeing eventual collapse at sufficiently long horizons. The effect is sharpest on the agentic task, where every model tested, including widely deployed systems, falls from near-perfect success to near zero within sixteen steps of (n=10,664 analyzed trajectories. Degradation is driven by step count rather than context length: bounding the context window steepens decay rather than easing it (logit slope -0.69 vs. -0.44), p=3x10-6), contradicting a lost-in-the-middle explanation and warning against a common production shortcut. Projecting measured reliability onto representative benchmark horizons quantifies a substantial gap between benchmark and production conditions, from 0.42 at GAIA-length horizons to 0.24 at hundred-step production horizons. For teams responsible for agent orchestration and reliability at scale, these results argue for horizon-aware evaluation and reliability budgeting in place of aggregate pass-rate metrics. Code, prompts, seeds, and raw trajectories are released.

physics.soc-ph

Improving Clinical Target Volume Segmentation Accuracy using Anatomical Priors and Active Learning for the AGITG TOPGEAR Clinical Trial

Training deep learning-based medical image segmentation models is challenging with limited curated datasets. For AGITG TOPGEAR, a gastric cancer trial, the Clinical Target Volume (CTV) is complex and defined by multiple anatomical landmarks, making upfront training data preparation difficult for an automated contour QA segmentation model. We investigate anatomical priors, derived from surrounding organ segmentations, to provide spatial context and improve TOPGEAR CTV segmentation accuracy. We also evaluate active learning, iteratively expanding the training dataset by selecting cases expected to improve performance. One hundred TOPGEAR CT scans were retrospectively analyzed. An initial set of 10 expert-contoured cases was used to train an nnU-Net model. TotalSegmentator generated a voxel-wise anatomical prior map from surrounding structures as an additional input channel. Active learning was simulated over four iterations, selecting cases by model uncertainty and segmentation performance. All models used five-fold cross-validation for an ensemble uncertainty measure. Evaluation used a hold-out testing set of 50 cases. The anatomical prior improved CTV segmentation accuracy, increasing mean Dice Similarity Coefficient (DSC) from 0.84 to 0.86. Active learning similarly improved performance to 0.86, with greatest benefit in the final round. Combining the anatomical prior with active learning achieved the highest accuracy, with a DSC of 0.87. Model uncertainty correlated with DSC, supporting its use in identifying suboptimal predictions and guiding active learning. Anatomical priors and active learning each improved CTV segmentation accuracy and generalizability, with their combination achieving the best performance, supporting integration into segmentation model development for automated contour QA in radiotherapy clinical trials.

physics.med-ph