Search arXivSearch

arXiv subjects

Christopher Kelly

Publications and source records attributed to Christopher Kelly.

At least 19 recordsLinked to original sources

Machine learning assisted Bayesian calibration of an accelerator digital twin from orbit response data

Digital twins of particle accelerators are used to plan and control operations and to design data collection campaigns. Accurate modeling typically requires knowledge of quantities that are hard to measure directly, e.g., magnet alignments, transfer functions relating power supply currents to magnetic fields, magnet nonlinearities, and stray fields. In this work we introduce multiplicative parameters to the quadrupole transfer functions to parametrize these effects. We use Bayesian methods to probabilistically estimate these parameters and their uncertainties by calibrating the Bmad digital twin to beam measurements performed at the AGS Booster at Brookhaven National Laboratory. The inference is computationally accelerated using a machine learning emulator of the physical accelerator digital twin trained to a perturbed-parameter ensemble of Bmad simulations. The result is a joint posterior distribution over the parameters constrained by the data, taking into account beam monitor errors. Incorporating estimates of the parameters into the digital twin is shown to result in a significant improvement in the quality of the model and provides error bars on the model parameters and predictions.

physics.acc-ph

High-Stakes Decisions with Language Models: Insights from Emergency Triage

High-stakes decisions under uncertainty, such as medical emergency triage, require more than accurate predictions. They depend on estimating the likelihood of alternative outcomes while explicitly weighing the consequences of different actions, principles that have long formed the foundation of medical diagnosis and decision making. Yet language models are increasingly used for high-stakes clinical recommendations without explicit specification of the utilities governing these decisions. Here we show that emergency triage with language models can be understood within a probabilistic decision framework, providing a case study of a broader decision-analytic paradigm for steering, evaluating, and deploying language models in high-stakes settings. Using clinical vignettes from a structured evaluation of a consumer triage system, we analyze recommendations for treatment under alternative utility functions that specify the relative costs of missed emergencies and unnecessary escalation. We find that capable language models adjust recommendations in response to stated utilities, revealing that the same underlying predictions can support markedly different decision policies. These findings show that effective deployment depends not only on improving predictions but also on making decision objectives explicit. More broadly, they suggest that language models for high-stakes applications should be understood and evaluated as probabilistic decision systems whose recommendations depend jointly on predictive performance and explicit utilities.

cs.AI

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation

This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established practices from disciplines with established RCT traditions, including software engineering, economics, clinical and health sciences, and psychology, we synthesize five principles drawn from established validity frameworks and open-science standards on transparency, repeatability, and verification, which together serve as the conceptual foundation for 33 actionable guidelines adapted for AI evaluation RCT contexts, expressed as requirements with rationales, implementation instructions, and evidence bases. We position the principles and guidelines as serving three key roles for AI evaluation RCTs: a design tool for planning studies, an evaluation rubric for assessing existing work, and a blueprint for standard setting as the field converges on norms. AI evaluation research currently lacks common standards and shared vocabulary for producing cumulative, comparable, policy-ready evidence. This framework is a contribution toward that foundation, providing evaluative criteria and a shared conceptual language alongside actionable guidelines.

cs.CY

How people use Copilot for Health

We analyze over 500,000 de-identified health-related conversations with Microsoft Copilot from January 2026 to characterize what people ask conversational AI about health. We develop a hierarchical intent taxonomy of 12 primary categories using privacy-preserving LLM-based classification validated against expert human annotation, and apply LLM-driven topic-clustering for prevalent themes within each intent. Using this taxonomy, we characterize the intents and topics behind health queries, identify who these queries are about, and analyze how usage varies by device and time of day. Five findings stand out. First, nearly one in five conversations involve personal symptom assessment or condition discussion, and even the dominant general information category (40%) is concentrated on specific treatments and conditions, suggesting that this is a lower bound on personal health intent. Second, one in seven of these personal health queries concern someone other than the user, such as a child, a parent, a partner, suggesting that conversational AI can be a caregiving tool, not just a personal one. Third, personal queries about symptoms and emotional health queries increase markedly in the evening and nighttime hours, when traditional healthcare is most limited. Fourth, usage diverges sharply by device: mobile concentrates on personal health concerns, while desktop is dominated by professional and academic work. Fifth, a substantial share of queries focuses on navigating healthcare systems such as finding providers, and understanding insurance, highlighting friction in the delivery of existing healthcare. These patterns have direct implications for platform-specific design, safety considerations, and the responsible development of health AI.

cs.HC

Sequential Diagnosis with Language Models

Artificial intelligence holds great promise for expanding access to expert medical knowledge and reasoning. However, most evaluations of language models rely on static vignettes and multiple-choice questions that fail to reflect the complexity and nuance of evidence-based medicine in real-world settings. In clinical practice, physicians iteratively formulate and revise diagnostic hypotheses, adapting each subsequent question and test to what they've just learned, and weigh the evolving evidence before committing to a final diagnosis. To emulate this iterative process, we introduce the Sequential Diagnosis Benchmark, which transforms 304 diagnostically challenging New England Journal of Medicine clinicopathological conference (NEJM-CPC) cases into stepwise diagnostic encounters. A physician or AI begins with a short case abstract and must iteratively request additional details from a gatekeeper model that reveals findings only when explicitly queried. Performance is assessed not just by diagnostic accuracy but also by the cost of physician visits and tests performed. We also present the MAI Diagnostic Orchestrator (MAI-DxO), a model-agnostic orchestrator that simulates a panel of physicians, proposes likely differential diagnoses and strategically selects high-value, cost-effective tests. When paired with OpenAI's o3 model, MAI-DxO achieves 80% diagnostic accuracy--four times higher than the 20% average of generalist physicians. MAI-DxO also reduces diagnostic costs by 20% compared to physicians, and 70% compared to off-the-shelf o3. When configured for maximum accuracy, MAI-DxO achieves 85.5% accuracy. These performance gains with MAI-DxO generalize across models from the OpenAI, Gemini, Claude, Grok, DeepSeek, and Llama families. We highlight how AI systems, when guided to think iteratively and act judiciously, can advance diagnostic precision and cost-effectiveness in clinical care.

cs.CL

A Tool to Facilitate Web-Browsing

Search engine results often misalign with users' goals due to opaque algorithms, leading to unhelpful or detrimental information consumption. To address this, we developed a Google Chrome plugin that provides "content labels" for webpages in Google search results, assessing Actionability (guiding actions), Knowledge (enhancing understanding), and Emotion. Using natural language processing and machine learning, the plugin predicts these properties from webpage text based on models trained on participants' ratings, effectively reflecting user perceptions. The implications include enhanced user control over information consumption and promotion of healthier engagement with online content, potentially improving decision-making and well-being.

cs.HC

Bootstrap-determined p-values in Lattice QCD

We present a general method to determine the probability that stochastic Monte Carlo data, in particular those generated in a lattice QCD calculation, would have been obtained were that data drawn from the distribution predicted by a given theoretical hypothesis. Such a probability, or p-value, is often used as an important heuristic measure of the validity of that hypothesis. The proposed method offers the benefit that it remains usable in cases where the standard Hotelling $T^2$ methods based on the conventional $\chi^2$ statistic do not apply, such as for uncorrelated fits. Specifically, we analyze a general alternative to the correlated $\chi^2$ statistic referred to as $q^2$, and show how to use the bootstrap as a data-driven method to determine the expected distribution of $q^2$ for a given hypothesis with minimal assumptions. This distribution can then be used to determine the p-value for a fit to the data. We also describe a bootstrap approach for quantifying the impact upon this p-value of estimating population parameters from a single ensemble of $N$ samples. The overall method is accurate up to a $1/N$ bias which we do not attempt to quantify.

hep-lat

Advancing Multimodal Medical Capabilities of Gemini

Many clinical tasks require an understanding of specialized data, such as medical images and genomics, which is not typically found in general-purpose large multimodal models. Building upon Gemini's multimodal models, we develop several models within the new Med-Gemini family that inherit core capabilities of Gemini and are optimized for medical use via fine-tuning with 2D and 3D radiology, histopathology, ophthalmology, dermatology and genomic data. Med-Gemini-2D sets a new standard for AI-based chest X-ray (CXR) report generation based on expert evaluation, exceeding previous best results across two separate datasets by an absolute margin of 1% and 12%, where 57% and 96% of AI reports on normal cases, and 43% and 65% on abnormal cases, are evaluated as "equivalent or better" than the original radiologists' reports. We demonstrate the first ever large multimodal model-based report generation for 3D computed tomography (CT) volumes using Med-Gemini-3D, with 53% of AI reports considered clinically acceptable, although additional research is needed to meet expert radiologist reporting quality. Beyond report generation, Med-Gemini-2D surpasses the previous best performance in CXR visual question answering (VQA) and performs well in CXR classification and radiology VQA, exceeding SoTA or baselines on 17 of 20 tasks. In histopathology, ophthalmology, and dermatology image classification, Med-Gemini-2D surpasses baselines across 18 out of 20 tasks and approaches task-specific model performance. Beyond imaging, Med-Gemini-Polygenic outperforms the standard linear polygenic risk score-based approach for disease risk prediction and generalizes to genetically correlated diseases for which it has never been trained. Although further development and evaluation are necessary in the safety-critical medical domain, our results highlight the potential of Med-Gemini across a wide range of medical tasks.

cs.CV

An Evaluation of Real-time Adaptive Sampling Change Point Detection Algorithm using KCUSUM

Detecting abrupt changes in real-time data streams from scientific simulations presents a challenging task, demanding the deployment of accurate and efficient algorithms. Identifying change points in live data stream involves continuous scrutiny of incoming observations for deviations in their statistical characteristics, particularly in high-volume data scenarios. Maintaining a balance between sudden change detection and minimizing false alarms is vital. Many existing algorithms for this purpose rely on known probability distributions, limiting their feasibility. In this study, we introduce the Kernel-based Cumulative Sum (KCUSUM) algorithm, a non-parametric extension of the traditional Cumulative Sum (CUSUM) method, which has gained prominence for its efficacy in online change point detection under less restrictive conditions. KCUSUM splits itself by comparing incoming samples directly with reference samples and computes a statistic grounded in the Maximum Mean Discrepancy (MMD) non-parametric framework. This approach extends KCUSUM's pertinence to scenarios where only reference samples are available, such as atomic trajectories of proteins in vacuum, facilitating the detection of deviations from the reference sample without prior knowledge of the data's underlying distribution. Furthermore, by harnessing MMD's inherent random-walk structure, we can theoretically analyze KCUSUM's performance across various use cases, including metrics like expected delay and mean runtime to false alarms. Finally, we discuss real-world use cases from scientific simulations such as NWChem CODAR and protein folding data, demonstrating KCUSUM's practical effectiveness in online change point detection.

cs.LG

ELIXR: Towards a general purpose X-ray artificial intelligence system through alignment of large language models and radiology vision encoders

In this work, we present an approach, which we call Embeddings for Language/Image-aligned X-Rays, or ELIXR, that leverages a language-aligned image encoder combined or grafted onto a fixed LLM, PaLM 2, to perform a broad range of chest X-ray tasks. We train this lightweight adapter architecture using images paired with corresponding free-text radiology reports from the MIMIC-CXR dataset. ELIXR achieved state-of-the-art performance on zero-shot chest X-ray (CXR) classification (mean AUC of 0.850 across 13 findings), data-efficient CXR classification (mean AUCs of 0.893 and 0.898 across five findings (atelectasis, cardiomegaly, consolidation, pleural effusion, and pulmonary edema) for 1% (~2,200 images) and 10% (~22,000 images) training data), and semantic search (0.76 normalized discounted cumulative gain (NDCG) across nineteen queries, including perfect retrieval on twelve of them). Compared to existing data-efficient methods including supervised contrastive learning (SupCon), ELIXR required two orders of magnitude less data to reach similar performance. ELIXR also showed promise on CXR vision-language tasks, demonstrating overall accuracies of 58.7% and 62.5% on visual question answering and report quality assurance tasks, respectively. These results suggest that ELIXR is a robust and versatile approach to CXR AI.

cs.CV

$\Delta I = 3/2$ and $\Delta I = 1/2$ channels of $K\to\pi\pi$ decay at the physical point with periodic boundary conditions

We present a lattice calculation of the $K\to\pi\pi$ matrix elements and amplitudes with both the $\Delta I = 3/2$ and 1/2 channels and $\varepsilon'$, the measure of direct $CP$ violation. We use periodic boundary conditions (PBC), where the correct kinematics of $K\to\pi\pi$ can be achieved via an excited two-pion final state. To overcome the difficulty associated with the extraction of excited states, our previous work \cite{Bai:2015nea,RBC:2020kdj} successfully employed G-parity boundary conditions, where pions are forced to have non-zero momentum enabling the $I=0$ two-pion ground state to express the on-shell kinematics of the $K\to\pi\pi$ decay. Here instead we overcome the problem using the variational method which allows us to resolve the two-pion spectrum and matrix elements up to the relevant energy where the decay amplitude is on-shell. In this paper we report an exploratory calculation of $K\to\pi\pi$ decay amplitudes and $\varepsilon'$ using PBC on a coarser lattice size of $24^3\times64$ with inverse lattice spacing $a^{-1}=1.023$ GeV and the physical pion and kaon masses. The results are promising enough to motivate us to continue our measurements on finer lattice ensembles in order to improve the precision in the near future.

hep-lat

Isospin 0 and 2 two-pion scattering at physical pion mass using all-to-all propagators with periodic boundary conditions in lattice QCD

A study of two-pion scattering for the isospin channels, $I=0$ and $I=2$, using lattice QCD is presented. M\"obius domain wall fermions on top of the Iwasaki-DSDR gauge action for gluons with periodic boundary conditions are used for the lattice computations which are carried out on two ensembles of gauge field configurations generated by the RBC and UKQCD collaborations with physical masses, inverse lattice spacings of 1.023 and 1.378 GeV, and spatial extents of $L=4.63$ and 4.58 fm, respectively. The all-to-all propagator method is employed to compute a matrix of correlation functions of two-pion operators. The generalized eigenvalue problem (GEVP) is solved for a matrix of correlation functions to extract phase shifts with multiple states, two pions with a non-zero relative momentum as well as two pions at rest. Our results for phase shifts for both $I=0$ and $I=2$ channels are consistent with and the Roy Equation and chiral perturbation theory, though at this preliminary stage our errors for $I=0$ are large. An important outcome of this work is that we are successful in extracting two-pion excited states, which are useful for studying $K\to\pi\pi$ decay, on physical-mass ensembles using GEVP.

hep-lat

Report of the Snowmass 2021 Topical Group on Lattice Gauge Theory

Lattice gauge theory continues to be a powerful theoretical and computational approach to simulating strongly interacting quantum field theories, whose applications permeate almost all disciplines of modern-day research in High-Energy Physics. Whether it is to enable precision quark- and lepton-flavor physics, to uncover signals of new physics in nucleons and nuclei, to elucidate hadron structure and spectrum, to serve as a numerical laboratory to reach beyond the Standard Model, or to invent and improve state-of-the-art computational paradigms, the lattice-gauge-theory program is in a prime position to impact the course of developments and enhance discovery potential of a vibrant experimental program in High-Energy Physics over the coming decade. This projection is based on abundant successful results that have emerged using lattice gauge theory over the years: on continued improvement in theoretical frameworks and algorithmic suits; on the forthcoming transition into the exascale era of high-performance computing; and on a skillful, dedicated, and organized community of lattice gauge theorists in the U.S. and worldwide. The prospects of this effort in pushing the frontiers of research in High-Energy Physics have recently been studied within the U.S. decadal Particle Physics Planning Exercise (Snowmass 2021), and the conclusions are summarized in this Topical Report.

hep-lat

Lattice QCD and Particle Physics

Contribution from the USQCD Collaboration to the Proceedings of the US Community Study on the Future of Particle Physics (Snowmass 2021).

hep-lat

Algorithms for Domain Wall Fermions

We discuss algorithms for domain wall fermions focussing on accelerating Hybrid Monte Carlo sampling of gauge configurations. Firstly a new multigrid algorithm for domain wall solvers and secondly a domain decomposed hybrid monte carlo approach applied to large subvolumes and optimised for GPU accelerated nodes. We propose a formulation of DD-RHMC that is suitable for the simulation of odd numbers of fermions.

hep-lat

Lattice QCD and the Computational Frontier

The search for new physics requires a joint experimental and theoretical effort. Lattice QCD is already an essential tool for obtaining precise model-free theoretical predictions of the hadronic processes underlying many key experimental searches, such as those involving heavy flavor physics, the anomalous magnetic moment of the muon, nucleon-neutrino scattering, and rare, second-order electroweak processes. As experimental measurements become more precise over the next decade, lattice QCD will play an increasing role in providing the needed matching theoretical precision. Achieving the needed precision requires simulations with lattices with substantially increased resolution. As we push to finer lattice spacing we encounter an array of new challenges. They include algorithmic and software-engineering challenges, challenges in computer technology and design, and challenges in maintaining the necessary human resources. In this white paper we describe those challenges and discuss ways they are being dealt with. Overcoming them is key to supporting the community effort required to deliver the needed theoretical support for experiments in the coming decade.

hep-lat

Discovering new physics in rare kaon decays

The decays and mixing of $K$ mesons are remarkably sensitive to the weak interactions of quarks and leptons at high energies. They provide important tests of the standard model at both first and second order in the Fermi constant $G_F$ and offer a window into possible new phenomena at energies as high as 1,000 TeV. These possibilities become even more compelling as the growing capabilities of lattice QCD make high-precision standard model predictions possible. Here we discuss and attempt to forecast some of these capabilities.

hep-lat

Testing And Hardening IoT Devices Against the Mirai Botnet

A large majority of cheap Internet of Things (IoT) devices that arrive brand new, and are configured with out-of-the-box settings, are not being properly secured by the manufactures, and are vulnerable to existing malware lurking on the Internet. Among them is the Mirai botnet which has had its source code leaked to the world, allowing any malicious actor to configure and unleash it. A combination of software assets not being utilised safely and effectively are exposing consumers to a full compromise. We configured and attacked 4 different IoT devices using the Mirai libraries. Our experiments concluded that three out of the four devices were vulnerable to the Mirai malware and became infected when deployed using their default configuration. This demonstrates that the original security configurations are not sufficient to provide acceptable levels of protection for consumers, leaving their devices exposed and vulnerable. By analysing the Mirai libraries and its attack vectors, we were able to determine appropriate device configuration countermeasures to harden the devices against this botnet, which were successfully validated through experimentation.

cs.CY