Search arXivSearch

arXiv subjects

Matthew Ball

Publications and source records attributed to Matthew Ball.

9 recordsLinked to original sources

Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI

Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in high-stakes settings. Yet this work has not produced a market that rewards trustworthiness. Firms that invest seriously in safety, fairness, and oversight cannot consistently prove to consumers, regulators, and shareholders that their systems go beyond the bare minimum of compliance. What is missing is a way for society to recognize or compare the difference. The result is a trust gap: a structural condition in which responsible development efforts happen inside organizations but produce no external, independently recognized and verifiable signal of trustworthy outcomes. We argue this gap is sustained in part because of a focus on responsible AI (a matter of internal process) as opposed to trustworthy AI (a matter of independently verifiable real-world outcomes), and that it persists because of three compounding failures: (1) the market cannot distinguish trustworthy systems from their imitations; (2) evaluation targets models and outputs rather than deployed sociotechnical systems and their outcomes; (3) the measurement ecosystem is oriented toward avoiding harm rather than demonstrating benefit. Reviewing existing AI governance instruments and comparing them to certification regimes in healthcare, sustainability, and security, we show that none integrate a governance baseline, independently verified positive-outcome evidence, and market signaling in a single framework. We propose independent, outcome-oriented certification as the connective layer that can close the trust gap, complementing regulation and internal governance by making trustworthiness measurable, comparable, and commercially rewarded.

cs.AI

Harmonizing AI Safety Thresholds

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.

cs.AI

Designing escalation criteria for international AI incident response: criteria, triggers, and thresholds

AI incident reporting requirements are emerging in regulation and policy, yet no operational criteria exist for determining when a detected AI incident warrants escalation beyond national handling to international coordination. This paper proposes an escalation framework to address this gap, intended as a common reference point across jurisdictions that enables aligned escalation while preserving flexibility in how actors respond within their own legal and policy contexts. We review SB 53, the EU AI Act, the GPAI Code of Practice, and incident frameworks from other industries to derive eight criteria for assessing whether an incident warrants escalation, translated into a sequential flowchart with gated decision points and threshold checks. For each criterion, we map how it interplays with these regulatory frameworks, identifying where their design choices support or undermine effective detection. We test the framework against ten documented AI incidents and structured variants to identify where criteria under-detect or misclassify incidents in practice. We find three design patterns that may lead to systematic under-detection in regimes where model developers are responsible for escalation: a. where escalation requires confirmed harm, events such as model weight exfiltration risk detection only after severe, irreversible harm has propagated; b. where incidents are assessed individually, systemic harms emerging from accumulation risk being under-detected; and c. where thresholds align with legal instruments rather than quantitatively testable terms, criteria risk being impractical to apply under time pressure. We also find that escalation rules are only one component of a broader framework: the underlying definitions against which thresholds are set, and the data available to the responsible actor, create interdependencies that can themselves drive under-detection.

cs.CY

Methods for Estimating Neutron Star Parameters using Multiple Mechanisms for Gravitational Wave Emission Associated with Pulsar Glitches

Several mechanisms for gravitational wave (GW) emission are believed to be associated with pulsar glitches. This emission may be split between long duration continuous waves and short duration bursts. In the Advanced LIGO era, searches for GWs associated with pulsar glitches have only considered continuous wave emission. The increasing sensitivity of the detectors and the prospects for future detectors suggest that astrophysically motivated analyses involving multiple mechanisms may be possible. Here, we present a framework for combining two simple models for GW emission - long duration continuous waves and short duration bursts - to derive more constraining astrophysical implications than a single model would allow. The best limits arise from using models that predict a specific amount of GW emission; however, there are relatively few models that make such predictions. We apply these methods to the December 2016 Vela pulsar glitch and make predictions for how well future observing runs and detectors would improve results. As part of this analysis, we performed a targeted search for GW bursts associated with this glitch and find no signal.

astro-ph.HE

Squeezing the quantum noise of a gravitational-wave detector below the standard quantum limit

Precision measurements of space and time, like those made by the detectors of the Laser Interferometer Gravitational-wave Observatory (LIGO), are often confronted with fundamental limitations imposed by quantum mechanics. The Heisenberg uncertainty principle dictates that the position and momentum of an object cannot both be precisely measured, giving rise to an apparent limitation called the Standard Quantum Limit (SQL). Reducing quantum noise below the SQL in gravitational-wave detectors, where photons are used to continuously measure the positions of freely falling mirrors, has been an active area of research for decades. Here we show how the LIGO A+ upgrade reduced the detectors' quantum noise below the SQL by up to 3 dB while achieving a broadband sensitivity improvement, more than two decades after this possibility was first presented.

gr-qc

Prospects for Neutron Star Parameter Estimation using Gravitational Waves from f-modes Associated with Magnetar Flares

Magnetar vibrational modes are theorized to be associated with energetic X-ray flares. Regular searches for gravitational waves from these modes have been performed by Advanced LIGO and Advanced Virgo, with no detections so far. Presently, search results are given in limits on the root-sum-square of the integrated gravitational-wave strain. However, the increased sensitivity of current detectors and the promise of future detectors invite the consideration of more astrophysically motivated methods. We present a framework for augmenting gravitational wave searches to measure or place direct limits on magnetar astrophysical properties in various search scenarios using a set of phenomenological and analytic models.

astro-ph.HE

Correlated 1-1000 Hz magnetic field fluctuations from lightning over earth-scale distances and their impact on gravitational wave searches

We report Earth-scale distance magnetic correlations from lightning strokes in the frequency range 1-1000 Hz at several distances ranging from 1100 to 9000 km. Noise sources which are correlated on Earth-scale distances can affect future searches for gravitational-wave signals with ground-based gravitational-wave interferometric detectors. We consider the impact of correlations from magnetic field fluctuations on gravitational-wave searches due to Schumann resonances ($<$50 Hz) as well as higher frequencies ($>$100 Hz). We demonstrate that individual lightning strokes are a likely source for the observed correlations in the magnetic field fluctuations at gravitational-wave observatories and discuss some of their characteristics. Furthermore, we predict their impact on searches for an isotropic gravitational-wave background, as well as for searches looking for short-duration transient gravitational waves, both unmodeled signals (bursts) as well as modeled signals (compact binary coalescence). Whereas the recent third observing run by LIGO and Virgo was free of an impact from correlated magnetic field fluctuations, future runs could be affected. For example, at current magnetic coupling levels, neutron star inspirals in third generation detectors are likely to be contaminated by multiple correlated lightning glitches. We suggest that future detector design should consider reducing lightning coupling by, for example, reducing the lightning-induced beam tube currents that pass through sensitive magnetic coupling regions in current detectors. We also suggest that the diurnal and seasonal variation in lightning activity may be useful in discriminating between detector correlations that are produced by gravitational waves and those produced by lightning.

gr-qc

Point absorbers in Advanced LIGO

Small, highly absorbing points are randomly present on the surfaces of the main interferometer optics in Advanced LIGO. The resulting nano-meter scale thermo-elastic deformations and substrate lenses from these micron-scale absorbers significantly reduces the sensitivity of the interferometer directly though a reduction in the power-recycling gain and indirect interactions with the feedback control system. We review the expected surface deformation from point absorbers and provide a pedagogical description of the impact on power build-up in second generation gravitational wave detectors (dual-recycled Fabry-Perot Michelson interferometers). This analysis predicts that the power-dependent reduction in interferometer performance will significantly degrade maximum stored power by up to 50% and hence, limit GW sensitivity, but suggests system wide corrections that can be implemented in current and future GW detectors. This is particularly pressing given that future GW detectors call for an order of magnitude more stored power than currently used in Advanced LIGO in Observing Run 3. We briefly review strategies to mitigate the effects of point absorbers in current and future GW wave detectors to maximize the success of these enterprises.

physics.ins-det

Toward Automated Virtual Assembly for Prefabricated Construction: Construction Sequencing through Simulated BIM

To adhere to the stringent time and budget requirements of construction projects, contractors are utilizing prefabricated construction methods to expedite the construction process. Prefabricated construction methods require an adequate schedule and understanding by the contractors and constructors to be successful. The specificity of prefabricated construction often leads to inefficient scheduling and costly rework time. The designer, contractor, and constructors must have a strong understanding of the assembly process to experience the full benefits of the method. At the root of understanding the assembly process is visualizing how the process is intended to be performed. Currently, a virtual construction model is used to explain and better visualize the construction process. However, creating a virtual construction model is currently time consuming and requires experienced personnel. The proposed simulation of the virtual assembly will increase the automation of virtual construction modeling by implementing the data available in a building information modeling (BIM) model. This paper presents various factors (i.e., formalization of construction sequence based on the level of development (LOD)) that needs to be addressed for the development of automated virtual assembly. Two case studies are presented to demonstrate these factors.

cs.AI