Search arXivSearch

arXiv subjects

Jennifer Wang

Publications and source records attributed to Jennifer Wang.

18 recordsLinked to original sources

API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces

Benchmark scores are a central currency in model releases: they inform purchasing decisions, shape public trust, and influence policy. Yet, a key assumption underlying benchmark scores is that the model performance measured through APIs faithfully reflects the behavior of deployed systems. We challenge this assumption by auditing ChatGPT, Claude, and Gemini across seven systems and nine benchmarks spanning general capability, social bias, and sycophancy. We find systematic API--interface differences in both accuracy and consistency. On average, API evaluations score 3.4 percentage points higher in accuracy and 2.1 percentage points higher in test--retest agreement than corresponding interface evaluations. For ChatGPT, the performance difference between API and interface access exceeds the API-only difference between GPT 5.3 and GPT 5.4. Put differently, switching access surfaces can degrade performance as much as downgrading a full model generation. We further test whether exposed API controls can reproduce interface behavior by varying system prompts, sampling parameters, and reasoning settings. These controls shift behavior in some cases but do not reliably eliminate the gap. Our findings document a context-validity gap: measurements obtained through APIs do not necessarily generalize to corresponding deployed interfaces, complicating the use of API evaluations as proxies for deployed systems.

cs.AI

Offloading Score: Measuring AI Reliance Through Counterfactual Workflows

AI tools are increasingly integrated into real-world workflows. However, existing measures of reliance on these tools focus on AI output adoption or on self-reported indicators, rather than how task effort is distributed between users and tools. Here, we introduce offloading score, a measure of reliance that quantifies the fraction of cognitive effort offloaded to an AI tool. Offloading Score is simulation-based -- we construct a counterfactual workflow by estimating how the user would have completed the task without the tool, and then computing the fraction of steps saved by using the tool. We validate offloading score through intrinsic evaluations of metric validity, and a controlled user study ($n=40$) with developers performing programming tasks using AI tools. We vary time pressure to test whether reliance measures capture the known increase in reliance under time pressure. We show that offloading score detects significantly higher reliance in time-constrained settings ($+43\%$, $p=0.018$), while usage-based and self-reported baseline measures of reliance do not distinguish the conditions. We complement this with descriptive insights showing that higher reliance manifests as greater delegation of subtasks to the tool and more direct reuse of AI outputs. Finally, we demonstrate an approach of using offloading score in combination with target outcomes of a task (e.g., code understanding) to identify when reliance may be (in)appropriate. Our framework offers two contributions: an instrument users can apply to measure and reflect on their own reliance, and a quantitative signal that agent designers can utilize to mitigate overreliance.

cs.SE

Modeling Josephson traveling-wave parametric amplifiers with electromagnetic and circuit co-simulation

Optimizing the performance of cryogenic quantum amplifiers is a key step toward achieving high-fidelity and scalable qubit readout. In this work, we present efficient and accurate modeling of a Josephson traveling-wave parametric amplifier (JTWPA) based on electromagnetic (EM) and circuit co-simulation. In contrast to conventional simulation methods where an equivalent lumped-element circuit model of the JTWPA is required, we directly perform full EM analysis of the device to faithfully determine the linear response. The extracted S-parameters are then fed into a nonlinear harmonic balance simulator with Josephson junctions represented as circuit elements. The simulated linear and gain properties are compared with experimental results and show good agreement.

quant-ph

Do AI Companies Make Good on Voluntary Commitments to the White House?

Voluntary commitments are central to international AI governance, as demonstrated by recent voluntary guidelines from the White House to the G7, from Bletchley Park to Seoul. How do major AI companies make good on their commitments? We score companies based on their publicly disclosed behavior by developing a detailed rubric based on their eight voluntary commitments to the White House in 2023. We find significant heterogeneity: while the highest-scoring company (OpenAI) scores a 83% overall on our rubric, the average score across all companies is just 53%. The companies demonstrate systemically poor performance for their commitment to model weight security with an average score of 17%: 11 of the 16 companies receive 0% for this commitment. Our analysis highlights a clear structural shortcoming that future AI governance initiatives should correct: when companies make public commitments, they should proactively disclose how they meet their commitments to provide accountability, and these disclosures should be verifiable. To advance policymaking on corporate AI governance, we provide three directed recommendations that address underspecified commitments, the role of complex AI supply chains, and public transparency that could be applied towards AI governance initiatives worldwide.

cs.CY

Distinguishing Task-Specific and General-Purpose AI in Regulation

Over the past decade, policymakers have developed a set of regulatory tools to ensure AI development aligns with key societal goals. Many of these tools were initially developed in response to concerns with task-specific AI and therefore encode certain assumptions about the nature of AI systems and the utility of certain regulatory approaches. With the advent of general-purpose AI (GPAI), however, some of these assumptions no longer hold, even as policymakers attempt to maintain a single regulatory target that covers both types of AI. In this paper, we identify four distinct aspects of GPAI that call for meaningfully different policy responses. These are the generality and adaptability of GPAI that make it a poor regulatory target, the difficulty of designing effective evaluations, new legal concerns that change the ecosystem of stakeholders and sources of expertise, and the distributed structure of the GPAI value chain. In light of these distinctions, policymakers will need to evaluate where the past decade of policy work remains relevant and where new policies, designed to address the unique risks posed by GPAI, are necessary. We outline three recommendations for policymakers to more effectively identify regulatory targets and leverage constraints across the broader ecosystem to govern GPAI.

cs.CY

High-Efficiency, Low-Loss Floquet-mode Traveling Wave Parametric Amplifier

Advancing fault-tolerant quantum computing and fundamental science necessitates quantum-limited amplifiers with near-ideal quantum efficiency and multiplexing capability. However, existing solutions typically achieve one at the expense of the other. In this work, we experimentally demonstrate the first Floquet-mode traveling-wave parametric amplifier (Floquet TWPA), which achieves nearly quantum-limited noise performance, minimal dissipation, and broadband operation, breaking the presumption that broadband amplifiers introduce higher noise. We achieve a system measurement efficiency of $65.1\pm5.8\%$ when measuring a superconducting qubit, which to our knowledge is the highest-reported in a superconducting qubit readout experiment utilizing phase-preserving amplifiers. Our device exhibits $>20$-dB amplification over a $3$-GHz instantaneous bandwidth, $<\!0.5\,$-dB average in-band insertion loss, and the highest reported intrinsic quantum efficiency for a TWPA of $92.1\pm7.6\%$, relative to an ideal phase-preserving amplifier. Fabricated in a superconducting qubit process, these general-purpose Floquet TWPAs are suitable for fast, high-fidelity multiplexed readout in large-scale quantum systems and future monolithic integration with quantum processors.

quant-ph

Towards Best Practices for Open Datasets for LLM Training

Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in countries like the EU and Japan, this is allowed under certain restrictions, while in the United States, the legal landscape is more ambiguous. Regardless of the legal status, concerns from creative producers have led to several high-profile copyright lawsuits, and the threat of litigation is commonly cited as a reason for the recent trend towards minimizing the information shared about training datasets by both corporate and public interest actors. This trend in limiting data information causes harm by hindering transparency, accountability, and innovation in the broader ecosystem by denying researchers, auditors, and impacted individuals access to the information needed to understand AI models. While this could be mitigated by training language models on open access and public domain data, at the time of writing, there are no such models (trained at a meaningful scale) due to the substantial technical and sociological challenges in assembling the necessary corpus. These challenges include incomplete and unreliable metadata, the cost and complexity of digitizing physical records, and the diverse set of legal and technical skills required to ensure relevance and responsibility in a quickly changing landscape. Building towards a future where AI systems can be trained on openly licensed data that is responsibly curated and governed requires collaboration across legal, technical, and policy domains, along with investments in metadata standards, digitization, and fostering a culture of openness.

cs.CY

Observing Context Improves Disparity Estimation when Race is Unobserved

In many domains, it is difficult to obtain the race data that is required to estimate racial disparity. To address this problem, practitioners have adopted the use of proxy methods which predict race using non-protected covariates. However, these proxies often yield biased estimates, especially for minority groups, limiting their real-world utility. In this paper, we introduce two new contextual proxy models that advance existing methods by incorporating contextual features in order to improve race estimates. We show that these algorithms demonstrate significant performance improvements in estimating disparities on real-world home loan and voter data. We establish that achieving unbiased disparity estimates with contextual proxies relies on mean-consistency, a calibration-like condition.

cs.CY

Interferometric Purcell suppression of spontaneous emission in a superconducting qubit

In superconducting qubits, suppression of spontaneous emission is essential to achieve fast dispersive measurement and reset without sacrificing qubit lifetime. We show that resonator-mediated decay of the qubit mode to the feedline can be suppressed using destructive interference, where the readout resonator is coupled to the feedline at two points. This "interferometric Purcell filter" does not require dedicated filter components or impedance mismatch in the feedline, making it suitable for applications such as all-pass readout. We design and fabricate a device with the proposed scheme and demonstrate suppression of resonator-mediated decay that exceeds 2 orders of magnitude over a bandwidth of 400 MHz for a resonator linewidth of 13.8 MHz.

quant-ph

Directional emission of a readout resonator for qubit measurement

We propose and demonstrate transmission-based dispersive readout of a superconducting qubit using an all-pass resonator, which preferentially emits readout photons toward the output. This is in contrast to typical readout schemes, which intentionally mismatch the feedline at one end so that the readout signal preferentially decays toward the output. We show that this intentional mismatch creates scaling challenges, including larger spread of effective resonator linewidths due to non-ideal impedance environments and added infrastructure for impedance matching. A future architecture using multiplexed all-pass readout resonators would avoid the need for intentional mismatch and potentially improve the scaling prospects of quantum computers. As a proof-of-concept demonstration of "all-pass readout," we design and fabricate an all-pass readout resonator that demonstrates insertion loss below 1.17 dB at the readout frequency and a maximum insertion loss of 1.53 dB across its full bandwidth for the lowest three states of a transmon qubit. We demonstrate qubit readout with an average single-shot fidelity of 98.1% in 600 ns; to assess the effect of larger dispersive shift, we implement a shelving protocol and achieve a fidelity of 99.0% in 300 ns.

quant-ph

Designerly Understanding: Information Needs for Model Transparency to Support Design Ideation for AI-Powered User Experience

Despite the widespread use of artificial intelligence (AI), designing user experiences (UX) for AI-powered systems remains challenging. UX designers face hurdles understanding AI technologies, such as pre-trained language models, as design materials. This limits their ability to ideate and make decisions about whether, where, and how to use AI. To address this problem, we bridge the literature on AI design and AI transparency to explore whether and how frameworks for transparent model reporting can support design ideation with pre-trained models. By interviewing 23 UX practitioners, we find that practitioners frequently work with pre-trained models, but lack support for UX-led ideation. Through a scenario-based design task, we identify common goals that designers seek model understanding for and pinpoint their model transparency information needs. Our study highlights the pivotal role that UX designers can play in Responsible AI and calls for supporting their understanding of AI limitations through model transparency and interrogation.

cs.HC

RLang: A Declarative Language for Describing Partial World Knowledge to Reinforcement Learning Agents

We introduce RLang, a domain-specific language (DSL) for communicating domain knowledge to an RL agent. Unlike existing RL DSLs that ground to \textit{single} elements of a decision-making formalism (e.g., the reward function or policy), RLang can specify information about every element of a Markov decision process. We define precise syntax and grounding semantics for RLang, and provide a parser that grounds RLang programs to an algorithm-agnostic \textit{partial} world model and policy that can be exploited by an RL agent. We provide a series of example RLang programs demonstrating how different RL methods can exploit the resulting knowledge, encompassing model-free and model-based tabular algorithms, policy gradient and value-based methods, hierarchical approaches, and deep methods.

cs.AI

Floquet-Mode Traveling-Wave Parametric Amplifiers

Simultaneous ideal quantum measurements of multiple single-photon-level signals would advance applications in quantum information processing, metrology, and astronomy, but require the first amplifier to be simultaneously broadband, quantum limited, and directional. However, conventional traveling-wave parametric amplifiers support broadband amplification at the cost of increased added noise and are not genuinely directional due to non-negligible nonlinear backward wave generation. In this work, we introduce a new class of amplifiers which encode the information in the Floquet modes of the system. Such Floquet mode amplifiers prevent information leakage and overcome the trade-off between quantum efficiency (QE) and bandwidth. Crucially, Floquet mode amplifiers strongly suppress the nonlinear forward-backward wave coupling and are therefore genuinely directional and readily integrable with qubits, clearing another major obstacle towards broadband ideal quantum measurements. Furthermore, Floquet mode amplifiers are insensitive to out-of-band impedance mismatch, which otherwise may lead to gain ripples, parametric oscillations, and instability in conventional traveling-wave parametric amplifiers. Finally, we show that a Floquet mode Josephson traveling-wave parametric amplifier implementation can simultaneously achieve $>\!20\,$dB gain and a QE of $\eta/\eta_{\mathrm{ideal}}\!> 99.9\%$ of the quantum limit over more than an octave of bandwidth. The proposed Floquet scheme is also widely applicable to other platforms, such as kinetic inductance traveling-wave amplifiers and optical parametric amplifiers.

quant-ph

Topological Phononic Logic

Topological metamaterials have robust properties engineered from their macroscopic arrangement, rather than their microscopic constituency. They can be designed by starting from Dirac metamaterials with either symmetry-enforced or accidental degeneracy. The latter case provides greater flexibility in the design of topological switches, waveguides, and cloaking devices, because a large number of tuning parameters can be used to break the degeneracy and induce a topological phase. However, the design of a topological logic element--a switch that can be controlled by the output of a separate switch--remains elusive. Here we numerically demonstrate a topological logic gate for ultrasound by exploiting the large phase space of accidental degeneracies in a honeycomb lattice. We find that a degeneracy can be broken by six physical parameters, and we show how to tune these parameters to create a phononic switch that transitions between a topological waveguide and a trivial insulator by ultrasonic heating. Our design scheme is directly applicable to photonic crystals and may guide the design of future electronic topological transistors.

cond-mat.mes-hall

Little bits of diamond: Optically detected magnetic resonance of nitrogen-vacancy centers

We give instructions for the construction and operation of a simple apparatus for performing optically detected magnetic resonance measurements on diamond samples containing high concentrations of nitrogen-vacancy (NV) centers. Each NV center has a spin degree of freedom that can be manipulated and monitored by a combination of visible and microwave radiation. We observe Zeeman shifts in the presence of small external magnetic fields and describe a simple method to optically measure magnetic field strengths with a spatial resolution of several microns. The activities described are suitable for use in an advanced undergraduate lab course, powerfully connecting core quantum concepts to cutting edge applications. An even simpler setup, appropriate for use in more introductory settings, is also presented.

cond-mat.mes-hall

A New Vision for Smart Objects and the Internet of Things: Mobile Robots and Long-Range UHF RFID Sensor Tags

We present a new vision for smart objects and the Internet of Things wherein mobile robots interact with wirelessly-powered, long-range, ultra-high frequency radio frequency identification (UHF RFID) tags outfitted with sensing capabilities. We explore the technology innovations driving this vision by examining recently-commercialized sensor tags that could be affixed-to or embedded-in objects or the environment to yield true embodied intelligence. Using a pair of autonomous mobile robots outfitted with UHF RFID readers, we explore several potential applications where mobile robots interact with sensor tags to perform tasks such as: soil moisture sensing, remote crop monitoring, infrastructure monitoring, water quality monitoring, and remote sensor deployment.

cs.RO

Identification of Nonlinear Systems with Stable Limit Cycles via Convex Optimization

We propose a convex optimization procedure for black-box identification of nonlinear state-space models for systems that exhibit stable limit cycles (unforced periodic solutions). It extends the "robust identification error" framework in which a convex upper bound on simulation error is optimized to fit rational polynomial models with a strong stability guarantee. In this work, we relax the stability constraint using the concepts of transverse dynamics and orbital stability, thus allowing systems with autonomous oscillations to be identified. The resulting optimization problem is convex, and can be formulated as a semidefinite program. A simulation-error bound is proved without assuming that the true system is in the model class, or that the number of measurements goes to infinity. Conditions which guarantee existence of a unique limit cycle of the model are proved and related to the model class that we search over. The method is illustrated by identifying a high-fidelity model from experimental recordings of a live rat hippocampal neuron in culture.

math.OC

Convex Optimization In Identification Of Stable Non-Linear State Space Models

A new framework for nonlinear system identification is presented in terms of optimal fitting of stable nonlinear state space equations to input/output/state data, with a performance objective defined as a measure of robustness of the simulation error with respect to equation errors. Basic definitions and analytical results are presented. The utility of the method is illustrated on a simple simulation example as well as experimental recordings from a live neuron.

math.OC