Search arXiv⌕ Search

arXiv subjects

Ricardo Correia

Publications and source records attributed to Ricardo Correia.

4 recordsLinked to original sources

CausalVerify: End-to-End Verification of Causal Analyses by Language Models

Language models increasingly perform empirical analyses end to end, yet existing evaluations assess the written explanation or whether generated code executes, not whether the executed workflow recovers the intended causal estimand. We introduce CausalVerify, an execution-grounded benchmark for end-to-end causal analysis that follows a model from research-context interpretation to estimand recovery. It scores this workflow at four distinct layers: method recognition, design specification, executable implementation, and estimand recovery. It combines 259 real-paper contexts, 100 fixed-seed synthetic scenarios with executable reference estimates, and 23 paper-twin pairs in which a model commits to a design before seeing the data and its executed analysis is scored against a canonical estimator on the same realised dataset. Execution is not correctness. Among 426 model-written workflows that run without error, 15.5% fail verification, and a keyword score of the effect direction stated in the text is a poor proxy for recovery. In a 23-pair, nine-model paired study, replacing a model's committed design with the reference design and its execution conventions raises joint recovery of the point estimate and standard error from 15.0% to 51.5%; yet 48.5% of eligible seeds still fail under the reference design. The direction replicates on six pairs built afterwards under a frozen construction protocol, although on the two newest pairs the gain is confined to models from the family that built the references. Design specification is consequential but not sufficient: a plausible method and runnable code do not guarantee recovery, and even supplying the reference design and its conventions leaves substantial downstream failure. CausalVerify evaluates the executed workflow rather than its surface plausibility, and every reported number is recomputed from frozen artifacts by a single script.

cs.AI↗

Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations

Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Across 26 economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence. In a five-paper pilot with labels reconciled by two annotators, it assigns a higher risk level than the labels on 25 of the 33 assessments it completes. A rule fixed before the labels arrived removes most of this in-sample; weighted agreement stays low. ARGUS provides evidence-linked risk reports that localize potential weaknesses for expert review, without adjudicating causal claims. Code and data: https://github.com/yonghongzhang-io/ARGUS

cs.CL↗

The Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the Mean

Agent benchmarks are increasingly used to compare large language models (LLMs) and guide deployment decisions, yet benchmark scores are meaningful only if they measure model capability rather than properties of the evaluation pipeline. We identify a double measurement confound: execution-critical decisions are performed by a fixed scaffold instead of the model, while the scorer evaluates outputs using criteria that may not reflect task correctness. We unify these issues within a measurement-theoretic framework that characterizes when benchmark scores can be interpreted as evidence of model capability, and instantiate it with an audit-and-repair protocol that (i) transfers execution-critical decisions from the scaffold to the model, (ii) replaces shape-based evaluation with seeded ground-truth scoring, and (iii) reports reliability beyond the mean through worst-case and tail-risk metrics. Experiments on ComtradeBench show that the joint intervention transforms a nearly flat leaderboard into a reliability spectrum that distinguishes both average performance and robustness across seeds. Applying the audit to existing benchmarks further shows that scorer validity is benchmark-specific, whereas scaffold ownership is an uncontrolled axis wherever we probed it. Our results suggest that benchmark scores should be interpreted together with their scaffolding level, scoring criterion, and reliability profile, providing a practical framework for more valid evaluation of LLM agents.

cs.SE↗

Formal Simulation and Visualisation of Hybrid Programs

The design and analysis of systems that combine computational behaviour with physical processes' continuous dynamics - such as movement, velocity, and voltage - is a famous, challenging task. Several theoretical results from programming theory emerged in the last decades to tackle the issue; some of which are the basis of a proof-of-concept tool, called Lince, that aids in the analysis of such systems, by presenting simulations of their respective behaviours. However being a proof-of-concept, the tool is quite limited with respect to usability, and when attempting to apply it to a set of common, concrete problems, involving autonomous driving and others, it either simply cannot simulate them or fails to provide a satisfactory user-experience. The current work complements the aforementioned theoretical approaches with a more practical perspective, by improving Lince along several dimensions: to name a few, richer syntactic constructs, more operations, more informative plotting systems and errors messages, and a better performance overall. We illustrate our improvements via a variety of examples that involve both autonomous driving and electrical systems.

eess.SY↗