Search arXivSearch

arXiv · 2604.03338

The Ideation Bottleneck: Decomposing the Quality Gap Between AI-Generated and Human Economics Research

Abstract

Autonomous AI systems can now generate complete economics research papers, but they substantially underperform human-authored publications in head-to-head comparisons. This paper decomposes the quality gap into two independent components: research idea quality and execution quality. Using a two-model ensemble of fine-tuned language models trained on publication decisions (Gong, Li, and Zhou, 2026) to evaluate idea quality and a comprehensive six-dimension rubric assessed by Gemini 3.1 Flash Lite -- the same model family used as the APE tournament judge, ensuring methodological consistency -- to evaluate execution quality, we analyze 953 economics papers -- 912 AI-generated papers from the APE project and 41 human papers published in the American Economic Review and AEJ: Economic Policy. The idea quality gap is large (Cohen's d = 2.23, p < 0.001), with human papers achieving 47.1% mean ensemble exceptional probability versus 16.5% for AI. The execution quality gap is also significant but smaller (d = 0.90, p < 0.001), with human papers scoring 4.38/5.0 versus 3.84. Idea quality accounts for approximately 71% of the overall quality difference, with execution contributing 29%. The largest execution weakness is mechanism analysis depth (d = 1.43); no significant difference is found on robustness. We document that 74% of AI papers employ difference-in-differences, and only 7 AI papers (0.8%) surpass the median human paper on both idea and execution quality simultaneously. The primary bottleneck to competitive AI-generated economics research remains ideation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ning Li. 2026-04-03. The Ideation Bottleneck: Decomposing the Quality Gap Between AI-Generated and Human Economics Research. https://arxiv.org/abs/2604.03338

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The time interpretation of expected utility theory

Economic models often maximise expectation values of wealth or utility. In non-ergodic settings, these can differ from time-averages, so that maximising expected outcomes need not maximise – and can systematically reduce – long-run wealth or utility. Ergodicity economics highlights this problem and models individual agents as maximising wealth in the long run, known as growth optimality. Two instances where expected utility maximisation maps to growth optimality are known: linear utility does this for additive wealth dynamics; and logarithmic utility for multiplicative wealth dynamics. Here we show that the mapping holds more generally when the utility function coincides with the ergodicity transformation in the growth optimal model. This mapping offers a theoretical basis for choosing utility functions and suggests the testable hypothesis that wealth dynamics are predictive of risk preferences.

econ.GN

Monetary Regimes and Trade before the Classical Gold Standard: Evidence from the Latin Monetary Union

This paper reexamines the trade effects of the Latin Monetary Union (LMU), a 19th century agreement to standardize gold and silver coinage among several European countries. The LMU provides a useful setting for studying whether monetary arrangements fostered trade before the classical gold standard, when gold, silver, bimetallic, and paper regimes coexisted. Because some countries already shared other monetary standards, treating all non-member pairs as a single control group mixes pairs with and without alternative forms of monetary coordination. I classify pairs by standard and estimate the LMU effect relative to pairs without a common standard, bringing the comparison closer to those used in the literature on the gold standard and contemporary currency unions. The results suggest that the LMU increased trade between its members by approximately 30\% during its early years, when bimetallism was still credible. These effects subsequently faded, converging to zero by the end of the 1870s. More broadly, these findings also highlight the importance of accounting for the existing monetary regimes when estimating the trade effects of other international policies.

econ.GN

Access to Live AI Advice and Behavior Under Risk: An Incentivized Experiment

Generative AI has become an everyday advisor, and the systems people consult are live and interactive, not pre-scripted. We ask whether access to such a system changes behavior under risk. In an incentivized experiment (N = 158), participants made lottery choices with an optional decision aid presented as a conventional pre-written tool, a live one-shot AI, or a live interactive AI they could query, with information format held equivalent across conditions. Risk preferences are elicited via DOSE. We find no evidence that access to a live AI advisor changes risk aversion.

econ.GN