Search arXiv⌕ Search

arXiv · 2509.25721

The AI Productivity Index (APEX)

Abstract

We present an extended version of the AI Productivity Index (APEX-v1-extended), a benchmark for assessing whether frontier models are capable of performing economically valuable tasks in four jobs: investment banking associate, management consultant, big law associate, and primary care physician (MD). This technical report details the extensions to APEX-v1, including an increase in the held-out evaluation set from n = 50 to n = 100 cases per job (n = 400 total) and updates to the grading methodology. We present a new leaderboard, where GPT5 (Thinking = High) remains the top performing model with a score of 67.0%. APEX-v1-extended shows that frontier models still have substantial limitations when performing typical professional tasks. To support further research, we are open sourcing n = 25 non-benchmark example cases per role (n = 100 total) along with our evaluation harness.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Bertie Vidgen, Abby Fennelly, Evan Pinnix, Julien Benchek, Daniyal Khan, Zach Richards, Austin Bridges, Calix Huang, Kanishka Sahu, Abhishek Kottamasu, Bo Ma, Ben Hunsberger, Isaac Robinson, Akul Datta, Chirag Mahapatra, Dominic Barton, Cass R. Sunstein, Eric Topol, Brendan Foody, Osvald Nitski. 2025-12-16. The AI Productivity Index (APEX). https://arxiv.org/abs/2509.25721

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Critical Mathematical Economics and Progressive Computer Science

The aim of this article is to present elements and discuss the potential of a research program at the intersection between mathematics and heterodox economics, which we call Criticial Mathematical Economics (CME). We propose to focus on the mathematical and model-theoretic foundations of controversies in economic policy, and aim at providing an entrance to the literature as an invitation to mathematicians that are potentially interested in such a project. From our point of view, mathematics has been partly misused in mainstream economics to justify `unregulated markets'. We identify two key parts of CME, which leads to a natural structure of this article: The first part focusses on an analysis and critique of mathematical models used in mainstream economics, like e.g. the Dynamic Stochastic General Equilibrium (DSGE) in Macroeconomics and the so-called ``Sonnenschein-Mantel-Debreu''-Theorems. The aim of the second part is to improve and extend heterodox models using ingredients from modern mathematics and computer science, a method with strong relation to Complexity Economics. We exemplify this idea by describing how methods from Non-Linear Dynamics have been used in Post-Keynesian Macroeconomics', and also discuss (Pseudo-) Goodwin cycles and possible Micro- and Mesofoundations. Finally, we outline in which areas a collaboration between mathematicians and heterodox economists could be most promising, and discuss both existing projects in such a direction as well as areas where new models for policy advice are most needed. In an outlook, we discuss the role of (ecological) data, AI and the need for what we call Progressive Computer Science.

econ.GN↗

Don't Fake It If You Can't Make It: Driver Misconduct in Last-Mile Delivery

In the last two decades, last-mile delivery (LMD) firms have seen immense growth fueled by the success of e-commerce, leading to faster and cheaper deliveries. Operating on thin margins, LMD firms strive for successful first-time deliveries to avoid the financial and reputational costs of reattempts. Delivery Agents (DAs) are integral to LMD efficiency, influencing customer experience, delivery success, and productivity. However, most LMD performance enhancement research focuses on process, technology, and incentives, which presume workers will conform to procedures and monitoring tools will function flawlessly. Nevertheless, in practice, DAs deviate from expected behaviors, i.e., indulge in misconduct, negatively affecting delivery efficiency, often resulting in returned parcels. One of the major misconducts is fake remarked deliveries, wherein DAs intentionally do not deliver the parcels and provide a fake reason for it. For instance, even without reaching a delivery address, a DA remarks 'customer unavailable' and records a delivery failure. In this study, we collaborated with a leading Indian LMD firm and, using instrumental variable regression, find that such misconduct leads to a spillover productivity loss. This effect reduces the next day's successful deliveries by 1.60% and first-time-right deliveries by 1.86%. We discuss misconduct's correlation with factors such as task complexity and offer novel insights into how opportunistic circumstances can influence worker behavior.

econ.GN↗

Decomposing Wage Stagnation: Employment Reallocation, Wage Structure,and Demographics

Japan's average log real hourly wages rose until the mid-1990s, declined through the mid-2010s, and partially recovered thereafter. This paper decomposes these changes over 1980-2024 into four components: demographic change across worker types, changes in relative employment shares across job types, changes in relative log wages across job types, and unweighted mean wage growth. The framework combines a shift-share decomposition across worker types with an extension of the Olley-Pakes decomposition across job types within worker types, separating employment reallocation from changes in relative wage structure. The four components vary across periods. Before 1996, unweighted mean wage growth and changes in relative wage structure contribute positively, while demographic change and employment reallocation contribute negatively. During 1996-2014, all four components are negative. After 2014, the recovery mainly reflects unweighted mean wage growth. Employment reallocation and changes in relative wage structure contribute differently across dimensions of job type.

econ.GN↗