Search arXivSearch

arXiv · 2601.07573

A Model of Artificial Jagged Intelligence

Abstract

Generative AI systems often display highly uneven performance across tasks that appear ``nearby'': they can be excellent on one prompt and confidently wrong on another with only small changes in wording or context. We call this phenomenon Artificial Jagged Intelligence (AJI). This paper develops a tractable economic model of AJI that treats adoption as an information problem: users care about \emph{local} reliability, but typically observe only coarse, global quality signals. In a baseline one-dimensional landscape, truth is a rough Brownian process, and the model ``knows'' scattered points drawn from a Poisson process. The model interpolates optimally, and the local error is measured by posterior variance. We derive an adoption threshold for a blind user, show that experienced errors are amplified by the inspection paradox, and interpret scaling laws as denser coverage that improves average quality without eliminating jaggedness. We then study mastery and calibration: a calibrated user who can condition on local uncertainty enjoys positive expected value even in domains that fail the blind adoption test. Modelling mastery as learning a reliability map via Gaussian process regression yields a learning-rate bound driven by information gain, clarifying when discovering ``where the model works'' is slow. Finally, we study how scaling interacts with discoverability: when calibrated signals and user mastery accelerate the harvesting of scale improvements, and when opacity can make gains from scaling effectively invisible.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Joshua Gans. 2026-01-12. A Model of Artificial Jagged Intelligence. https://arxiv.org/abs/2601.07573

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Log-concave functions and transformations thereof

I summarize Bagnoli and Bergstrom (2005)'s review on log-concave functions, make several corrections, and augment the discussion with further results that can be useful in establishing monotone hazard rates. I also provide an application to monopoly pricing, where log-concavity of the demand curve implies strict concavity of the revenue function in quantity.

econ.TH

Audit the Auditors: Commitment versus Professional Judgment

This paper provides a theoretical framework to evaluate the trade-off between the self-regulated peer review system and independent government inspection (PCAOB) in the auditing profession. We model the peer review system as a Judgment Regime, where a stakeholder utilizes professional expertise, captured as a private signal, to make ex-post decisions on verifying audit failures. In contrast, PCAOB inspection is modeled as a Commitment Regime, where the stakeholder lacks private information but can commit ex-ante to a predetermined level of verification. We find that the Judgment Regime benefits from a resource-allocation effect and a deterrence effect driven by informed verification, whereas the Commitment Regime deters audit failures through the first-mover advantage of ex-ante commitment. Our analysis demonstrates that the stakeholder prefers the peer review system if and only if the private signal is sufficiently informative. Furthermore, comparative statics reveal that higher verification costs or stronger audit incentives shift the stakeholder's preference toward PCAOB inspection.

econ.TH

Reputation, Disclosure, and the Scope of Entry

This paper studies how learning about an incumbent affects the scope of entry when competitive responses use resources shared across markets. An entrant chooses whether to launch in neither, one, or both of two markets. Entry into the second market reduces the incumbent's cost-reducing response in the first and can make one-market entry unattractive. The entrant learns about the incumbent's capability from a record of its response to an earlier rival. More frequent publication encourages a less capable incumbent to imitate a more capable one. An observed response then becomes less informative, and entry after that record expands. We compare publication of conduct with a public audit of capability. For an open set of parameters with uniform setup costs, full publication maximizes total surplus within the specified policy class when publication costs are low. Removing the interaction between response costs across markets reverses this choice, while preserving all early and singleton-market payoffs. A capability audit is dominated in both economies. Expected entry scope is constant across the considered policies within each technology, although productive investment and the allocation of entry change.

econ.TH