Search arXiv⌕ Search

arXiv subjects

John J. Horton

Publications and source records attributed to John J. Horton.

9 recordsLinked to original sources

Chaining Tasks, Redefining Work: A Theory of AI Automation

We study the automation and productivity effects of AI when production is a sequence of interdependent steps with a Leontief structure. Each step can be executed manually by labor, augmented with AI, or automated within contiguous AI-executed steps called "chains." By characterizing the firm's cost-minimizing assignment of steps to humans and AI, we show that chaining can overturn comparative advantage logic, since human-advantaged steps may be optimally assigned to AI as part of a chain. Chaining also makes gains from automation depend on how AI-exposed steps cluster in the workflow beyond just how many there are, and can make the returns to improving AI quality non-monotone in a pattern resembling the productivity J-curve. Empirically, we show that existing automation patterns are consistent with the model's predictions: a step is more likely to be AI-executed when its neighbors are, and AI execution is more common where AI-exposed steps cluster in the workflow.

econ.GN↗

LLM Bidders Preserve the Mechanism-Level Orderings of Human Bidders

Training on vast amounts of human-generated data has motivated growing interest in using large language models (LLMs) to simulate human behavior. We ask which features of human behavior general-purpose models preserve when used out of the box in auctions, where multiple bidders interact under explicit rules and incentives. We evaluate five LLMs across seven laboratory settings against human benchmarks reconstructed from published experiments, with uncertainty bands for the private-value comparisons. Our main focus is on three large models without extended test-time reasoning: GPT-4o, Claude~3.5 Haiku, and Gemini~2.0 Flash. LLM and human deviations from theory differ in magnitude and often in direction: humans overbid in second-price auctions, whereas most models that deviate underbid. Surprisingly, without task-specific fine-tuning or calibration to human bids, the three non-reasoning large models robustly preserve key orderings of auction formats by deviation from theory. First-price auctions are harder than second-price, and ascending clocks reduce deviations relative to sealed bids wherever data are adequate. Kendall's $τ_b$ between the human and GPT-4o difficulty rankings is $0.60$ and positive in every joint bootstrap draw. The reasoning model bids almost at equilibrium in the observed private-value settings, leaving little variation in errors to compare; the small model's large errors yield an inverted ranking. All five models nevertheless reproduce the stronger first-price winner's curse. Clock framing improves bidding for two of the three non-reasoning large models, and GPT-4o recovers the ordering of last-minute bidding across closing rules in an eBay-style marketplace.

cs.GT↗

General Social Agents

Useful social science theories predict behavior across settings. However, applying a theory to make predictions in new settings is challenging: rarely can it be done without ad hoc modifications to account for setting-specific factors. We argue that AI agents put in simulations of those novel settings offer an alternative for applying theory, requiring minimal or no modifications. We present an approach for building such "general" agents that use theory-grounded natural language instructions, existing empirical data, and knowledge acquired by the underlying AI during training. To demonstrate the approach in settings where no data from that data-generating process exists--as is often the case in applied prediction problems--we design a heterogeneous population of 883,320 novel games. AI agents are constructed using human data from a small set of conceptually related but structurally distinct "seed" games. In preregistered experiments, on average, agents predict initial human play in a random sample of 1,500 games from the population better than (i) a cognitive hierarchy model, (ii) game-theoretic equilibria, and (iii) out-of-the-box agents. For a small set of separate novel games, these simulations predict responses from a new sample of human subjects better even than the most plausibly relevant published human data.

econ.GN↗

Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?

We argue that newly-developed large language models (LLMs), because of how they are trained and designed, are implicit computational models of humans -- a Homo silicus. LLMs can be used like economists use Homo economicus: they can be given endowments, information, preferences, and so on, and then their behavior can be explored in scenarios via simulation. Experiments using this approach, derived from Charness and Rabin (2002), Kahneman et al. (1986), Samuelson and Zeckhauser (1988), Oprea (2024b), and Horton (2025), show qualitatively similar results to the original, and when they differ, it is often generative for future research. We discuss potential applications, conceptual issues, and why this approach can inform the study of humans.

econ.GN↗

Automated Social Science: Language Models as Scientist and Subjects

We present an approach for automatically generating and testing, in silico, social scientific hypotheses. This automation is made possible by recent advances in large language models (LLM), but the key feature of the approach is the use of structural causal models. Structural causal models provide a language to state hypotheses, a blueprint for constructing LLM-based agents, an experimental design, and a plan for data analysis. The fitted structural causal model becomes an object available for prediction or the planning of follow-on experiments. We demonstrate the approach with several scenarios: a negotiation, a bail hearing, a job interview, and an auction. In each case, causal relationships are both proposed and tested by the system, finding evidence for some and not others. We provide evidence that the insights from these simulations of social interactions are not available to the LLM purely through direct elicitation. When given its proposed structural causal model for each scenario, the LLM is good at predicting the signs of estimated effects, but it cannot reliably predict the magnitudes of those estimates. In the auction experiment, the in silico simulation results closely match the predictions of auction theory, but elicited predictions of the clearing prices from the LLM are inaccurate. However, the LLM's predictions are dramatically improved if the model can condition on the fitted structural causal model. In short, the LLM knows more than it can (immediately) tell.

econ.GN↗

Algorithmic Writing Assistance on Jobseekers' Resumes Increases Hires

There is a strong association between the quality of the writing in a resume for new labor market entrants and whether those entrants are ultimately hired. We show that this relationship is, at least partially, causal: a field experiment in an online labor market was conducted with nearly half a million jobseekers in which a treated group received algorithmic writing assistance. Treated jobseekers experienced an 8% increase in the probability of getting hired. Contrary to concerns that the assistance is taking away a valuable signal, we find no evidence that employers were less satisfied. We present a model in which better writing is not a signal of ability but helps employers ascertain ability, which rationalizes our findings.

econ.GN↗

The Tragedy of Your Upstairs Neighbors: Is the Airbnb Negative Externality Internalized?

A commonly expressed concern about the rise of the peer-to-peer rental market Airbnb is that hosts---those renting out their properties---impose costs on their unwitting neighbors. I consider the question of whether apartment building owners will, in a competitive rental market, set a building-specific Airbnb hosting policy that is socially efficient. I find that if tenants can sort across apartments based on the owners policy then the equilibrium fraction of buildings allowing Airbnb listing would be socially efficient.

q-fin.GN↗

Employer Expectations, Peer Effects and Productivity: Evidence from a Series of Field Experiments

This paper reports the results of a series of field experiments designed to investigate how peer effects operate in a real work setting. Workers were hired from an online labor market to perform an image-labeling task and, in some cases, to evaluate the work product of other workers. These evaluations had financial consequences for both the evaluating worker and the evaluated worker. The experiments showed that on average, evaluating high-output work raised an evaluator's subsequent productivity, with larger effects for evaluators that are themselves highly productive. The content of the subject evaluations themselves suggest one mechanism for peer effects: workers readily punished other workers whose work product exhibited low output/effort. However, non-compliance with employer expectations did not, by itself, trigger punishment: workers would not punish non-complying workers so long as the evaluated worker still exhibited high effort. A worker's willingness to punish was strongly correlated with their own productivity, yet this relationship was not the result of innate differences---productivity-reducing manipulations also resulted in reduced punishment. Peer effects proved hard to stamp out: although most workers complied with clearly communicated maximum expectations for output, some workers still raised their production beyond the output ceiling after evaluating highly productive yet non-complying work products.

cs.HC↗

The Online Laboratory: Conducting Experiments in a Real Labor Market

Online labor markets have great potential as platforms for conducting experiments, as they provide immediate access to a large and diverse subject pool and allow researchers to conduct randomized controlled trials. We argue that online experiments can be just as valid---both internally and externally---as laboratory and field experiments, while requiring far less money and time to design and to conduct. In this paper, we first describe the benefits of conducting experiments in online labor markets; we then use one such market to replicate three classic experiments and confirm their results. We confirm that subjects (1) reverse decisions in response to how a decision-problem is framed, (2) have pro-social preferences (value payoffs to others positively), and (3) respond to priming by altering their choices. We also conduct a labor supply field experiment in which we confirm that workers have upward sloping labor supply curves. In addition to reporting these results, we discuss the unique threats to validity in an online setting and propose methods for coping with these threats. We also discuss the external validity of results from online domains and explain why online results can have external validity equal to or even better than that of traditional methods, depending on the research question. We conclude with our views on the potential role that online experiments can play within the social sciences, and then recommend software development priorities and best practices.

cs.HC↗