Search arXiv⌕ Search

arXiv · 2609.35779

Large Language Models Exhibit Human-Like Bayesian Hypocrisy

Abstract

Given recent achievements of large language models (LLMs), frontier models are expected to perform well on Bayesian reasoning tasks, at least as well as humans. Furthermore, there is no reason to expect that LLMs will condemn others who offer those very same Bayesian judgments, a fallibility observed in human decision-making (Cao, et al., 2019). In 5 experiments with 48 experimental conditions employing over 5,000 trials, GPT-4o and Claude 3.7 Sonnet were tested on two variations of a Bayesian reasoning task. We also assessed LLM evaluation of the competence and morality of a hypothetical person who had offered the same reasoning task as them. LLMs hovered near human performance on the Bayesian task, though their reasoning was more rule-based and rigid. Surprisingly, like humans but to a greater extent, LLMs also demonstrated the same hypocrisy in condemning others who, like them, had deployed Bayes' rule. In demonstrating Bayesian hypocrisy, LLMs highlight a humanlike error of a dissociation between self-performance and other-judgment, and caution against their use in domains where statistical fidelity and fairness norms collide.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Nykko Vitali, Mahzarin R. Banaji. 2026-08-03. Large Language Models Exhibit Human-Like Bayesian Hypocrisy. https://arxiv.org/abs/2609.35779

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Characterizing Creativity in Data Visualization: Reflections and Future Directions

Creativity is valued in data visualization design and research, yet how it manifests remains unclear. Understanding creativity is essential for developing expressive visualization tools and representations. In this paper, we present two complementary studies to characterize creativity in data visualization. First, a systematic review of 104 papers yields a design space spanning four themes: design frameworks incorporating divergent and convergent thinking activities; creative representations developing unorthodox visualizations; visualization-enabled creativity support tools supporting creative tasks (e.g., writing); and understanding creativity, creative processes, and influencing factors. Second, qualitative interviews with 11 visualization practitioners and researchers explore practical challenges and contrast them with current academic framing through our design space. Findings indicate that artifacts or final products are often disproportionately considered as the primary indicator of creativity, whereas the design process remains undervalued in practical. We conclude by presenting a definition of creativity in data visualization and directions for future research.

cs.HC↗

Evaluation format, not model capability, drives measured triage failure in the assessment of consumer health AI

A recent Nature Medicine study reported that ChatGPT Health under-triages 51.6% of emergencies and concluded that consumer-facing AI triage poses safety risks. Its protocol, however, was an exam-style scaffold (forced A/B/C/D output, knowledge suppression, no clarifying questions) unlike how consumers use health chatbots. We ask whether the headline error rate is a property of the models or of the measurement. In a first, mechanistic study, five frontier LLMs on a 17-scenario bank scored 6.4 points higher under naturalistic patient-style messages than under the constrained scaffold (p=0.015), and on one vignette three models went from 0-24% with forced choice to 100% with free text. In a second, faithful replication we ran the authors' own 60 released vignettes through six frontier models under four matched formats, with clinician validation of the rewrites and a blinded clinician audit of the LLM adjudicators. Here the direction reversed: free-text rewrites scored slightly below the exact structured prompt (78.7% vs 81.8%, p=0.020), removing only the answer scaffold changed little (80.6% vs 81.4%, p=0.81), and naturalistic input with a forced categorical answer beat both the exact prompt (84.4%, p=0.023) and free text (p=0.0001). On the four vignettes defining the original emergency rate, under-triage was 17% with the exact scaffold, 50% in free text, and 21% when the same message was answered with a forced letter; every free-text "under-triage" was a same-day recommendation scored C rather than D, and 60% made escalation conditional on information the patient was asked to check, behavior a single-turn benchmark cannot score. The headline rate is therefore largely a property of output format and of mapping prose onto a four-point scale. Benchmark scaffolds are behaviorally active instruments; safety claims should report sensitivity to input wording, output format and adjudication.

cs.HC↗

HandMade: Spatial Prompting for Generative 3D Creation with Part-Labeled VR Sketches

Text-to-3D generation lowers the barrier to 3D content creation, but text alone is a weak interface for specifying spatial intent: where parts should be placed, how they relate, and how an object should be organized in 3D. We present HandMade, a workflow that combines VR 3D sketching and language for open-domain 3D asset generation. HandMade treats coarse, part-labeled 3D sketches not as incomplete geometry to reconstruct directly, but as spatial prompts for existing generative models. It converts segmented VR strokes into multi-view part guidance and structured prompts, allowing users to specify object layout and part relationships through 3D sketching while using language for identity, material, style, and local details. A technical evaluation shows that HandMade better preserves user-authored spatial scaffolds than text-only and sketch-based baselines on 20 varied examples. A user study with eight participants characterizes how users make use of 3D sketching for spatial layout and language for identity, materials, and details across initial authoring and subsequent revision. HandMade contributes an interaction paradigm and interface-to-generation pipeline for spatially guided 3D creation.

cs.HC↗