Search arXivSearch

arXiv · 2607.02325

Personality Without Persons? A Psychometric Critique of Big Five Testing in Large Language Models

Abstract

Human personality inventories are increasingly used to characterize large language models (LLMs), compare systems, and inform downstream governance claims. Yet, these inventories were developed and validated for humans, and it remains unclear whether they are valid for non-human systems. We present a systematic psychometric evaluation of Big Five personality measurement in LLMs. We ask three research questions: Do Big Five inventories a) appropriately describe LLMs, b) capture meaningful differences between models, and c) reflect internal factors consistent with human personality? We assess the content validity of five candidate Big Five inventories and administer the best-performing inventory to N = 264 LLMs spanning 50 model families. Our findings are threefold. First, Big Five items adapted for LLMs achieve acceptable content validity, whereas the original human-developed items do not. Second, Big Five inventories fail to capture meaningful differences across LLMs: between-model variance accounts for only 7% - 17% of the total score variance. Third, LLMs responses do not reproduce the canonical Big Five five-factor structure of human personality, with four of the five personality facets collapsing into one (r >= .90). Moreover, comparisons between base and instruction-tuned variants suggest that alignment training shifts Big Five scores toward socially desirable profiles. These findings demonstrate that Big Five inventories do not measure a construct equivalent to human personality in LLMs. Thus, using human personality frameworks to characterize, benchmark, compare, or govern LLMs risks producing misleading conclusions. We highlight the need for evaluation frameworks that are specifically designed and validated for LLMs, rather than transferring human psychological constructs without first establishing their validity.

Explore related subjects

Keep this discovery

BibTeXRIS

Kim Zierahn, Cristina Cachero, Anna Korhonen, Nuria Oliver. 2026-09-02. Personality Without Persons? A Psychometric Critique of Big Five Testing in Large Language Models. https://arxiv.org/abs/2607.02325

Cite the original work for its findings. Save a collection to share your selection of sources.

Discover connections

Connections use source metadata and explicit phrase matches, not verified experimental comparisons.

KEEP EXPLORING

Related papers

GRAND-HC: Graph-Refined Author Name Disambiguation

From-Scratch Name Disambiguation (SND) groups papers sharing an ambiguous name into clusters of distinct real-world authors. Existing methods suffer from two critical limitations: (1) inherent long-tailed author distribution biases representation learning, causing over-merging of tail authors; (2) existing cluster number estimation methods are unreliable for long paper sequences, hindering large-scale deployment. We propose \textbf{GRAND-HC}, a complete end-to-end SND framework. We construct a heterogeneous paper graph via co-author, co-organization, and co-venue relations, using a graph attention network as the embedding backbone. \textbf{Harmony Contrastive Learning (HCL)} dynamically reweights training loss to suppress overfitting to prolific authors, learning discriminative embeddings. A \textbf{Graph-Refined Distance Matrix (GRDM)} leverages graph topology to optimize pairwise distances, further preventing tail author over-merging. Meanwhile, a lightweight \textbf{Paper Compression Module (PCM)} achieves accurate cluster number estimation across varying scales. Finally, Hierarchical Agglomerative Clustering outputs the final clusters. Extensive experiments demonstrate state-of-the-art macro F1 performance. GRAND-HC has been deployed in a billion-scale academic database. Source code: https://github.com/baokou-fw2/GRAND-HC.

cs.IR

FocusAdapt: Context-aware Adaptive Focus Assistance in Diminished Reality

Diminished Reality (DR) can reduce visual clutter by removing irrelevant objects. However, removing all task-irrelevant objects may eliminate useful contextual information and reduce situational awareness. We present FocusAdapt, a context-aware DR system that predicts object-level distraction by integrating visual saliency, semantic relevance, and gaze behavior. Based on findings from a formative study, FocusAdapt selectively diminishes highly distracting objects while preserving useful context, enabling adaptive focus assistance during procedural tasks.

cs.HC

TSExplorer: An interactive data annotation and exploration tool for time-series data

We present TSExplorer, a cross-platform tool for interactive annotation and exploration of time-series data. The tool enables users to inspect high-dimensional datasets through multiple complementary 2D visualizations derived from high-dimensional feature representations. TSExplorer is designed as a general-purpose research tool supporting a wide range of workflows, including exploratory data analysis, annotation of unlabeled or partially-labeled datasets, comparison of feature representations, and post-hoc inspection and refinement of existing labels with interactive visual feedback.

cs.HC