arXiv · 2610.07936
Pseudowords as probes: Large Language Models show little of the sublexical sensitivity that governs human pseudoword processing
Abstract
Systematicity, the probabilistic mapping of form to meaning, permeates language at all levels, and sublexical cues have been shown to govern human pseudoword processing. Yet whether LLMs exhibit comparable sensitivity to these cues remains unclear. We tested five LLMs on two Italian two-alternative forced-choice pseudoword experiments and compared their responses with a human behavioural baseline. LLMs aligned more reliably with humans when real-word options provided a lexical familiarity cue than in the pseudoword-only condition, where they fell substantially below fastText, a character-n-gram model. In addition, the sublexical cosine-similarity cue that reliably drove human--fastText agreement did not consistently transfer to human--LLM alignment, and reasoning-token expenditure bore no consistent relation to human processing difficulty. These findings suggest that LLMs do not necessarily share the sublexical cues that govern human pseudoword processing; we discuss tokenization and training-data coverage as candidate explanations.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jing Chen, Giulia Loca, Simona Amenta, Marco Marelli. 2026-10-06. Pseudowords as probes: Large Language Models show little of the sublexical sensitivity that governs human pseudoword processing. https://arxiv.org/abs/2610.07936
Cite the original work for its findings. Save a collection to share your selection of sources.