arXiv · 2609.31192
Learning a non-linguistic code for inferred rules from reward
Abstract
How can a rule inferred from examples reach someone who never saw them, without a shared code? Patients with severe aphasia do it by gesture or sketch. One network sees worked examples and emits eight invented symbols; a second, blind to them, applies them to a new input. Rewarded for the second's success, the first learns a code carrying rules to three-step transformations training never presents, which new learners acquire. Like invented human languages, the code has two regimes: under reward alone the speaker drifts to one message, as human languages lose words under plain transmission; expressive pressure keeps messages differentiated. Success on new rules tracks how much the message says about the rule, not how varied messages are. The learning signal shapes the code: reward sorts many rules under few fixed labels; the listener's error gradient gives each rule a region of similar messages, telling rules apart far better.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Cristiano Capone. 2026-09-25. Learning a non-linguistic code for inferred rules from reward. https://arxiv.org/abs/2609.31192
Cite the original work for its findings. Save a collection to share your selection of sources.