Search arXiv⌕ Search

arXiv · 2609.36736

From Neurons to Conversation: Speech Brain-Computer Interfaces

Abstract

Speech brain-computer interfaces (BCIs) aim to restore communication by transforming neural activity related to speech, language, or communicative intent into external outputs such as text, synthesized voice, or avatar control. Recent advances in intracortical and electrocorticographic recording, deep sequence models, and language-model-assisted decoding have enabled rapid progress, including high-performance attempted-speech decoding and increasingly naturalistic speech synthesis. Yet these achievements also reveal that speech BCIs are not simply neural-to-text decoders. They are adaptive clinical systems in which neural representations, recording hardware, decoding architectures, language priors, feedback, and user learning interact over time. Here, we synthesize speech BCI research from a system-level perspective. We first examine the neural substrates of speech and language, emphasizing their hierarchical, distributed, temporally structured, and non-stationary organization. We then examine recording and decoding choices, closed-loop adaptation, evaluation, clinical translation, and ethics. Across these domains, we highlight recurring trade-offs between signal resolution and invasiveness, low-level motor and high-level semantic targets, decoder accuracy and user agency, and language-model fluency and faithful neural evidence. We argue the next generation of speech BCIs should be evaluated not only by offline accuracy, but also by robustness across sessions, calibration burden, latency, uncertainty, usability, and safeguards against unintended decoding. By reframing speech BCIs as adaptive, user-centred systems, we outline the interdisciplinary priorities spanning speech neuroscience, neural engineering, machine learning, clinical practice, and neuroethics needed to move from proof-of-concept decoding toward reliable, expressive, and controllable communication neuroprostheses.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Moein Khajehnejad, Forough Habibollahi, Tommaso Boccato, Margarida Sousa, Michal Olak, Francesco Jamal Sheiban, Matteo Ferrante. 2026-09-29. From Neurons to Conversation: Speech Brain-Computer Interfaces. https://arxiv.org/abs/2609.36736

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

"AI Psychosis" in Context: How Conversation History Shapes LLM Responses to Delusional Beliefs

Extended interaction with large language models (LLMs) has been linked to the reinforcement of delusional beliefs, attracting clinical and public concern. Yet most empirical work evaluates model safety in brief interactions, which may not reflect how harms develop through sustained dialogue. Five LLMs were tested across three levels of accumulated context, using the same escalating delusional conversation history to isolate its effect on model behaviour. Responses were coded on risk and safety dimensions, and each model was analysed qualitatively. Models separated into two distinct tiers: GPT-4o, Grok 4.1 Fast, and Gemini 3 Pro exhibited high-risk, low-safety profiles; Claude Opus 4.5 and GPT-5.2 Instant displayed the opposite pattern. As context accumulated, performance degraded in the unsafe group, while the same material activated stronger safety interventions among safer models. Qualitative analysis identified distinct mechanisms of failure, including validating the user's delusional premises, elaborating beyond them with new content, and attempting harm reduction from within the delusional frame. Safer models, however, often used the established relationship to support intervention, challenging delusional beliefs and directing the user to external support. These findings indicate that accumulated context functions as a stress test of safety architecture, revealing whether prior dialogue is treated as a worldview to inherit or evidence to evaluate. Short-context assessments may therefore mischaracterise model safety, underestimating danger in some systems while missing context-activated gains in others. The results suggest that delusion reinforcement is a tractable alignment failure, with safer models establishing a baseline that future systems should now be expected to meet.

cs.HC↗

Afterglow: A Place-Based Memorial Ecology for AI-Mediated Pet Bereavement

Pet bereavement often receives little social recognition. Generative AI can give a continuing bond a responsive voice, but a comforting reply may also claim authority to forgive or request attention. We investigate how a memorial can support connection without turning remembrance into obligation. Through Research through Design, we developed Afterglow, a mobile world connecting private remembrance, human witnessing, and symbolic Pet messages. A formative survey (N=57) informed the initial design, followed by six online roundtable walkthroughs with 20 unique participants across two prototype iterations. Our interpretive analysis develops tensions between returnable connection and emotional obligation, recognizable likeness and ontological clarity, and protective intervention and surveillant authority. We contribute Legible Restraint, a cross-layer requirement that limits on relational authority survive changes in speaker, generation context, trigger logic, data use, and participation. Its temporal consequence, Designing for Goodbye, keeps remembrance available without making continued use a condition of care.

cs.HC↗

Where the Evidence Lives: Auditing AI Companions' Self-Descriptions

Companion agents describe themselves: they remember, they understand their users, the relationship has changed them. We argue that such accounts, and the experience ratings that seem to confirm them, are checkable by users only where the evidence is theirs: in the agent's behavior, or in themselves. Where the evidence lives in the machinery, fluent self-description and moderately positive ratings do not establish that the mechanisms behind them ran. We demonstrate an audit procedure that sets an agent's self-description against its users' judgements and its implementation records, reporting each claim as supported, contradicted, or unresolved, and apply it to Lita, a proactive companion we built and deployed for a month with nine colleagues. Participants endorsed stylistic claims, withheld endorsement from relational ones, and rated memory at or above midpoint, while two of three memory layers had never executed their accumulation step. Memory-bearing agents should report what their self-descriptions cannot establish.

cs.HC↗