Search arXiv⌕ Search

arXiv · 2609.30780

Sampling Safe Futures: Multimodal Trajectory Planning for Personalized Safety in Anthropomorphic AI

Abstract

Anthropomorphic artificial intelligence systems increasingly remember personal details, display empathy, and are engaged with as social counterparts, creating forms of risk that emerge from the evolution of the user-system relationship over time. Existing safeguards largely operate at the level of individual conversational turns and cannot determine whether a sequence of seemingly acceptable interactions is cumulatively moving a particular user toward harm. This paper introduces personalized trajectory-level safety, a framework that treats relational safety as a sequential decision problem over a latent escalation state inferred from the user's messages and influenced by the system's responses. At each turn, a screening step first discards any response strategy that does not preserve at least one safe continuation of the interaction under every plausible model of the user. Among the remaining strategies, we formulate action selection as multimodal trajectory sampling, and use a Generative Flow Network to generate diverse future evolutions in proportion to their plausibility, safety, and utility. The system then selects the strategy that preserves the largest fraction of safe and useful continuations. We evaluate the framework in simulation, calibrated on statistics reported for real human-chatbot interactions, and using response strategies derived from public benchmarks. Results show that trajectory-aware decision making substantially reduces the frequency of harmful states while keeping helpful interaction. This work reframes safety for anthropomorphic AI from response-level filtering to personalized control over the future evolution of human-AI relationships. The source code is available at https://github.com/benedettapicano/ANTHROPOMORPHIC_SAFETY_TRAJ.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Benedetta Picano, Dusit Niyato. 2026-09-25. Sampling Safe Futures: Multimodal Trajectory Planning for Personalized Safety in Anthropomorphic AI. https://arxiv.org/abs/2609.30780

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

"If You're Not Doing It, Somebody Else Is": Active Negotiation and the Invisible Labor of Sustained LLM Use

Large language models (LLMs) have become fixtures of academic work even as their users describe them as degrading their writing, thinking, and skills. Dominant adoption frameworks read continued use as evidence of satisfaction, and cannot explain continued use of a distrusted tool. We interviewed 36 graduate student workers, balanced between English-as-a-foreign-language (EFL) and non-EFL speakers, and introduce the Active Negotiation framework: a model of sustained LLM use as a recurring cycle of risk, mitigation, and justification. A failure surfaces a risk, mitigation labor addresses it, and a justification renders the residual risk tolerable until the next failure reopens the cycle. The cycle runs across three dimensions: practical, auditing output; internal, auditing one's own cognition and identity; and social, managing how peers and institutions perceive use. EFL participants invoke linguistic parity as a further justification. We reframe continued adoption as compliance sustained by invisible labor.

cs.HC↗

Skill Profiling with Attributable Reasoning (SPAR): A Wearable Analysis System for Boxing

A punch is a ballistic, full-body action driven by a kinetic chain running from the legs through the trunk to the arm, where a small sequencing error separates a scoring strike from a miss. Wearable sensors can capture this movement in the gym, but most deployable systems only classify which punch was thrown rather than assess how well it was thrown. We present Skill Profiling with Attributable Reasoning (SPAR), an eight-IMU garment and pressure-insole system that classifies each punch as expert or novice and treats an explanation of that prediction as feedback. Feedback is only useful if the person receiving it can act on it, so SPAR explains the prediction at three tiers, a per-joint attribution for the analyst, a counterfactual over kinetic-chain layers for the coach, and a plain-language narrative of the two for the athlete. Across 17 participants and 4,713 punches, SPAR reaches a leave-one-participant-out AUC of 0.842 (95% CI [0.769, 0.907] over participants). A frozen time-series foundation model encodes the joint-angle and plantar-force series, and a small transformer trained on the cohort classifies the encoding. We audit the two quantitative tiers and report six themes from a thematic analysis of interviews with six practicing boxing coaches.

cs.HC↗

CDBG: Causally Motivated Dual-Invariance Learning against Topological and Predictive Shifts in EEG Workload Recognition

Generalizing Electroencephalography (EEG)-based mental workload recognition to unseen subjects remains a formidable challenge due to severe inter-subject variability. While functional brain graphs effectively model distributed cognitive dynamics, their inherent subject-specificity induces two coupled distribution shifts: a class-conditional topological shift in the underlying functional connectivity, and a predictive mechanism shift in the learned representation-to-label mapping. Motivated by the subject-induced distribution shifts, we propose CDBG, a Causally motivated Dual-invariance learning framework for Brain Graphs. CDBG disentangles and mitigates these shifts via a two-stage rationale learning pipeline. First, it employs stochastic edge masking to extract sparse, workload-predictive graph rationales, regularized by workload-conditional Laplacian spectral alignment to enforce topological invariance across subjects. Second, it applies subject-wise Invariant Risk Minimization (IRM) to the graph representations, ensuring environment-wise risk stationarity. Extensive experiments on a self-built air traffic controller EEG cognitive workload dataset and multiple public datasets under a strict leave-one-subject-out protocol demonstrate that CDBG significantly outperforms state-of-the-art cross-subject and graph-based baselines, improving the Macro-F1 score by up to 4.23%, while simultaneously providing neurophysiologically interpretable functional rationales.

cs.HC↗