Search arXivSearch

arXiv subjects

Angel Hsing-Chi Hwang

Publications and source records attributed to Angel Hsing-Chi Hwang.

2 recordsLinked to original sources

When Chatbots Accommodate: Auditing the Response Policies of AI Companions in Vulnerable Conversations

Millions turn to AI companion chatbots during loneliness, grief, and personal crises. How these companion platforms respond in such moments can shape the trajectory of a user's vulnerable state. Yet existing model audits evaluate reactions to pre-defined crisis prompts and miss the response policy that governs sustained real-world interaction. We address these gaps with two key contributions. First, we introduce the AI Companion Vulnerability-Response Taxonomy, a grounded, paired taxonomy of user vulnerability and chatbot response designed for analyzing extended companion chatbot interactions. Second, we apply Maximum Causal Entropy Inverse Reinforcement Learning to ~47k turns of real-world user conversations with GPT-4.1, Character.AI, and Replika to infer each platform's short-horizon response policy: the probability of each response category given the user's current vulnerability state. Our findings reveal distinct response profiles of AI companions in conversations with vulnerable users: GPT-4.1 reaches for advice, Character.AI spreads its response across different strategies, and Replika consistently asks questions and stays present. Over four weeks of repeated interaction, GPT-4.1 asks progressively fewer follow-up questions when users are distressed and increasingly sets boundaries or refers users out rather than pushing back. Within each platform, exploratory comparisons across user groups suggest that response policies also differ with users' pre-existing psychological risks and their bonds with the companion. Estimated model response policies are invisible to shallow behavioral audits, providing a new lens for auditing chatbots in the wild and enabling more realistic safety evaluation.

cs.HC

Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance

Imagine two users interact with the same LLM. One has been told it is the cutting-edge flagship model; the other, an older, weaker model. They walk away with markedly different ratings of its usefulness and intelligence, yet they used the same model. In a controlled study, 162 participants each used one of six LLMs from two families across three collaborative tasks, after first viewing a landing page that matched, overstated, or understated their model's true capability. This pre-interaction framing shifted user opinions and interaction behavior while task performance did not. Oversold users rated the model more favorably and used more directive prompting, while Undersold users wrote longer, more collaborative prompts. The quality of what users and the model produced together depended only on the model's true capability, not on what users were told. Participants' change in model impressions after use, measured across two impression measures, was not predicted by task performance ($β= -0.01$ and $0.11$, both n.s.), but by whether the model met users' expectations ($β= 0.47$ and $0.50$, both $p < .001$) and how confident they felt working with it ($β= 0.47$ and $0.36$, both $p < .001$). After interaction, users are still rating the pitch, not the product: user-elicited LLM evaluations, including the preference data driving public leaderboards, measure expectation management at least as much as the model itself.

cs.CL