Search arXiv⌕ Search

arXiv · 2609.33321

Separating Memory and Workflow Effects in Predicting Individual Answers

Abstract

Personalized language agents choose both what to remember about a person and how to use that memory. We separate these choices when predicting unseen answers to known interview questions. On 1,768 tasks from 188 people, a concrete memory built from a verified interview prefix outscores a trait description by 0.0158 (95% whole-person interval [0.0044, 0.0271]). Crossing both memories with one-shot generation and three-answer fusion, fusion lowers concrete-memory scores by 0.0123 ([-0.0189, -0.0056]); prompted and trained selectors do not detectably beat a random candidate. One call on the longer, unrewritten source record outscores every memory condition. Under a limited context budget, OwnWords retrieves the person's sentences with BM25 and answers in one call. It outperforms the written memory on 500 people outside the benchmark (+0.0127, [+0.0037, +0.0217]; an earlier held-out test was inconclusive) and across four budgets on 300 people (mean +0.0218, [+0.0138, +0.0298]), with the latter result repeated on 114 people. It does not detectably outperform recency truncation. These results compare evidence-construction procedures; they do not isolate the effect of verbatim wording. On Twin-2K-500, OwnWords predicts ordinal survey answers more closely than the written memory, but does not improve exact-choice accuracy and lowers it in one of two samples. Interview scores use a model-based content rubric without human ratings, and the original benchmark's participants were seen during development. These results characterize the tested procedures, not a general human-prediction ceiling.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tianzhu Qin, Leo Yang Yang, Lee Wei Jun, Kun Chen, Ramit Debnath, Davin Youchao Dong. 2026-09-27. Separating Memory and Workflow Effects in Predicting Individual Answers. https://arxiv.org/abs/2609.33321

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Beyond Judgment: Exploring Large Language Models as Non-Judgmental Support for Maternal Mental Health

In the age of Large Language Models (LLMs), much work has already been done on how LLMs support medication advice and serve as information providers; however, how mothers use these tools for emotional and informational support to avoid social judgment remains underexplored. This study conducted a 10-day mixed-methods exploratory survey ($N=107$) to investigate how mothers use LLMs as a non-judgmental resource for emotional support and regulation, and for situational reassurance. Our findings show that mothers are asking LLMs various questions about childcare to reassure themselves and avoid judgment, particularly around childcare decisions, maternal guilt, and late-night caregiving. Open-ended responses also show that mothers are comfortable with LLMs because they do not have to think about social consequences or judgment. Although mothers use LLMs for quick information or reassurance to avoid judgment, over half of the participants value human warmth more than LLMs; however, a significant minority, especially those in joint families, consider LLMs to avoid human judgment. These findings help understand how LLMs can be framed as low-risk interaction support rather than a replacement for human support, and highlight the role of social context in shaping emotional technology use.

cs.HC↗

Avoiding Social Judgment, Seeking Privacy: Investigating why Mothers Shift from Facebook Groups to Large Language Models

Social media platforms, especially Facebook parenting groups, have long been used as informal support networks for mothers seeking advice and reassurance. However, growing concerns about social judgment, privacy exposure, and unreliable information are changing how mothers seek help. This exploratory mixed-method study examines why mothers are moving from Facebook parenting groups to large language models such as ChatGPT and Gemini. We conducted a cross-sectional online survey of 109 mothers. Results show that 41.3% of participants avoided Facebook parenting groups because they expected judgment from others. This difference was statistically significant across location and family structure. Mothers living in their home country and those in joint families were more likely to avoid Facebook groups. Qualitative findings revealed three themes: social judgment and exposure, LLMs as safe and private spaces, and quick and structured support. Participants described LLMs as immediate, emotionally safe, and reliable alternatives that reduce social risk when asking for help. Rather than replacing human support, LLMs appear to fill emotional and practical gaps within existing support systems. These findings show a change in maternal digital support and highlight the need to design LLM systems that support both information and emotional safety.

cs.HC↗

Learning to Assign Prediction Tasks to Agents with Capacity Constraints

We address the problem of learning to assign prediction tasks to one agent from a set of available agents, including human decision-makers and AI models. We focus on sequential learning of agent expertise and assignment policies where each agent is constrained to handle a fraction of tasks. We provide a general theoretical characterization of this problem in terms of agent capacities, differences in agent expertise, and task context. We then develop a framework of sequential explore-exploit policy-learning algorithms that seek to maximize overall performance. Experimental results over a variety of tabular, image, and text prediction tasks demonstrate systematic gains from our policy-learning algorithms relative to non-contextual baselines across different types of agents, including LLMs and humans.

cs.HC↗