Search arXivSearch

arXiv subjects

Boden Moraski

Publications and source records attributed to Boden Moraski.

2 recordsLinked to original sources

What Makes a Terminal-Bench Task Hard? Separating Genuine Hardness from Fake-Hardness on an Adjudicated Agentic Corpus

Frontier benchmarks need tasks that current models cannot solve. But a task that no model solves is not automatically a hard task. The same zero pass rate can come from a real capability gap, but it can also come from missing context, a broken reference solution, infrastructure failure, or a verifier that can be bypassed. In this paper, we study this issue using a frozen Terminal-Bench 3 / Frontier-Bench 0.1 production record with 1,081 pull requests, 639 scored tasks, 28,801 trials, and $105,933 in logged agent spend. We ask what an all-fail task actually certifies. For the 125 tasks with no honest pass, we combine task artifacts, reference-solution runs, empty-solution controls, adversarial trials, trajectories, telemetry, and review records, and apply an ordered validity screen. Only 78 of the 125 tasks survive as certified-unsolved candidates. The remaining tasks include 14 with broken oracles, 8 dominated by infrastructure failures, 4 that are only passable through verifier bypasses, and 21 whose solvability is not certified by the available evidence. Thus, lack of saturation and genuine difficulty are not the same thing. The certified-unsolved label is also narrow: it means that the authored route passed, infrastructure did not dominate, no strict bypass was observed, and all evaluated agents failed. It does not prove intrinsic hardness, verifier completeness, or failure at the intended capability. We further analyze rejected submissions and passing tasks to show that pass rate alone cannot explain why a task is difficult. Overall, our results suggest that frontier benchmarks should report the evidence behind their all-fail tasks before using them as capability claims.

cs.LG

Anthropomorphism in AI Companion Communities: Age, Gender, and Emotional Correlates

Artificial intelligence (AI) systems are increasingly integrated into daily life, with millions now using AI chatbots built on Large Language Models (LLMs) for companionship. Both humanlike AI qualities and user predispositions to anthropomorphize relate to social consequences, such as increased trust, social health benefits, and psychological harms. Populations such as children, older adults, or those with mental health vulnerabilities may be particularly susceptible to anthropomorphism and its detriments, but mixed findings complicate the role of demographics. We used publicly available Reddit data from three popular AI companion subreddits to assess relationships between gender, age, anthropomorphism, and elicited emotions, to better understand how different people perceive and are affected by AI companions. We investigated three questions: How do age and gender relate to anthropomorphization of AI?, How does emotional expression relate to anthropomorphization?, and How do age and gender moderate emotion-anthropomorphization relationships? We found that adults and women anthropomorphize AI chatbots more than teens and men, and that positive emotional expression, particularly joy, is positively associated with anthropomorphization, while neutrality is negatively associated with anthropomorphism. Both relationships were stronger in adults than teens. Our findings suggest that the tendency to anthropomorphize may be more broadly distributed across age groups than previously expected, thereby prompting the reevaluation of existing digital safety norms.

cs.HC