Search arXivSearch

arXiv subjects

Daniel Khashabi

Publications and source records attributed to Daniel Khashabi.

4 recordsLinked to original sources

Continual Learning Mechanisms Compose for Long-Horizon Memorization

Language models may need to internalize information that arrives over time and retain it through many subsequent updates. To study this challenge, we introduce long-horizon memorization, a setting in which a model learns 100 query-answer tasks through continual supervised fine-tuning without retaining earlier training examples or receiving task identifiers at inference. Sequential updates cause catastrophic forgetting, and no single continual learning mechanism we evaluate maintains strong retention at this horizon. We hypothesize that mechanisms addressing complementary sources of forgetting will be more effective when composed. We organize these compositions along two design dimensions. Data, function, and weight anchors specify what prior information each update should preserve, while low-rank allocation rules determine where successive updates are retained. To test this hypothesis systematically, we construct three distinct 100-task memorization datasets. We introduce task-level successive halving to search the combinatorial design space and use a factorial experiment to measure individual and interaction effects. Our best method combines all three anchors with merged LoRA, ranks among the top 3 methods in all datasets, and raises average final retention from 1.2% under naive sequential fine-tuning to 34.9%, a 28-fold improvement. The data anchor and merged LoRA provide the largest average gains and interact super-additively on all three datasets. Together, these results show that composing complementary mechanisms substantially improves long-horizon memorization beyond what any individual mechanism achieves.

cs.LG

The Flaw of Averages: Measuring Benchmark-Level Distributional Robustness

Benchmarks are central to measuring progress in language models, but aggregate scores can obscure substantial variation across subdomains, making models appear broadly competent despite concentrated strengths and weaknesses. We study this issue as benchmark-level distributional robustness: whether aggregate scores faithfully reflect performance across benchmark subdomains. We operationalize this notion with benchmark Harmony, an entropy-based measure of how uniformly model performance is distributed across subdomains. Measuring Harmony on 19 language model benchmarks across five model families, we find substantial variation in benchmark-level distributional robustness. Low-Harmony benchmarks are more likely to yield aggregate scores that overstate broad competence, whereas high-Harmony benchmarks provide more representative summaries of model capability. Rebalancing benchmarks by pruning overrepresented subdomains to increase Harmony substantially shifts aggregate scores for low-Harmony benchmarks, but leaves high-Harmony benchmarks comparatively stable. For example, while BoolQ remains comparatively stable as Harmony increases, PubMedQA, which evaluates performance in a medically consequential domain, exhibits substantial, often statistically significant, shifts in aggregate accuracy. Together, these findings show that aggregate scores can misrepresent broad competence when performance is unevenly distributed. We therefore recommend reporting benchmark Harmony alongside aggregate accuracy as a diagnostic of benchmark representativeness when interpreting claims about broad model competence.

cs.CL

Many-Tier Instruction Hierarchy in LLM Agents

Large language model agents receive instructions from many sources-system messages, user prompts, tool outputs, other agents, and more-each carrying different levels of trust and authority. When these instructions conflict, agents must reliably follow the highest-privilege instruction to remain safe and effective. The dominant paradigm, instruction hierarchy (IH), assumes a fixed, small set of privilege levels (typically fewer than five) defined by rigid role labels (e.g., system > user). This is inadequate for real-world agentic settings, where conflicts can arise across far more sources and contexts. In this work, we propose Many-Tier Instruction Hierarchy (ManyIH), a paradigm for resolving instruction conflicts among instructions with arbitrarily many privilege levels. We introduce ManyIH-Bench, the first benchmark for ManyIH. ManyIH-Bench requires models to navigate up to 12 levels of conflicting instructions with varying privileges, comprising 853 agentic tasks (427 coding and 426 instruction-following). ManyIH-Bench composes constraints developed by LLMs and verified by humans to create realistic and difficult test cases spanning 46 real-world agents. Our experiments show that even the current frontier models perform poorly (~40% accuracy) when instruction conflict scales. This work underscores the urgent need for methods that explicitly target fine-grained, scalable instruction conflict resolution in agentic settings.

cs.CL

Can Coding Agents Reproduce Findings in Computational Materials Science?

Large language models are increasingly deployed as autonomous coding agents and have achieved remarkably strong performance on software engineering benchmarks. However, it is unclear whether such success transfers to computational scientific workflows, where tasks require not only strong coding ability, but also the ability to navigate complex, domain-specific procedures and to interpret results in the context of scientific claims. To address this question, we present AutoMat, a benchmark for evaluating LLM-based agents' ability to reproduce claims from computational materials science. AutoMat poses three interrelated challenges: recovering underspecified computational procedures, navigating specialized toolchains, and determining whether the resulting evidence supports a claim. By working closely with subject matter experts, we curate a set of claims from real materials science papers to test whether coding agents can recover and execute the end-to-end workflow needed to support (or undermine) such claims. We then evaluate multiple representative coding agent settings across several foundation models. Our results show that current LLM-based agents obtain low overall success rates on AutoMat, with the best-performing setting achieving a success rate of only 53%. Error analysis further reveals that agents perform worst when workflows must be reconstructed from paper text alone and that they fail primarily due to incomplete procedures, methodological deviations, and execution fragility. Taken together, these findings position AutoMat as both a benchmark for computational scientific reproducibility and a tool for diagnosing the current limitations of agentic systems in AI-for-science settings.

cs.CL