Search arXiv⌕ Search

arXiv subjects

Dongha Lee

Publications and source records attributed to Dongha Lee.

At least 19 recordsLinked to original sources

JudgeCast: Time Series Forecasting with Experience-Informed Covariate Judgements

Covariate effects vary across contexts and shift over time, requiring forecasters to assess how to use them for each forecasting context. As forecasting proceeds, observations for earlier forecasts become available, providing feedback on past covariate use for subsequent forecasts. However, when multiple covariates act together, the forecast error reveals the numerical discrepancy from the observation but not how the covariates should have been used. We introduce JudgeCast, an experience-based framework for time series forecasting with covariates. Following the judgmental adjustment practice, a frozen TSFM provides the base forecast, while a frozen LLM uses the current context and relevant experience to adjust it. Within the adjustment, assessing covariate effects and determining the numerical adjustment serve distinct roles, so JudgeCast first forms explicit covariate-wise judgments and then determines the adjustment. After observation, JudgeCast uses the observed residual of the base forecast to reconstruct alternative judgments and evaluates the original and alternatives through their resulting adjustments. The best-performing decision is selected and retained as validated experience for subsequent forecasts. Across diverse real-world datasets, JudgeCast outperforms strong baselines. Ablations show that explicit covariate-wise judgment can improve forecast-time adjustment, while residual-guided experience construction yields more reliable forecasting gains than retaining raw decisions as experience.

cs.LG↗

In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners

Stronger learning signals do not necessarily produce a stronger small reasoning model. Off-policy supervised fine-tuning (SFT) supplies stronger traces that may be incompatible with the student. On-policy post-training instead remains constrained by the quality of the student's own trajectories: reinforcement learning (RL) provides little learning signal because a small student rarely produces a correct rollout. Consequently, effective post-training must supply stronger reasoning in a form the student can learn from. We hypothesize that a small number of tokens with very low probability under the student can make an otherwise strong reasoning trace difficult to learn. To test this hypothesis, we introduce Interleaved-Policy Distillation (IPD), which generates SFT data balancing teacher guidance and student compatibility. During generation, the student replaces teacher proposals that have low probability under its own policy, allowing the teacher to continue reasoning from a prefix the student can support. IPD improves reasoning across model families, datasets, and varying degrees of student intervention; for example, it raises the seven-benchmark average from 24.55 to 29.72 on Qwen3-0.6B/s1K. It outperforms baselines with comparable training budgets and broadens problem-solving coverage, while ablations show that filtering traces or masking losses cannot match the gains from revising trajectories during generation. Analyses link transfer failures to the rare tokens the student finds least likely and show that the source of the tokens also matters. Small models reason better when they learn strong solutions in a form they can reach in their own words.

cs.CL↗

Hindsight Compacts but Does Not Repair: Rethinking On-Policy Self-Distillation in Reasoning Models

On-Policy Self-Distillation (OPSD) has emerged as a promising post-training method: requiring only privileged context, it enables token-level credit assignment from a self-teacher with hindsight, yielding higher accuracy with shorter responses. However, in thinking-enabled mathematical reasoning, where solving a problem takes long reasoning traces that explore, reflect, and backtrack, its length reductions persist but its accuracy gains largely do not. Does OPSD's hindsight signal, then, still repair failed trajectories, or does it only compact viable ones? We hypothesize compaction: in long traces, hindsight reveals which steps were unnecessary more readily than which steps would have repaired the solution. To test this, we apply OPSD separately to correct and incorrect rollouts. Training only on correct rollouts shortens responses by 18-29% while largely preserving accuracy, whereas training only on incorrect rollouts degrades it. The accuracy gap holds across three models, six benchmarks, and three seeds, and under a same-prompt control for difficulty. Neither branch raises the pass@k ceiling, and both suppress exploration and reflection markers, which viable traces can spare but failed ones need. Changing the divergence, enriching or reinjecting the privileged context, and training longer only move OPSD along the same accuracy-length tradeoff. Hindsight compacts reasoning the model can already produce but does not repair reasoning it cannot.

cs.AI↗

ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis

Learning from experience in LLM agents has become a key paradigm for developing self-evolving agents that continuously learn and expand their capabilities. Within this paradigm, synthesizing the agent skill has emerged as a promising solution for transforming accumulated experience into reusable procedural knowledge, serving as an important layer for the harness system that supplies agents at runtime. Despite its potential, existing approaches largely abstract past experience into fixed procedural knowledge before downstream demands are known, which risks discarding knowledge that later becomes critical while retaining instance-specific details irrelevant to future tasks. In this paper, we reframe agent skill synthesis as a dynamic navigation problem over past experience, where agents actively explore accumulated trajectories on demand for the current task with targeted and fine-grained access to experience knowledge. To this end, we propose ExpVoyager, a novel framework in which a skill curator navigates raw experience across different views and resolutions, continually identifying reusable procedural knowledge from what it observes while tracking remaining knowledge needs that guide where to navigate next. Extensive experiments demonstrate both the effectiveness and versatility of ExpVoyager, showing consistent improvements in downstream task performance, continual gains as the experience space scales, and practical compatibility with existing skills under efficient experience access.

cs.CL↗

Self-Evolving Search Index

Information retrieval is increasingly important as LLM agents tackle complex tasks involving diverse information needs. Because retrieval relies on an index that represents each document through index keys, retrieval quality depends heavily on how effectively these keys expose the knowledge contained in each document. However, effective index representations vary across retrieval environments, making it difficult for any fixed optimization strategy to perform consistently. Yet evolving an index to its retrieval environment remains largely human-driven, requiring humans to diagnose retrieval failures, refine the optimization strategy, and reprocess the index accordingly. We propose SELF-INDEX, a framework that enables an index to self-evolve without human intervention. Its Optimizer autonomously diagnoses retrieval shortfalls, selectively revises the responsible index keys, and validates each revision before updating the index. Beyond reacting to observed retrieval demands, SELF-INDEX proactively explores additional demands through a Query Simulator, allowing the index to evolve beyond the queries already available for optimization. Across diverse corpora and retrievers, SELF-INDEX consistently improves retrieval performance while outperforming existing index optimization methods. We further show that these benefits extend to downstream applications, improving the effectiveness and efficiency of search agents and helping agent memory systems retrieve useful past interactions.

cs.IR↗

A Ruh-Vilms theorem for hypersurfaces in Weitzenböck geometry

A well-known theorem by Ruh and Vilms states that the Laplacian of the Gauss map for a smooth immersion into Euclidean space is given by the covariant derivative of the mean curvature vector field. For hypersurfaces, this implies that the Gauss map is harmonic iff the mean curvature is constant. In this paper, we extend this result to hypersurfaces in Weitzenböck geometry. While Riemannian geometry corresponds to the curved geometry without torsion, Weitzenböck geometry is a flat geometry with torsion. They represent two opposite extremes of Riemann-Cartan geometry.

math.DG↗

Kirigami Meta-Sheet for Enhanced Impact Absorption

Impact absorbers based on mechanical metamaterials often use bulky, vertically stacked architectures, limiting large area deployment and scalable manufacturing. Here, we propose a kirigami meta-sheet as a planar absorber that uses transitions between positive and negative stiffness regimes rather than sacrificial crushing. Guided by an analysis of a simple mass-spring-damper model, we program the stiffness of kirigami meta-sheets through the hinge ratio connecting the unit cells. Quasistatic indentation experiments confirm that the meta-sheet with a low hinge ratio most clearly exhibits the negative stiffness transition. The drop-tower tests show that it reduces rebound, decreases the first impact force, and increases dissipation. Unlike polyethylene mesh and styrofoam, this kirigami meta-sheet is shown to be effective in protecting a falling egg. Its planar geometry enables area scaling by tiling and is compatible with sheet level manufacturing routes such as cutting, molding, and lamination, establishing kirigami meta-sheets as practical impact absorbers.

cs.CE↗

Activation Boundary Matching: Task-Informed Initialization for Low-Rank Adaptation

Low-Rank Adaptation (LoRA) is highly sensitive to initialization, yet existing schemes construct the initial subspace from statistics at the pretrained point, capturing pre-adaptation geometry rather than how the adapter must move during learning. We examine the early adaptation trajectory and uncover a temporal asymmetry: task-induced activation boundaries---the signs of layer-wise pre-activations---recover markedly faster than activation values or effective low-rank updates and become reusable well before the probe adapter converges. Motivated by this, we propose Activation Boundary Matching for LoRA (ABM-LoRA), which briefly trains a standard adapter as an early task probe, then uses its pre-activation signs as layer-wise targets for a fresh adapter under a margin-based hinge objective. Both are discarded before otherwise unchanged downstream fine-tuning. We show that this boundary-supervision lifting can expose LoRA directions attenuated or locally unobservable through the downstream Jacobian, and that what transfers is the boundary side---not the exact activation value or the output-level signal. Because a partial boundary trace suffices, ABM-LoRA recovers most of the benefit of a fully trained reference at a fraction of the overhead, requiring only probe forward passes. It improves over standard LoRA on T5-base/GLUE, ConvNeXt-T and Swin-T fine-grained classification, and instruction tuning with Qwen2.5-1.5B and LLaMA2-7B, and matches or surpasses SVD- and gradient-based initializers without their preprocessing.

cs.CV↗

SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization

Search-Augmented Generative Engines (SAGE) have emerged as a new paradigm for information access, bridging web-scale retrieval with generative capabilities to deliver synthesized answers. This shift has fundamentally reshaped how web content gains exposure online, giving rise to Search-Augmented Generative Engine Optimization (SAGEO), the practice of optimizing web documents to improve their visibility in AI-generated responses. Despite growing interest, no evaluation environment currently supports comprehensive investigation of SAGEO. Specifically, existing benchmarks lack end-to-end visibility evaluation of optimization strategies, operating on pre-determined candidate documents that abstract away retrieval and reranking preceding generation. Moreover, existing benchmarks discard structural information (e.g., schema markup) present in real web documents, overlooking the rich signals that search systems actively leverage in practice. Motivated by these gaps, we introduce SAGEO Arena, a realistic and reproducible environment for stage-level SAGEO analysis. Our objective is to jointly target search-oriented optimization (SEO) and generation-centric optimization (GEO). To achieve this, we integrate a full generative search pipeline over a large-scale corpus of web documents with rich structural information. Our findings reveal that existing approaches remain largely impractical under realistic conditions and often degrade performance in retrieval and reranking. We also find that structural information helps mitigate these limitations, and that effective SAGEO requires tailoring optimization to each pipeline stage. Overall, our benchmark paves the way for realistic SAGEO evaluation and optimization beyond simplified settings.

cs.IR↗

Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding

User requests serve as research specifications for deep research agents, shaping what evidence to seek and how to synthesize it. In personalized deep research, these specifications must additionally reflect user goals, constraints, preferences, and evaluation criteria. User context can be incorporated either within the deep research pipeline or into the research specification provided as its input. We focus on the latter, refining the user request into a personalized research specification before passing it to an unchanged deep research agent. This requires resolving three coupled decisions: which framing factors are relevant, whether the available user context sufficiently supports them, and whether to retrieve user memory, ask the user, or stop and refine the query. For training, G-STEER organizes framing factors as elicitation targets in an Intent Elicitation Graph that captures their dependencies. It learns a clarification policy from graph-scaffolded trajectories spanning diverse factor dependencies and evidence conditions. The policy produces a refined query while balancing target coverage against the costs of evidence acquisition. Experiments show that G-STEER achieves the strongest overall weighted target coverage and the highest downstream report personalization across both evaluated DRAs, while asking roughly one third as many user questions as a strong clarification baseline.

cs.AI↗

WatchLens: A Configurable Platform for Online Video Recommendation Experiments

Studying how video recommender systems shape user behavior requires online experiments that link playback behavior with the recommendation conditions that produced it. Existing user-study infrastructure provides one or the other, but not both within a single experimentation workflow. We present WatchLens, an open-source platform that fills this gap. WatchLens adopts a modular architecture in which user interfaces, content sources, and recommendation policies are independently configurable, with policies assignable separately to the feed and the watch page, while a standardized logging layer attaches the recommendation policy and ranking position to every event at recording time. This design enables researchers to analyze how recommendation policies and ranking positions shape downstream playback behavior, session continuation, and navigation between the feed and the watch page, with the linkage between policy and outcome available in each event rather than reconstructed afterwards. We demonstrate WatchLens with a short-form video case study that holds the interface, feed policy, and content pool constant while varying only the watch-page policy, showing how the platform supports session-level comparison of recommendation effects on real viewing behavior. WatchLens is released as a publicly available, single-server deployable system for reproducible online video recommendation research.

cs.IR↗

PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization

A significant hurdle for current LLMs is the execution of complex, multi-stage tasks. Group Relative Policy Optimization (GRPO) has been emerging as a leading choice, but its reliance on sparse outcome rewards severely limits credit assignment across intermediate steps. Existing remedies such as running full rollouts to assign step-level advantages, calling external LLM judges at each step, or computing intrinsic rewards that require ground-truth answers at every evaluation introduce significant costs or practical constraints. We hypothesize that internal correctness probing over LLM hidden states can be repurposed as a step-level reward signal, potentially addressing all of these limitations at once. However, existing probing research assumes clean inputs, and we first show that this assumption breaks down in multi-step settings: hidden-state probes degrade severely under prefix contamination tracking coherence with the (possibly corrupted) prefix rather than grounded correctness, while attention-based features remain robust to contamination but underperform on clean prefixes. Building on this complementary relationship, we propose the Prefix-Aware Internal Reward (PAIR), a two-stage model with a frozen hidden-state probe estimating belief-consistency and a lightweight attention-based head correcting it toward grounded correctness. Experimental results show that PAIR achieves the highest AUROC on contaminated trajectories while operating at negligible inference cost, enabling dense step-level reward signals for GRPO training without external model calls, ground-truth dependencies, or full-trajectory rollouts.

cs.AI↗

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist

Recent approaches in personalized reward modeling have primarily focused on leveraging user interaction history to align model judgments with individual preferences. However, existing approaches largely treat user context as a static or implicit conditioning signal, failing to capture the dynamic and multi-faceted nature of human judgment. In this paper, we propose P-Check, a novel personalized reward modeling framework, designed to train a plug-and-play checklist generator that synthesizes dynamic evaluation criteria for guiding the reward prediction. To better align these checklists with personalized nuances, we introduce Preference-Contrastive Criterion Weighting, a training strategy that assigns saliency scores to criteria based on their discriminative power for personalized judgment. We conduct extensive experiments and demonstrate that P-Check not only improves reward accuracy but also enhances downstream personalized generation, and remains robust in OOD scenarios.

cs.CL↗

Improving Scientific Document Retrieval with Academic Concept Index

Adapting general-domain retrievers to scientific domains is challenging due to the scarcity of large-scale domain-specific relevance annotations and the substantial mismatch in vocabulary and information needs. Recent approaches address these issues through two independent directions that leverage large language models (LLMs): (1) generating synthetic queries for fine-tuning, and (2) generating auxiliary contexts to support relevance matching. However, both directions overlook the diverse academic concepts embedded within scientific documents, often producing redundant or conceptually narrow queries and contexts. To address this limitation, we introduce an academic concept index, which extracts key concepts from papers and organizes them guided by an academic taxonomy. This structured index serves as a foundation for improving both directions. First, we enhance the synthetic query generation with concept coverage-based generation (CCQGen), which adaptively conditions LLMs on uncovered concepts to generate complementary queries with broader concept coverage. Second, we strengthen the context augmentation with concept-focused auxiliary contexts (CCExpand), which leverages a set of document snippets that serve as concise responses to the concept-aware CCQGen queries. Extensive experiments show that incorporating the academic concept index into both query generation and context augmentation leads to higher-quality queries, better conceptual alignment, and improved retrieval performance.

cs.IR↗

Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History

Large language models have advanced web agents, yet current agents lack personalization capabilities. Since users rarely specify every detail of their intent, practical web agents must be able to interpret ambiguous queries by inferring user preferences and contexts. To address this challenge, we present Persona2Web, the first benchmark for evaluating personalized web agents on the real open web, built upon the clarify-to-personalize principle, which requires agents to resolve ambiguity based on user history rather than relying on explicit instructions. Persona2Web consists of: (1) user histories that reveal preferences implicitly over long time spans, (2) ambiguous queries that require agents to infer implicit user preferences, and (3) a reasoning-aware evaluation framework that enables fine-grained assessment of personalization. We conduct extensive experiments across various agent architectures, backbone models, history access schemes, and queries with varying ambiguity levels, revealing key challenges in personalized web agent behavior. For reproducibility, our codes and datasets are publicly available at https://serin-kimm.github.io/Persona2Web/.

cs.CL↗

Verified Misguidance: Measuring Structural Citation Failures in Search-Augmented LLMs

Users of search-augmented LLMs rely on citations as evidence that responses are grounded in real sources, and rarely verify the cited pages themselves. Millions of queries per day now pass through these systems, making citation quality a silent determinant of whether users are informed or misled-yet existing benchmarks each address one facet in isolation, leaving the joint structure that determines citation trustworthiness unmeasured. We construct CITETRACE, a large-scale dataset that traces the full citation chain from user query through retrieved source to generated answer: 11,200 real-world queries from 28 communities paired with 112,000 responses from ten models across five providers, yielding 761,495 evaluable citation pairs. We design a three-dimension evaluation framework that scores each citation on intent-purpose alignment, source suitability, and answer-source fidelity, using expert-validated predefined matrices and a five-level fidelity rubric; the framework applies to any system that produces citation-bearing responses. Applying this framework at scale, we identify a systematic pattern we call VERIFIED MISGUIDANCE (VM): models cite real, accessible sources yet fail along one or more dimensions, producing a fidelity-suitability trade-off in which faithful models select inappropriate sources and vice versa. Across our pool, 30.6% of citations distort their sources and 27.1% originate from domain-inappropriate sources; at the response level, up to 96% of users encounter at least one structurally misleading citation. Provider-level differences explain 88-96% of citation-quality variance, suggesting that source selection is governed more by factors beyond individual model capability than by the LLMs themselves. Together, CITETRACE and its evaluation framework provide the first resource for diagnosing structural citation failures in deployed search-augmented systems.

cs.DL↗

BESPOKE: Benchmark for Search-Augmented Large Language Model Personalization via Diagnostic Feedback

Search-augmented large language models (LLMs) have advanced information-seeking tasks by integrating retrieval into generation, reducing users' cognitive burden compared to traditional search systems. Yet they remain insufficient for fully addressing diverse user needs, which requires recognizing how the same query can reflect different intents across users and delivering information in preferred forms. While recent systems such as ChatGPT and Gemini attempt personalization by leveraging user histories, systematic evaluation of such personalization is under-explored. To address this gap, we propose BESPOKE, the realistic benchmark for evaluating personalization in search-augmented LLMs. BESPOKE is designed to be both realistic, by collecting authentic chat and search histories directly from humans, and diagnostic, by pairing responses with fine-grained preference scores and feedback. The benchmark is constructed through long-term, deeply engaged human annotation, where human annotators contributed their own histories, authored queries with detailed information needs, and evaluated responses with scores and diagnostic feedback. Leveraging BESPOKE, we conduct systematic analyses that reveal key requirements for effective personalization in information-seeking tasks, providing a foundation for fine-grained evaluation of personalized search-augmented LLMs. Our code and data are available at https://augustinlib.github.io/BESPOKE/.

cs.CL↗

Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents

Although recent tool-augmented benchmarks involve complex requests, evaluation remains limited to answer matching, neglecting critical trajectory aspects like efficiency, hallucination, and adaptivity. The most straightforward method for evaluation is to compare an agent's trajectory with the ground-truth, but annotating all valid ground-truth trajectories is prohibitively expensive. In this manner, we introduce TRACE, a reference-free framework for the multi-dimensional evaluation of tool-augmented LLMs. By incorporating an evidence bank which accumulates knowledge from preceding steps, TRACE assesses an agent's reasoning trajectory effectively. To validate our framework, we develop a new meta-evaluation dataset with diverse and flawed trajectories, each labeled with multi-faceted performance scores. Our results confirm that TRACE accurately evaluates complex trajectories even with small open-source LLMs. Furthermore, we apply our method to evaluate the trajectories that agents produce while solving tool-augmented tasks, presenting previously unreported observations and their corresponding insights.

cs.AI↗