Search arXivSearch

arXiv · 2602.05810

Bifrost: Steering Strategic Trajectories to Bridge Contextual Gaps for Self-Improving Agents

Abstract

Autonomous agents excel in self-improvement through reflection and iterative refinement, which reuse successful task trajectories as in-context examples to assist subsequent reasoning. However, shifting across tasks often introduces a context mismatch. Hence, existing approaches either discard the trajectories or manipulate them using heuristics, leading to a non-negligible fine-tuning cost or unguaranteed performance. To bridge this gap, we reveal a context-trajectory correlation, where shifts of context are highly parallel with shifts of trajectory. Based on this finding, we propose BrIdge contextual gap FoR imprOvised trajectory STeering (Bifrost), a training-free method that leverages context differences to precisely guide the adaptation of previously solved trajectories towards the target task, mitigating the misalignment caused by context shifts. Our trajectory adaptation is conducted at the representation level using agent hidden states, ensuring trajectory transformation accurately aligns with the target context in a shared space. Across diverse benchmarks, Bifrost consistently outperforms existing trajectory reuse and finetuned self-improvement methods, demonstrating that agents can effectively leverage past experiences despite substantial context shifts.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Quan M. Tran, Zhuo Huang, Wenbin Zhang, Bo Han, Koji Yatani, Masashi Sugiyama, Tongliang Liu. 2026-02-05. Bifrost: Steering Strategic Trajectories to Bridge Contextual Gaps for Self-Improving Agents. https://arxiv.org/abs/2602.05810

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Online Regularized Statistical Learning in Reproducing Kernel Hilbert Space With Non-Stationary Data

We study recursive regularized learning algorithms in the reproducing kernel Hilbert space (RKHS) with non-stationary online data streams. We introduce the concept of a random Tikhonov regularization path and decompose the tracking error of the algorithm's output for the regularization path into random difference equations in RKHS. We show that the tracking error vanishes in mean square and almost surely if the regularization path is slowly time-varying. Then, leveraging the monotonicity of inverse operators and the spectral decomposition of compact operators, and introducing the RKHS persistence of excitation condition, we develop a dominated convergence method to prove the mean square and almost sure consistency between the regularization path and the unknown function to be learned. Especially, for independent and non-identically distributed data streams, the mean square and almost sure consistency between the algorithm's output and the unknown function is achieved if the input data's marginal probability measures are slowly time-varying and the average measure over each fixed-length time period is uniformly above a strictly positive finite Borel measure.

cs.LG

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models

Direct Preference Optimization (DPO) improves relative preference by increasing the margin between chosen and rejected responses, but this objective does not specify how probability mass should be redistributed to achieve that margin. Consequently, preference separation can improve even as probability mass collapses away from both preferred and unrelated responses, a phenomenon associated with the squeezing effect. We show that this pathology is localized: continued negative updates become destructive when rejected responses enter low-probability valleys, where further suppression induces harmful redistribution through the coupled softmax geometry. We introduce Gradient-Gated Preference Optimization (Gate-DPO), which uses the current probability geometry of each rejected response to selectively attenuate its contribution, limiting excessive rejected suppression while delaying preference-loss saturation. Gate-DPO is objective-agnostic and composes with existing preference methods. Across four architectures, two preference datasets, and multiple objectives, gating consistently reduces squeezing and improves chosen-response likelihood, while exhibiting robust reductions in rejected-response collapse across architectures and objectives. Mass-dynamics analysis further shows that large preference margins can conceal substantial suppression of unrelated responses, while gating achieves healthier redistribution. Together, our results show that preference margin alone is insufficient to characterize optimization health and that selectively controlling probability dynamics provides a simple, general mechanism for stabilizing preference optimization.

cs.LG

Noise-Debiased Thermodynamic Variance for Local Learning Coefficient Probes

Local learning coefficient (LLC) probes offer a singularity-aware view of neural-network training, but mean-energy methods require a local loss baseline that is ambiguous at transient checkpoints. Thermodynamic variance avoids this input; under mini-batch evaluation, however, direct variance mixes cross-state loss fluctuations with same-state noise. We operationalize this route with the \emph{Shift-Invariant Variance Estimator} (SIVE), which estimates and subtracts the latter component using repeated evaluations. Conditional on any fixed retained path, unclipped SIVE is unbiased for noiseless path variance without requiring MCMC stationarity. The finite-scale diagnostic remains indexed by localization scale $h$---even a locally linear loss has tether-dependent variance---while interpretation as a Real Log Canonical Threshold (RLCT) requires additional stationary low-temperature conditions. Toy experiments recover calibrated finite-scale targets. At the primary localization scale, all five MNIST MLP trajectories exhibit a mid-training trough followed by a rebound in SIVE, while Raw Variance decreases from Epoch 40 to 100 in every trajectory. Across four localization scales, the joint early-drop/late-rise criterion is met in 19 of 20 trajectory--scale pairs. At Epoch 40, the estimated observation-noise correction accounts for $77.5\%$ of Raw Variance. Same-state debiasing thus reveals a reproducible turning structure masked by time-varying observation noise.

cs.LG