Search arXivSearch

arXiv subjects

Yuheng Wu

Publications and source records attributed to Yuheng Wu.

At least 19 recordsLinked to original sources

NavSight in the Wild: Understanding Real-World Use of a Mobile Augmented Reality Application for People with Low Vision in Outdoor Navigation

The ability to navigate outdoors safely and independently is crucial yet challenging for people with low vision (PLV). While various augmented reality (AR) systems for low vision have been designed and evaluated in ideal lab environments, no research has investigated their real-world feasibility and challenges. We present NavSight, a mobile AR application that assists PLV in outdoor navigation by recognizing important outdoor objects (e.g., curb, vehicle) and rendering real-time visual augmentations. Through a seven-day diary study with 12 PLV in real-world settings, we characterize the impact of NavSight on scene perception, users' configuration strategies on what objects to augment and how to augment them across scenarios, how users made sense of and responded to recognition errors, and the social acceptability of using NavSight in public. We further identify environmental factors affecting recognition, such as weather conditions, lighting and shadows, and nonstandard road markings and textures, as well as usability issues in daily use. We discuss these real-world challenges and derive design implications for future AI-powered assistive AR systems for outdoor use.

cs.HC

Exact Expressions of Entropy for Classical Non-interacting Many-body Systems

In the thermodynamic limit, the equilibrium state of a many-body system can be characterized by three pairs of conjugate thermodynamic variables: $E/T,V/P,N/\mu$. In this limit, the thermodynamic properties in all ensembles are equivalent up to the leading order of $E,V,N$. However, for systems of finite size, this ensemble equivalence is no longer exact, and the thermodynamic properties may differ substantially among ensembles. To quantify these finite-size effects rigorously, it is desirable to develop a universal ensemble theory applicable to systems of arbitrary size, providing exact expressions for entropy and, thereby, giving rise to the precise value of all equilibrium thermodynamic quantities. In this work, we propose a theory that determines the exact entropy expressions for classical non-interacting many-body systems of arbitrary size across all statistical ensembles, based on only two postulates: \textbf{stationarity}, requiring that the physical laws be invariant under time translation, and \textbf{unbiasedness}, requiring that the equilibrium mixed state maximize the entropy subject to the prescribed constraints. Moreover, we show that the entropy expressions obtained in different ensembles converge to the common asymptotic form $S \asymp \ln\!\left( \left( \frac{4\pi m e E}{3N} \right)^{3N/2} \cdot \left(\frac{V}{N}\right)^N \right)+N$, consistent with the predictions of the large deviation theory.

cond-mat.stat-mech

Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks

Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and Reviewer execute bounded missions over durable project state. Argus separates stable user intent from operational objectives, constraints, and verification criteria, and admits memories, skills, procedures, verifiers, routing decisions, and rejected routes only after role-owned review and, when available, task-native verification. Model weights remain fixed; self-evolution occurs through persistent runtime state and control policy, with autonomous execution between operator-owned escalation points. Across seven GPT-5.5 benchmark arenas, Argus achieves about 78% on SWE-Bench Pro versus 59% for Direct Copilot while using 1.41 times the aggregate tokens. After verification-gated self-evolution, mature SWE-Bench waves use 21% fewer solve-input tokens and 15% less active workflow time per task than startup waves, while recording 34 verifier recoveries and 22 strict review-loop rescues. Argus also reaches 76.8% on AARRI-Bench and a 28.0-point gap on mathematical data synthesis, with competitive GPU-kernel and language-model-training results. Beyond benchmarks, an optimized RWKV6 kernel was merged upstream; a multi-day mathematics campaign retained falsified routes and proof-backed frontier updates; and six paper pipelines completed 254 missions with 16 stage rollbacks. These results show that a fixed-weight, self-evolving harness can revise, recover, and accumulate verified approaches while producing structured trajectories for future supervised and reinforcement learning.

cs.AI

Visko Orbis 1.0: A Live Model for Real-Time Interactive Long Video Generation

We present Visko Orbis 1.0, a Live Model for real-time, interactive long video generation. Users can change the prompt at any moment during generation, and the update becomes visible in real time. Visko Orbis 1.0 supports long-form text-to-video, image-to-video, and video continuation, with multilingual prompts and prompt switching while generation is in progress. A bounded multi-scale memory preserves subjects, scenes, and style across chunks, sustaining hour-scale rollouts without evident quality or color drift. The generator is factorized causally in time, matching the causal structure of physical dynamics, and is aligned with a latent world-model reward for predictive consistency. Built on a distilled chunk-wise streaming generator and a streaming video upscaler, Visko Orbis 1.0 delivers 4K video generation at 24 FPS in real time, using an optimized GPU serving engine. In quantitative evaluations, Visko Orbis 1.0 achieves the best DOVER aesthetic and technical scores and the best VideoAlign visual and motion quality, and leads three physical-plausibility protocols (VideoPhy-2, Physics-IQ, and VBench-2.0 Physics); in long-form Arena comparisons, it obtains the highest overall-preference and temporal-stability ratings among all the state-of-the-art real-time interactive video generation systems.

cs.CV

A coupled-channel quark model study of possible $\Xi_{cc}^{(*)} K^{(*)}$ molecular states

Inspired by the recent experimental discovery of doubly charmed baryons, we investigate the possible $\Xi_{cc}^{(*)}K^{(*)}$ molecular systems within the framework of the quark delocalization color screening model. The energy spectra and scattering processes of the relevant baryon-meson systems are investigated to explore the dynamical properties of the possible molecular states. The spectrum calculations predict three bound states, namely the $I(J^P)=0(1/2^{-})$ $\Xi_{cc}K$, the $I(J^P)=0(3/2^{-})$ $\Xi_{cc}^{*}K$, and the $I(J^P)=0(5/2^{-})$ $\Xi_{cc}^{*}K^{*}$ molecular states. The scattering phase shift analysis further confirms two $\Xi_{cc}K^{*}$ resonance states with $I(J^P)=0(1/2^{-})$ and $0(3/2^{-})$, which originate from quasi-bound states through channel coupling. In particular, the $I(J^P)=0(1/2^{-})$ $\Xi_{cc}K$ bound state is consistent with previous theoretical studies, making it one of the most promising candidates for future experimental searches.

hep-ph

What to Distinguish and How? Opportunities and Challenges of Augmenting Multiple, Cluttered Objects in Complex Scenes for People with Low Vision

People with low vision (PLV) struggle to perceive complex scenes like busy kitchens and crowded streets, which contain many objects, visual clutter, and dynamic elements. Prior AR systems for low vision either enhance low-level visual features or augment task-relevant objects for single tasks in simple settings, leaving multi-object augmentation in complex scenes underexplored. Informed by a formative study characterizing important objects and their perceived importance for PLV, we built SceneGlance, a wearable AR system that recognizes important objects and visually distinguishes them by importance level. Through a controlled lab study with 12 PLV in a mock-up kitchen scene and a free-form think-aloud study with 13 PLV navigating an outdoor route, we found that AR distinction on object importance shifted PLV's attention toward objects of higher importance, and supported perception strategies such as building mental snapshots from the augmentation distribution and hierarchical scanning by importance. However, this attention shift came with a tradeoff of reduced overall scene recall. The studies also surfaced challenges posed by AR augmentations in complex scenes, such as adjacent augmentations blending or interfering with each other, yielding design implications for more practical AR vision enhancement systems in the complex real world.

cs.HC

Last But Not Least: Boundary Attention CalibratiON for Multimodal KV Cache Compression

Multimodal Large Language Models (MLLMs) achieve strong vision-language reasoning but incur large KV caches and high decoding latency with long visual contexts. Existing compression methods rely on observation window attention for stable token importance estimation, yet this aggregation can dilute sparse critical evidence and discard answer-relevant tokens under aggressive compression. We identify last query attention as a complementary signal for recovering such evidence, though its irrelevant signals may introduce additional noise. We propose BACON, a plug-and-play method that calibrates observation window attention with last query evidence while suppressing noise through intra-layer coherence and inter-layer persistence. Across diverse benchmarks, models, budgets, and compression methods, BACON improves multimodal KV-cache compression by 7.5% on average under the most aggressive budget, with gains up to 30.9%.

cs.CV

Beyond Uniform Token-Level Trust Region in LLM Reinforcement Learning

Reinforcement learning with verifiable rewards (RLVR) has become standard for improving LLM reasoning. However, existing PPO-style trust-region mechanisms remain position-agnostic by enforcing uniform thresholds across all tokens independently. This pointwise treatment conflicts with autoregressive generation in two critical ways. First, uniform thresholds ignore autoregressive asymmetry. Early-stage deviations produce compounding sequence-level drift, causing static thresholds to under-regulate early divergence and excessively constrain late-stage exploration. Second, evaluating token-level divergence in isolation overlooks cumulative prefix drift, granting the same divergence allowance regardless of how far the conditioning history has already deviated from the rollout policy. To address this limitation, we propose CPPO (Cumulative Prefix-divergence Policy Optimization), a token-level masking rule that aligns updates with a finite-horizon policy-improvement bound via two coupled mechanisms. First, a position-weighted threshold imposes stricter limits at early positions whose effects persist longer, relaxing constraints for late-stage tokens. Second, a cumulative prefix budget tracks historical deviations, dynamically restricting further token-level deviation to prevent compounding errors along the prefix. Empirically, CPPO enhances training stability and significantly improves reasoning accuracy across various model scales.

cs.LG

Investigation of fully heavy tetraquark within chiral quark model

In the framework of the Chiral quark model (ChQM), we investigate the fully charmed and fully bottomed tetraquark with $J^{PC}=2^{++}$ including two structures: $Q\bar{Q}-Q\bar{Q}$ and $QQ-\bar{Q}\bar{Q}$. The bound-state calculation shows that there is no bound state in either $cc\bar{c}\bar{c}$ or $bb\bar{b}\bar{b}$ systems. However, by using the real-scaling method, some resonance states are obtained. For the $cc\bar{c}\bar{c}$ system, when the channel-coupling includes only three $S$-wave channels, two resonant states are obtained: one with a mass around $7002$ MeV and decay width near $54$ MeV, and another with a mass around $7227$ MeV and a decay width near $66$ MeV. The former can be regarded as a candidate for the $X(6900)$, and the latter can be considered as a candidate for the $X(7200)$. Upon adding the $\chi_{c0}\chi_{c2}$, $\chi_{c1}\chi_{c1}$, $\chi_{c1}\chi_{c2}$, $\chi_{c2}\chi_{c2}$ channels, both resonant states still remain. For the $bb\bar{b}\bar{b}$ system, only one resonant state is obtained, regardless of whether the four channels composition of the excited mesons are included or excluded. The mass and width of this resonant state are around $19743$ MeV and $67$ MeV, respectively. We suggest that future experiments search for the possible resonance state in the invariant mass spectrum of $\Upsilon \Upsilon$ or $\Upsilon \Upsilon(2S)$.

hep-ph

LACO: Adaptive Latent Communication for Collaborative Driving

Collaborative driving aims to improve safety and efficiency by enabling connected vehicles to coordinate under partial observability. Recent approaches have evolved from sharing visual features for perception to exchanging language-based reasoning through foundation models for behavioral coordination. Though communicating in language provides intuitive information, it introduces two challenges: high latency caused by autoregressive decoding and information loss caused by compressing rich internal representations into discrete tokens. To address these challenges, we analyze latent communication in collaborative driving under inherent limitations of multi-agent settings. Our analysis reveals agent identity confusion, where direct fusion of latent states entangles decision representations across vehicles. Motivated by this, we propose LACO, a training-free \textbf{LA}tent \textbf{CO}mmunication paradigm that seamlessly adapts pretrained driving models to collaborative settings. LACO introduces Iterative Latent Deliberation (ILD) for latent reasoning, Cross-Horizon Saliency Attribution (CHSA) for communication-efficient information selection, and Structured Semantic Knowledge Distillation (SSKD) to stabilize ego-centric decision making. Closed-loop experiments in CARLA show that LACO notably reduces communication and inference latency while maintaining strong collaborative driving performance.

cs.AI

Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation

Interactive real-time autoregressive video generation is essential for applications such as content creation and world modeling, where visual content must adapt to dynamically evolving event conditions. A fundamental challenge lies in balancing reactivity and stability: models must respond promptly to new events while maintaining temporal coherence over long horizons. Existing approaches distill bidirectional models into autoregressive generators and further adapt them via streaming long tuning, yet often exhibit persistent drift after condition changes. We identify the cause as conditional bias, where the teacher may provide condition-aligned but trajectory-agnostic guidance, biasing generation toward locally valid yet globally inconsistent modes. Inspired by Trust Region Policy Optimization, we propose Delta Forcing, a simple yet effective framework that constrains unreliable teacher supervision within an adaptive trust region. Specifically, Delta Forcing estimates transition consistency from the latent delta between teacher and generator trajectories, and uses it to balance teacher supervision with a monotonic continuity objective. This suppress unreliable teacher-induced shifts while preserving responsiveness to new events. Extensive experiments demonstrate that Delta Forcing significantly improves consistency while maintaining event reactivity.

cs.CV

AgenticCache: Cache-Driven Asynchronous Planning for Embodied AI Agents

Embodied AI agents increasingly rely on large language models (LLMs) for planning, yet per-step LLM calls impose severe latency and cost. In this paper, we show that embodied tasks exhibit strong plan locality, where the next plan is largely predictable from the current one. Building on this, we introduce AgenticCache, a planning framework that reuses cached plans to avoid per-step LLM calls. In AgenticCache, each agent queries a runtime cache of frequent plan transitions, while a background Cache Updater asynchronously calls the LLM to validate and refine cached entries. Across four multi-agent embodied benchmarks, AgenticCache improves task success rate by 22% on average across 12 configurations (4 benchmarks x 3 models), reduces simulation latency by 65%, and lowers token usage by 50%. Cache-based plan reuse thus offers a practical path to low-latency, low-cost embodied agents. Code is available at https://github.com/hojoonleokim/MLSys26_AgenticCache.

cs.LG

NaviNote: Enabling In-situ Spatial Annotation Authoring to Support Exploration and Navigation for Blind and Low Vision People

GPS and smartphones enable users to place location-based annotations, capturing rich environmental context. Previous research demonstrates that blind and low vision (BLV) people can use annotations to explore unfamiliar areas. However, current commercial systems allowing BLV users to create annotations have never been evaluated, and current GPS-based systems can deviate several meters. Motivated by high-accuracy visual positioning technology, we first conducted a formative study with 24 BLV participants to envision a more accurate and inclusive annotation system. Surprisingly, many participants viewed the high-accuracy technology not just as an annotation system but also as a tool for precise last-few-meters navigation. Guided by participant feedback, we developed NaviNote, which combines vision-based high-precision localization with an agentic architecture to enable voice-based annotation authoring and navigation. Evaluating NaviNote with 18 BLV participants showed that it significantly improved navigation performance and supported users in understanding and annotating their surroundings. Based on these findings, we discuss design considerations for future accessible annotation authoring systems.

cs.HC

PISCO: Precise Video Instance Insertion with Sparse Control

The landscape of AI video generation is undergoing a pivotal shift: moving beyond general generation - which relies on exhaustive prompt-engineering and "cherry-picking" - towards fine-grained, controllable generation and high-fidelity post-processing. In professional AI-assisted filmmaking, it is crucial to perform precise, targeted modifications. A cornerstone of this transition is video instance insertion, which requires inserting a specific instance into existing footage while maintaining scene integrity. Unlike traditional video editing, this task demands several requirements: precise spatial-temporal placement, physically consistent scene interaction, and the faithful preservation of original dynamics - all achieved under minimal user effort. In this paper, we propose PISCO, a video diffusion model for precise video instance insertion with arbitrary sparse keyframe control. PISCO allows users to specify a single keyframe, start-and-end keyframes, or sparse keyframes at arbitrary timestamps, and automatically propagates object appearance, motion, and interaction. To address the severe distribution shift induced by sparse conditioning in pretrained video diffusion models, we introduce Variable-Information Guidance for robust conditioning and Distribution-Preserving Temporal Masking to stabilize temporal generation, together with geometry-aware conditioning for realistic scene adaptation. We further construct PISCO-Bench, a benchmark with verified instance annotations and paired clean background videos, and evaluate performance using both reference-based and reference-free perceptual metrics. Experiments demonstrate that PISCO consistently outperforms strong inpainting and video editing baselines under sparse control, and exhibits clear, monotonic performance improvements as additional control signals are provided. Project page: xiangbogaobarry.github.io/PISCO.

cs.CV

ARGaze: Autoregressive Transformers for Online Egocentric Gaze Estimation

Online egocentric gaze estimation predicts where a camera wearer is looking from first-person video using only past and current frames, a task essential for augmented reality and assistive technologies. Unlike third-person gaze estimation, this setting lacks explicit head or eye signals, requiring models to infer current visual attention from sparse, indirect cues such as hand-object interactions and salient scene content. We observe that gaze exhibits strong temporal continuity during goal-directed activities: knowing where a person looked recently provides a powerful prior for predicting where they look next. Inspired by vision-conditioned autoregressive decoding in vision-language models, we propose ARGaze, which reformulates gaze estimation as sequential prediction: at each timestep, a transformer decoder predicts current gaze by conditioning on (i) current visual features and (ii) a fixed-length Gaze Context Window of recent gaze target estimates. This design enforces causality and enables bounded-resource streaming inference. We achieve state-of-the-art performance across multiple egocentric benchmarks under online evaluation, with extensive ablations validating that autoregressive modeling with bounded gaze history is critical for robust prediction. We will release our source code and pre-trained models.

cs.CV

LLM-FSM: Scaling Large Language Models for Finite-State Reasoning in RTL Code Generation

Finite-state reasoning, the ability to understand and implement state-dependent behavior, is central to hardware design. In this paper, we present LLM-FSM, a benchmark that evaluates how well large language models (LLMs) can recover finite-state machine (FSM) behavior from natural-language specifications and translate it into correct register transfer-level (RTL) implementations. Unlike prior specification-to-RTL benchmarks that rely on manually constructed examples, LLM-FSM is built through a fully automated pipeline. LLM-FSM first constructs FSM with configurable state counts and constrained transition structures. It then prompts LLMs to express each FSM in a structured YAML format with an application context, and to further convert that YAML into a natural-language (NL) specification. From the same YAML, our pipeline synthesizes the reference RTL and testbench in a correct-by-construction manner. All 1,000 problems are verified using LLM-based and SAT-solver-based checks, with human review on a subset. Our experiments show that even the strongest LLMs exhibit sharply declining accuracy as FSM complexity increases. We further demonstrate that training-time scaling via supervised fine-tuning (SFT) generalizes effectively to out-of-distribution (OOD) tasks, while increasing test-time compute improves reasoning reliability. Finally, LLM-FSM remains extensible by allowing its FSM complexity to scale with future model capabilities.

cs.AI

Investigating $\Omega \phi$ Interaction and Correlation Functions

In this work, we investigate the interaction between the $\Omega$ baryon and the $s\bar{s}$ meson within the framework of the quark delocalization color screening model. The spectra calculations show that no bound state is formed in any of the considered channels, while the scattering indicates that the $\Omega\phi$ interaction with $J^{P}=1/2^{-}$ is weakly attractive. As for the $\Omega\phi$ interactions with $J^{P}=3/2^{-}$ and $5/2^{-}$, as well as the $\Omega\eta^{\prime}$ interaction with $J^{P}=3/2^{-}$, they are all repulsive. After an investigation on the femtoscopic correlation functions, we find that, due to the spin-averaging effect, the overall $\Omega\phi$ correlation function exhibits a weak dependence on the source size, which provides a crucial significance of our model for future experimental examinations in relativistic heavy-ion collisions.

hep-ph

AI-Driven Prediction of Cancer Pain Episodes: A Hybrid Decision Support Approach

Lung cancer patients frequently experience breakthrough pain episodes, with up to 91% requiring timely intervention. To enable proactive pain management, we propose a hybrid machine learning and large language model pipeline that predicts pain episodes within 48 and 72 hours of hospitalization using both structured and unstructured electronic health record data. A retrospective cohort of 266 inpatients was analyzed, with features including demographics, tumor stage, vital signs, and WHO-tiered analgesic use. The machine learning module captured temporal medication trends, while the large language model interpreted ambiguous dosing records and free-text clinical notes. Integrating these modalities improved sensitivity and interpretability. Our framework achieved an accuracy of 0.876 (48h) and 0.917 (72h), with improvements in sensitivity of 10.6% and 10.7%, respectively, attributable to large language model augmentation. This hybrid approach offers a clinically interpretable and scalable tool for early pain episode forecasting, with potential to enhance treatment precision and optimize resource allocation in oncology care.

cs.AI