Search arXiv⌕ Search

arXiv subjects

Yue Su

Publications and source records attributed to Yue Su.

At least 19 recordsLinked to original sources

Scalable Dynamic Pricing of Substitutable Products through Structure-Guided Policy Learning

Problem definition: We study dynamic pricing of substitutable products with finite, product-specific inventories. Customer substitution couples pricing decisions across products, while the inventory state makes exact dynamic programming intractable at realistic scale. Methodology / results: We develop two MNL-guided policy-learning approaches that replace the dynamic program with a statistical mapping from inventory states to pricing decisions. The first learns prices directly, while the second learns inventory opportunity costs and converts them into prices using the optimal MNL pricing rule. Both policies are trained using decision-focused learning. We also propose an efficient method to generate anticipative customer-choice targets to train our policies. Managerial implications: We conduct an extensive numerical evaluation of our approaches. On small instances for which the optimal dynamic program can be computed, the learned policies achieve average optimality gaps below 0.4%. On larger airline-motivated instances, they consistently improve on the tested revenue-management benchmarks. The comparison between the two architectures also highlights the role of model structure: using the MNL pricing characterization is particularly effective when demand is well described by MNL, while directly learning prices provides greater flexibility under heterogeneous mixed-MNL demand.

math.OC↗

Combinatorial Optimization Augmented Machine Learning for Dynamic Electric Autonomous Dial-a-Ride Problem

This study introduces a decision-epoch-based dynamic electric autonomous dial-a-ride problem (Dyn-EADARP), in which incoming requests are collected and processed at periodic decision epochs. A key decision is not only how to serve requests, but also when to dispatch them. At each decision epoch, the service provider decides which requests to serve and which to postpone, while jointly determining vehicle routes, schedules, and charging decisions. Unexecuted parts of existing plans can be revised as new information becomes available. To solve this problem, we develop an ML--CO policy following the combinatorial optimization augmented machine learning (COAML) framework, which combines a statistical model with a combinatorial optimization layer for decision making. The statistical model predicts prizes for available requests, and a prize-collecting E-ADARP uses these prizes to jointly determine request selection, routing, scheduling, and charging. The statistical model is trained directly to improve the decisions produced by the optimization layer. Computational experiments on 320 test instances demonstrate the efficiency of ML--CO, which solves instances with nearly 500 requests in about 6 seconds on average. It achieves 8.7%--14.3% lower objective values than benchmark policies and an average gap of 4.8% to the anticipative reference. The results provide several managerial insights. First, serving requests immediately is not always best, as selectively postponing some requests can create better ride-sharing opportunities. Second, revising existing plans preserves operational flexibility and substantially improves solution quality. Finally, more frequent decision making does not necessarily improve performance, highlighting the importance of choosing an appropriate decision frequency.

math.OC↗

SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space

World Action Models (WAMs) couple action generation with prediction of future states. Their effectiveness depends on whether future dynamics are modeled in a space that is both aligned with action generation and sufficiently geometry-aware to capture where and how actions change the scene. Existing WAMs typically satisfy only part of this requirement, relying on either perceptually heavy observation-space targets or auxiliary latent spaces that are not jointly structured for action relevance and geometry. We propose SG-WAM, a self-guided framework that learns geometry-aware action-conditioned dynamics directly in the policy-derived representation space. SG-WAM introduces learnable dynamics tokens and a Self-Guided World Predictor that forecasts their future latent states conditioned on intervening robot actions. Prediction targets are generated by an exponential moving average copy of the same policy backbone, providing stable supervision within the representation family used by the action expert. Geometric supervision further structures the policy image-token representations, providing spatially grounded context for the dynamics tokens and yielding a future-alignment space that is both action-relevant and geometry-aware. Latent future prediction, geometric grounding, and flow-matching action generation are jointly optimized end-to-end in a unified framework. Built on a 0.9B model without large-scale embodied pretraining, SG-WAM achieves 98.5% average success on LIBERO and 73% on LIBERO-Plus, while outperforming strong baselines in both in-distribution and out-of-distribution real-world evaluations.

cs.RO↗

CARA: Concept-Aware Risk Attention for Interpretable Collision Anticipation

Collision anticipation in autonomous driving requires not only accurate early warnings but also interpretable reasoning about what risk factors are being tracked and how risk evolves over time. Existing methods fall short in this regard: feature-driven models are opaque, post-hoc explanations often lack fidelity, and concept-based methods are mostly designed for static recognition rather than dynamic driving scenes. We propose CARA (Concept-Aware Risk Attention), an intrinsically interpretable spatio-temporal framework for collision anticipation. CARA derives domain-grounded risk concepts from accident narratives, aligns them with video frames via vision-language similarity, and organizes them into evolving concept trajectories. These trajectories provide explicit risk evidence that guides spatial attention, temporal attention, and anticipation, allowing semantic concepts to directly influence both where the model attends and how it predicts risk over time. By treating semantic risk factors as dynamic intermediate evidence rather than auxiliary post-hoc explanations, CARA tightly couples interpretability with the predictive process. Extensive experiments on three benchmarks show that CARA consistently improves anticipation accuracy and warning earliness over strong baselines, while providing sparse and semantically grounded concept evidence.

cs.MM↗

Momentum-resolved EELS study of collective charge excitations in 1$T$-TaS$_2$

We use momentum-resolved electron energy-loss spectroscopy (M-EELS) to study the low-energy charge excitations of 1$T$-TaS$_2$ across the nearly commensurate-to-commensurate charge-density-wave (CDW) transition. Single-crystal x-ray diffraction and elastic M-EELS measurements confirm the expected rotation of the CDW wave vector upon entering the commensurate phase. In the nearly commensurate phase, the low-energy M-EELS spectra reveal an acoustic phonon branch and two optical phonon features whose energies and dispersions are broadly consistent with previous calculations and inelastic x-ray measurements. Across the transition, the optical phonon energies remain nearly unchanged, while their spectral intensity develops a pronounced temperature dependence near the CDW ordering wave vector. At higher energies, the finite-momentum charge response undergoes a substantial redistribution of spectral weight below the transition, consistent with the opening of an energy gap. These results demonstrate that M-EELS provides simultaneous access to lattice dynamics and finite-momentum valence band charge excitations in 1$T$-TaS$_2$, revealing their evolution across the commensurate CDW transition.

cond-mat.str-el↗

Trust Stack for Mental Health AI: A Survey of Calibration across Human, Interaction, and AI Layers

Language-based AI is increasingly deployed for mental health support, yet trust is evaluated in interdisciplinary but operationally misaligned ways: NLP and AI work measures robustness, safety, privacy, and explanations, while psychotherapy, HCI, and regulatory work emphasize therapeutic fidelity, lived experience, empathy, and reliance. Empathetic chatbots can elicit strong user trust without commensurate safety, while safer systems are under-trusted when their boundaries are opaque, a calibration gap no single community owns. Through a structured scoping synthesis of 61 papers, we survey this landscape into a three-layer framework separating (L1) human-oriented trust, (L2) interaction-oriented trustworthiness, and (L3) AI-oriented trustworthiness, and map five stakeholder perspectives onto these layers. We outline a research agenda for building socio-technically aligned trustworthy AI for mental health support, highlighting that the central objective should shift from maximizing perceived trust to calibrating human trust to demonstrated interaction- and AI-level trustworthiness.

cs.CL↗

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants

Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precise object state detection are required. In this work, we explore LEGO Co-builder, a hybrid benchmark combining real-world LEGO assembly logic with programmatically generated multimodal scenes. The dataset captures stepwise visual states and procedural instructions, allowing controlled evaluation of instruction-following, object detection, and state detection. We introduce a unified framework and assess leading VLMs such as GPT-4o, Gemini, and Qwen-VL, under zero-shot and fine-tuned settings. We also evaluated the framework using a reasoning-focused model, GLM-4.1-thinking. Our results show that while object detection achieved high performance (98.16% with fine-tuned InstructBLIP), fine-grained scene understanding and assembly state detection remain challenging: Fine-tuned MiniGPT-v2 reached only 37.52% F1 for identifying theme entities, and even advanced models such as GPT-4o achieved just 40.54% F1 on state detection. This highlights gaps in fine-grained visual understanding among existing models. We release the benchmark, codebase, and generation pipeline to support future research on multimodal assembly assistants grounded in real-world workflows.

cs.AI↗

CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos

Generalist Vision-Language-Action models remain constrained by the scarcity of robotic data relative to the abundance of human video demonstrations. Existing Latent Action Models attempt to use video data but often suffer from visual entanglement, encoding noise rather than manipulation skills. To address this limitation, we propose Contrastive Latent Action Pretraining (CLAP), a framework that first uses Act-VAE to learn an executable action-token vocabulary from robot trajectories and then aligns human visual transitions with this vocabulary through contrastive learning. This alignment maps unlabeled human videos into a physically grounded latent action space rather than reconstructing appearance. Building on the aligned tokens, we train CLAP-NTP as an autoregressive VLA using robot demonstrations and pseudo-labeled human videos, preserving instruction following and object generalization. For deployment and target-domain adaptation, we further introduce a post-training strategy that combines CLAP-RF, a Rectified Flow action head for low-latency continuous action chunk prediction, with Knowledge Matching regularization to preserve pretrained semantic knowledge during fine-tuning. Extensive experiments show that CLAP achieves strong performance against competitive baselines while enabling effective skill transfer from human videos to robotic execution.

cs.RO↗

SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity

This work investigates the ability of large language models (LLMs) to generate mathematical equations from scientific texts. Prior work faces challenges in unstructured grounding, multi-equation dependency, and humanaligned evaluation. To this end, we construct a dataset of AI research papers, pairing contextual passages with ground-truth equations and variable descriptions. We develop an explainable equation generation workflow and evaluate it across diverse open- and closed-source LLM backbones. We introduce an evaluation protocol combining automatic metrics, LLM-based rubrics, and human judgments to assess accuracy, explainability, and human-LLM alignment. Results indicate that LLMs perform moderately on lexical- and syntactic-based similarity, while struggling with semantic accuracy. Comparisons between LLM-based evaluations and human judgments reveal limited alignment, highlighting challenges in using LLMs to assess equation quality. These findings offer insights for improving equation generation models and developing more reliable evaluation methods for scientific text. We provide code and data for reproducibility.

cs.AI↗

ClawTrace: Cost-Aware Tracing for LLM Agent Skill Distillation

Skill-distillation pipelines learn reusable rules from LLM agent trajectories, but they lack a key signal: how much each step costs. Without per-step cost, a pipeline cannot distinguish adding a missing step to fix a bug from removing an expensive step that never affected the outcome. We use the cost-attribution gap to ask whether the rule types inside a distilled skill transfer the same way to new tasks. ClawTrace records cost-attributed agent traces and compiles each session into a TraceCard; CostCraft reads TraceCards and writes three kinds of skill patches: preserve, prune, and repair. We find a pattern aggregate metrics hide. On 30 held-out SpreadsheetBench tasks across two seeds, removing prune patches roughly tripled the quality-regression count without lowering median cost. Across the full 84-task SkillsBench transfer, CostCraft saves no aggregate cost. All three quality regressions trace to the preserve lane, and both quality wins trace to the prune lane: prune patches act as quality guardrails while preserve patches drive regressions. We argue that reusable agent skills should be evaluated at the rule-type level, not as monolithic instruction packages. To support this, we release ClawTrace, the TraceCard schema, and the full set of typed skills.

cs.AI↗

Towards Generalist Game Players: An Investigation of Foundation Models in the Game Multiverse

The real world unfolds along a single set of physics laws, yet human intelligence demonstrates a remarkable capacity to generalize experiences from this singular physical existence into a multiverse of games, each governed by entirely different rules, aesthetics, physics, and objectives. This omni-reality adaptability is a hallmark of general intelligence. As Artificial Intelligence progresses towards Artificial General Intelligence, the multiverse of games has evolved from mere entertainment into the ultimate ground for training and evaluating AGI. The pursuit of this generality has unfolded across four eras: from environment-specific symbolic and reinforcement learning agents, to current large foundation models acting as generalist players, and toward a future creator stage where agent both creates new game worlds and continually evolves within them. We trace the full lifecycle of a generalist game player along four interdependent pillars: Dataset, Model, Harness, and Benchmark. Every advance across these pillars can be read as an attempt to break one of five fundamental trade-offs that currently bound the whole system. Building on this end-to-end view, we chart a five-level roadmap, progressing from single-game mastery to the ultimate creator stage in which the agent simultaneously creates and evolves within theoretical game multiverse. Taken together, our work offers a unified lens onto a rapidly shifting field,and a principled path toward the omnipotent generalist agent capable of seamlessly mastering any challenge within the multiverse of games, thereby paving the way for AGI.

cs.CV↗

StoryMI: Steerable Multi-Agent Therapeutic Dialogue Generation

Large language models (LLMs) can generate fluent dialogue, but prior works lack situational grounding, dynamic strategy control, and evaluation aligned with clinical standards in motivational interviewing (MI). We introduce StoryMI, a multi-LLM agent framework for controllable MI dialogue generation, where questionnaire-based client profiles are expanded into situational stories that provide narrative context for the dialogue. Therapist and client agents generate MI-coded utterances guided by MI codes selected by the interaction agent, while an interaction agent dynamically coordinates exchanges to control MI strategies during a multi-turn conversation. We propose a two-level evaluation protocol: lexical metrics and MI-specific measures of macro-level counseling strategies, alongside LLM-as-judge and human expert assessments. We construct a dataset of 6K simulated MI dialogues grounded in 1K questionnaire-story pairs, covering 12 MI codes and 13 symptom domains, and benchmark six open- and closed-source LLMs. Our results show that situational grounding and macro-level control can improve MI adherence and clinical plausibility, demonstrating the effectiveness of a structured multi-agent workflow for psychotherapy dialogue generation. We provide code and data for reproducibility.

cs.CL↗

Diagnosing Retrieval vs. Utilization Bottlenecks in LLM Agent Memory

Memory-augmented LLM agents store and retrieve information from prior interactions, yet the relative importance of how memories are written versus how they are retrieved remains unclear. We introduce a diagnostic framework that analyzes how performance differences manifest across write strategies, retrieval methods, and memory utilization behavior, and apply it to a 3x3 study crossing three write strategies (raw chunks, Mem0-style fact extraction, MemGPT-style summarization) with three retrieval methods (cosine, BM25, hybrid reranking). On LoCoMo, retrieval method is the dominant factor: average accuracy spans 20 points across retrieval methods (57.1% to 77.2%) but only 3-8 points across write strategies. Raw chunked storage, which requires zero LLM calls, matches or outperforms expensive lossy alternatives, suggesting that current memory pipelines may discard useful context that downstream retrieval mechanisms fail to compensate for. Failure analysis shows that performance breakdowns most often manifest at the retrieval stage rather than at utilization. We argue that, under current retrieval practices, improving retrieval quality yields larger gains than increasing write-time sophistication. Code is publicly available at https://github.com/boqiny/memory-probe.

cs.AI↗

NeuroWise: A Multi-Agent LLM "Glass-Box" System for Practicing Double-Empathy Communication with Autistic Partners

The double empathy problem frames communication difficulties between neurodivergent and neurotypical individuals as arising from mutual misunderstanding, yet most interventions focus on autistic individuals. We present NeuroWise, a multi-agent LLM-based coaching system that supports neurotypical users through stress visualization, interpretation of internal experiences, and contextual guidance. In a between-subjects study (N=30), NeuroWise was rated as helpful by all participants and showed a significant condition-time effect on deficit-based attributions (p=0.02): NeuroWise users reduced deficit framing, while baseline users shifted toward blaming autistic "deficits" after difficult interactions. NeuroWise users also completed conversations more efficiently (37% fewer turns, p=0.03). These findings suggest that AI-based interpretation can support attributional change by helping users recognize communication challenges as mutual.

cs.HC↗

World Guidance: World Modeling in Condition Space for Action Generation

Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision-Language-Action (VLA) models. However, existing approaches struggle to strike a balance between maintaining efficient, predictable future representations and preserving sufficient fine-grained information to guide precise action generation. To address this limitation, we propose WoG (World Guidance), a framework that maps future observations into compact conditions by injecting them into the action inference pipeline. The VLA is then trained to simultaneously predict these compressed conditions alongside future actions, thereby achieving effective world modeling within the condition space for action inference. We demonstrate that modeling and predicting this condition space not only facilitates fine-grained action generation but also exhibits superior generalization capabilities. Moreover, it learns effectively from substantial human manipulation videos. Extensive experiments across both simulation and real-world environments validate that our method significantly outperforms existing methods based on future prediction. Project page is available at: https://selen-suyue.github.io/WoGNet/

cs.RO↗

Combinatorial Optimization Augmented Machine Learning

Combinatorial optimization augmented machine learning (COAML) has recently emerged as a powerful paradigm for integrating predictive models with combinatorial decision-making. By embedding combinatorial optimization oracles into learning pipelines, COAML enables the construction of policies that are both data-driven and feasibility-preserving, bridging the traditions of machine learning, operations research, and stochastic optimization. This paper provides a comprehensive overview of the state of the art in COAML. We introduce a unifying framework for COAML pipelines, describe their methodological building blocks, and formalize their connection to empirical cost minimization. We then develop a taxonomy of problem settings based on the form of uncertainty and decision structure. Using this taxonomy, we review algorithmic approaches for static and dynamic problems, survey applications across domains such as scheduling, vehicle routing, stochastic programming, and reinforcement learning, and synthesize methodological contributions in terms of empirical cost minimization, imitation learning, and reinforcement learning. Finally, we identify key research frontiers. This survey aims to serve both as a tutorial introduction to the field and as a roadmap for future research at the interface of combinatorial optimization and machine learning.

cs.LG↗

$η$ and $η'$ mesons from $N_f = 2+1$ lattice QCD at the physical point using topological charge operators

By fitting the two-point correlation functions of topological charge density operators calculated on two $2+1$-flavor gauge ensembles with physical pion mass, we determine both the $η$ and $η'$ masses and also the mixing angle to be $m_η= 0.505(72)(75)$ GeV, $m_{η'}=0.952(47)(40)$ GeV, and $θ_1 = -8.9(2.1)(1.8)^\circ$, respectively, where the first error is the statistical uncertainty and the second one is the systematic uncertainty. This is the first extraction of both $η/η'$ masses and the mixing angle $θ_1$ using topological charge operators. Compared with previous studies using quark bilinear operators, the error of the $η$ mass is relatively large, but the mixing angle has comparable precision. This demonstrates that the topological charge operators are well suited to study the $η$ and $η'$ mesons.

hep-lat↗

The Bi-objective Electric Autonomous Dial-a-Ride Problem

The electric autonomous dial-a-ride problem (E-ADARP) introduces electric, autonomously driving vehicles and their unique requirements into the classic dial-a-ride problem, where people are transported between pickup and drop-off locations. Next to an electric autonomous vehicle fleet, in the literature, a weighted-sum objective function, which combines the classic routing cost-oriented objective with a user-oriented objective function, has usually been considered. The user-oriented objective function minimizes the total excess user ride time. In this work, we treat them as two separate objective functions, which are optimized concurrently. In order to address the resulting bi-objective E-ADARP, we develop a novel exact framework (called fragment-based checker), whose core part is a smart ``select-and-check" algorithm that iteratively constructs feasible solutions using fragments. Several enhancements are proposed to enforce the computational efficiency of the proposed method. In the computational experiments, we evaluate several variants of our checker algorithm by leveraging a previously developed branch-and-price algorithm. We benchmark the checker-based framework against state-of-the-art criterion space frameworks as well as a generalized branch-and-price algorithm. Numerical results on both bi-objective DARP and E-ADARP instances demonstrate the effectiveness of the proposed framework. With our proposed approaches, 21 out of 38 instances are solved optimally, where small-to-medium-sized instances are solved within seconds. On larger-scale instances, especially those requiring high battery end levels are computationally challenging to solve, our approaches provide high-quality approximations of the Pareto frontiers. Efficient solutions with varying energy restrictions are compared and we obtain valuable managerial insights for different kinds of service providers.

math.OC↗