Search arXivSearch

arXiv subjects

Shuo Wang

Publications and source records attributed to Shuo Wang.

At least 19 recordsLinked to original sources

Gaze responses to false-positive computer-aided detection prompts during colonoscopy: a paired-video and real-time eye-tracking study

False-positive computer-aided detection (CADe) prompts may divert endoscopists' attention during colonoscopy, yet the attentional impact of individual prompts remains unclear. We used event-locked eye tracking to quantify gaze attraction and attention occupation in complementary retrospective and prospective studies. In a retrospective paired-video experiment, 3 senior and 2 novice endoscopists viewed 60 colonoscopy videos with and without CADe. The prospective study recorded gaze during 42 real-time CADe-assisted colonoscopies performed by 9 senior endoscopists. Screened CADe prompts outside expert-annotated lesion windows were classified as false-positive artifact events. False-positive prompts attracted gaze in 48.6% (68/140) of retrospective observations and 65.2% (533/817) of prospective events. Among attraction events with complete recovery, median attention occupation lasted 1000 ms in the retrospective study and 1100 ms in the prospective study. Corresponding median prompt durations were 33 ms and 267 ms, with median time amplifications of 17.55-fold and 5.15-fold, respectively. In paired retrospective comparisons, visible artifact prompts drew gaze closer to the prompted region than did the same-coordinate unassisted reference. Secondary retrospective analyses showed high lesion gaze recognition without and with CADe (98.0% versus 99.0%). First gaze entry into lesion regions occurred 147.8 ms earlier with CADe. Across controlled and real-time clinical settings, false-positive CADe prompts frequently captured gaze, with attention persisting beyond prompt visibility. These findings support considering prompt-related attentional burden in CADe evaluation and design.

cs.AI

Preserving Geometric Integrity in Graph Prompting via Measure-Constrained Optimal Transport

Graph prompt learning enables parameter-efficient adaptation of frozen Graph Neural Networks to downstream tasks through lightweight prompt parameters. As routing becomes increasingly node-adaptive, however, independently optimized local decisions can collectively concentrate assignment mass on a small subset of a finite shared prompt bank, even when individual node--prompt matches remain locally meaningful. We propose MINT (Measure-INtegrity Transport), an entropically regularized optimal transport framework that formulates node-to-prompt adaptation as a globally coupled allocation problem. The transport cost favors local geometric compatibility, while a prescribed prompt-side marginal explicitly controls graph-wide prompt utilization. We further derive an exact variance decomposition that separates prompt-side geometric variance into retained prompt-update variation and within-node barycentric dispersion, together with a conditional stability bound for the frozen-encoder forward map. Across standard citation networks and additional heterophilic graphs, MINT remains competitive in few-shot adaptation. Controlled and end-to-end experiments further distinguish the roles of routing and topology: fixed-marginal routing controls graph-wide prompt utilization and has measurable end-to-end effects on citation networks, while topology augmentation provides a complementary, graph-dependent mechanism for addressing structural mismatch. Code is available at https://github.com/Ga1axy0051/MINT.

cs.LG

AquaWorld: Structure-Consistent Underwater World Generation for Robot Simulation

Underwater robot simulation requires diverse environments in which terrain, scene composition, tasks, and currents remain mutually consistent. We present AquaWorld, a world-generation framework that preserves these relationships through shared terrain structure. A language-conditioned plan generates 3D terrain and shared structural references that guide asset placement, task definition, and inflow specification under stochastic variation. The framework incorporates over 10,000 underwater-compatible assets, predicts reusable terrain-conditioned mean-flow fields through a CFD-supervised residual model, and supports conventional underwater vehicles and bio-inspired robotic fish. On 24 paired terrains, structure-consistent randomization produces substantially better cross-factor consistency than independent randomization. In a matched-budget policy-training comparison, structurally coherent randomization achieves a validation success rate 21% higher than independent randomization. In separate physical experiments, a simulation-trained visual navigation policy succeeds in 95% of physical tank trials without updating its perception or control modules. Overall, AquaWorld provides a practical way to generate varied underwater environments while retaining the structural relationships needed for flow simulation and robot learning.

cs.RO

AIMold: An Autonomous AI-based Pipeline for Complex Mold Design

Injection molding is the cornerstone of mass-producing plastic components. While current algorithms can automate mold design for basic geometries using standard two-piece molds, complex parts featuring undercuts, side holes, or re-entrant features present a significant challenge. These geometries often necessitate auxiliary components beyond the primary upper and lower molds. In practice, designing these intricate assemblies is a laborious process that relies heavily on expert knowledge. Furthermore, the scarcity of public datasets has hindered the development of effective learning-based solutions. To bridge these gaps, we introduce MoldCAD, a curated dataset that pairs complex single-body CAD parts with industry-standard mold assemblies. Each entry includes the upper and lower molds, parting surfaces, demolding orientations, and necessary auxiliary components. The dataset comprises 4,934 CAD models and over 3,850 mold assemblies, totaling more than 23k individual models. Building upon this dataset, we propose a comprehensive pipeline that predicts demolding orientations, identifies auxiliary components, and constructs parting surfaces to derive a complete, manufacturing-ready mold assembly for downstream CAD/CAM workflows. Our results demonstrate a promising path toward fully automated industrial mold design and contribute to the broader advancement of manufacturing-aware CAD generation.

cs.CV

Adaptive Chemotherapy Control under Tumor Heterogeneity via Reinforcement Learning

Designing effective chemotherapy regimens is hindered by tumor heterogeneity and drug resistance, which complicate the deployment of patient-specific model-based optimal control across diverse populations. We develop and compare closed-loop deep reinforcement learning (DRL) dosing policies with continuous (TD3) and discrete (DQN) action spaces trained on a high-dimensional heterogeneous tumor model. The DRL policies are benchmarked against a Pontryagin's Maximum Principle (PMP)-derived open-loop benchmark. We assess generalization under parametric heterogeneity using a 100-patient virtual cohort with plus or minus 10 percent uniform perturbations in growth and drug-sensitivity parameters. Across this cohort, TD3 achieves higher average tumor reduction, while DQN yields tighter inter-patient dosing consistency, revealing a clear efficacy-consistency trade-off in this study. Our simulations assume full observation of all tumor subpopulations; translation to sparse and noisy clinical measurements will require partial-observability formulations and/or state estimation. Overall, the results show that simulation-trained DRL can learn state-dependent feedback dosing policies that complement open-loop optimal control benchmarks.

cs.LG

Optimal Control for Cancer Chemotherapy Using Hybrid Quantum Particle Swarm Optimization

Optimal control in cancer chemotherapy is challenged by tumor heterogeneity and mutations, which complicate treatment effectiveness. Traditional methods, such as Pontryagin's maximum principle (PMP), are often hindered by their reliance on an initial guess for the costate equation, affecting accuracy and convergence. To address these limitations, this work introduces a hybrid quantum particle swarm optimization (QPSO) method based on regularization. QPSO is employed for global exploration to approximate the optimal control trajectory, followed by a regularization-based refinement to ensure smoothness and consistency with optimality conditions. The Hamiltonian function is used for first- and second-order optimality checks, verifying solution quality. Numerical case studies explore various drug effectiveness functions, demonstrating the role of periodic and localized drug delivery in achieving robust tumor suppression and providing insights into the impact of different drug combinations on optimal chemotherapy strategies.

math.OC

Optimal chemo-immunotherapy scheduling: A hybrid QPSO-SQP approach with Michaelis-Menten pharmacodynamics

Optimal scheduling of combined chemo-immunotherapy is often formulated as a control-affine optimal control problem, which generically yields boundary-selected (bang-bang-type) protocols unless singular arcs occur. This structure, while convenient, neglects saturating pharmacodynamics at high dose rates. We incorporate Michaelis-Menten saturation directly into the therapy channels, making the dynamics non-control-affine and the Hamiltonian nonlinear in each control. With strictly positive exposure penalties, the Hamiltonian admits an explicit three-regime pointwise minimization law: each input is chosen at the lower bound, the upper bound, or as a unique interior minimizer. On interior intervals the strict Legendre condition holds, so continuously modulated dosing arises as a regular interior extremal rather than a singular-arc or smoothing artifact. The resulting transcription is nonconvex; we therefore use a hybrid Quantum Particle Swarm Optimization (QPSO)-Sequential Quadratic Programming (SQP) pipeline, where QPSO provides a constraint-aware warm start and SQP enforces feasibility and Karush-Kuhn-Tucker (KKT) optimality on a collocation grid. Costate reconstruction corroborates the Pontryagin Minimum Principle (PMP) structure and the predicted boundary/interior regime transitions.

math.OC

A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM

Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics of LLM guided framework that models CoT as a hidden-state trajectory and replaces hard deletion with an explicit-implicit interleaved latent architecture. After projecting question, step, and solution representations into a 3D PCA space, it measures alignment between each local transition and global question-to-solution direction. Aligned steps remain explicit text, whereas deviating steps are compressed into continuous latent tokens. Directional angles capture both local semantics and reasoning dynamics: small angles indicate direct execution and answer formation, while large angles more frequently involve checking, correction, and branch exploration; their temporal variation reveals exploration, convergence, and refinement stages. To train this architecture, we introduce stepwise embedding forcing, which pools each redundant step into a single latent embedding, and label forcing, which supervises that latent token with a soft multi-modal vocabulary distribution instead of a hard one-hot label. Experiments on Qwen3.5-9B and Qwen3.6-27B across six in-domain and out-of-domain benchmarks show that A*-Thought-V2 improves average accuracy by up to 2.6% while reducing response length by up to half, increasing Accuracy per Computation Unit by 2.29$\times$, and reducing preprocessing and training time by 94.6% and up to 80.3%, respectively. Representation analyses suggest that latent states form a compact region distinct from textual states, while higher entropy at latent-token positions reflects broader soft targets that encourage richer step-level feature learning.

cs.CL

Experimental realization of a dusty plasma rocking ratchet with current reversal

A single dust particle confined in an asymmetric ratchet potential is periodically driven by two oppositely directed laser beams, forming an underdamped dusty plasma rocking ratchet. We experimentally investigate the transport dynamics of the particle under varying driving amplitudes and frequencies. Depending on the driving conditions, the particle exhibits positive, zero, or negative net currents, and current reversal is observed when the driving parameters cross critical thresholds. To interpret these transport behaviors, we develop a simplified model based on the competition between the driving force and the ratchet confinement. The model reveals that directional transport is governed by two requirements: the driving force must exceed the depinning threshold, and the duration of a driving semicycle must be longer than the uphill escape time from a ratchet well. The resulting dynamic phase diagram quantitatively reproduces the experimentally observed transport regimes and current reversals. These results demonstrate dusty plasma as a versatile platform for investigating nonequilibrium transport phenomena of underdamped particles in rocking ratchets.

physics.plasm-ph

CoLMIN: LLM-based Multi-Decision Path Negotiation for Cooperative Autonomous Driving

Multi-vehicle cooperative autonomous driving enhances the safety and reliability of autonomous driving systems through information sharing among connected vehicles, demonstrating significant potential for improving traffic safety. LLM-based approaches leverage strong reasoning capabilities of LLMs to enable effective inter-vehicle negotiation and improve cooperative driving performance. However, driving decisions in complex traffic scenarios are inherently multi-solution in nature. As a result, existing negotiation-based methods often converge prematurely to suboptimal solutions, hindering consensus formation and limiting the practical deployment of cooperative autonomous driving systems. To address this challenge, we propose CoLMIN, the LLM-based multi-decision path negotiation framework for cooperative autonomous driving, achieving stable decision consensus through multi-decision path negotiation and reflective reasoning. To achieve stable and high-quality consensus in cooperative autonomous driving, CoLMIN consists of three key components: (i) an LLM-based Multi-Intent Negotiation module (LMin), which adopts a Negotiator-Evaluator paradigm and generates multiple candidate driving intentions for joint evaluation; (ii) an Evaluation-based Shallow Reflection Module (ESRM), which analyzes negotiation outcomes and provides feedback to guide subsequent negotiations, thereby accelerating consensus formation; and (iii) an LLM-based Deep Reflection Module (LDRM), which performs long-term reflection over negotiation histories to mitigate cognitive fixation and prevent the system from converging to suboptimal solutions. Experimental results in the CARLA simulation environment demonstrate that CoLMIN significantly outperforms existing methods in challenging interactive driving scenarios.

cs.RO

Life Operators: a self-evolving framework for multiscale life modelling

Medical AI is moving beyond recognition towards clinical dialogue and longitudinal prediction. Yet a central question remains: how would a patient's state change under intervention? Statistical models learn future observations, whereas mechanistic models describe selected processes. Neither provides a common framework for representing patient state, coupling scales or revising failed assumptions. We propose Life Operators: task-bounded mappings that define three scientific roles. Perception operators infer task-relevant biological states from multimodal observations, Evolution operators propagate these states under natural or intervention-conditioned dynamics, and Generation operators map them to measurable signals. Each role may be realised by equations, statistical models, neural networks or hybrids. Bridge operators connect components with different variables, scales and time steps. Selected operators and bridges form task-specific Operator Graphs containing the smallest set of states and mechanisms sufficient for a declared claim. This modular structure also makes scientific revision localisable. An AI co-scientist may propose changes to states, operators, bridges or graph structure, while independent evidence determines which variants are retained, restricted or retired. Over time, validated components could accumulate into broader multiscale models of the human body and provide a computational foundation for medical artificial superintelligence.

cs.CL

Coronary Mask Guided Registration for Continuous Time 4D Cardiac CT Dataset Construction

Objective: Clinical cardiac CT multiphase reconstructions generally provide acceptable image quality in end-diastole (ED) or end-systole (ES) phases, but in other phases may exhibit motion artifacts, especially in the right coronary artery (RCA). This limits ground-truth availability in 4D cardiac CT imaging research. We aim to construct a 4D cardiac CT dataset that is generally suitable to serve as pseudo ground truth. Methods: We propose Coronary Mask Guided Registration (CMGR) to produce a motion-preserved, artifact-reduced, and continuous-time 4D cardiac CT sequence from the clinical multiphase reconstruction of each patient. For artifact reduction, CMGR uses the ED or ES phase as the reference phase and warps the reference volume with deformation fields to produce the sequence. For motion preservation, CMGR registers the reference phase to each non-reference phase of the multiphase reconstruction. To capture the motion of both the RCA and other cardiac structures in each registration, CMGR regularizes RCA masks and incorporates them into image-domain registration. Time-continuity is achieved by interpolating the deformation fields for non-reference phases to arbitrary times. Results: CMGR outperformed representative image-domain registration methods in capturing RCA motion and providing reasonable RCA shape, and showed competitive performance in capturing whole-heart motion. Additionally, CMGR reduced motion artifacts from clinical multiphase reconstructions, and intermediate CMGR frames generally provided plausible transitions between discrete cardiac phases. Conclusion: CMGR provides an effective approach for constructing continuous-time 4D cardiac CT datasets. Significance: The dataset can be used in system design simulations and in reconstruction algorithm development, thereby facilitating advances in cardiac CT imaging.

eess.IV

Symmetry Origins of the Field-Free Superconducting Diode Effect in the Kagome Superconductor CsV$_3$Sb$_5$

Field-free superconducting diode effects require both inversion-symmetry breaking and an internal time-reversal-symmetry (TRS) breaking field, making them sensitive probes of hidden order in superconductors. In centrosymmetric kagome AV$_3$Sb$_5$, the inversion symmetry generally should generally preclude the observation of the superconducting diode effect. Furthermore, though TRS breaking has been reported in the superconducting regime of CsV$_3$Sb$_5$, whether it is generated by superconductivity or inherited from charge-density-wave (CDW) order remains unresolved. Here we show that pristine CsV$_3$Sb$_5$ devices exhibit no intrinsic field-free superconducting diode effect, whereas surface oxidation or asymmetric etching activates a large nonreciprocal supercurrent. Moreover, the response is stochastic, with sweep-dependent polarity and magnitude, indicating metastable TRS-breaking domain configurations. Small out-of-plane magnetic fields stabilize the superconducting diode response, consistent with field selection of such domains. Finally, when long-range CDW order is suppressed by Ti doping, the SDE disappears. Our results establish the symmetry requirements for the field-free SDE in CsV$_3$Sb$_5$, reveal its stochastic domain-controlled character, and link superconducting-state TRS breaking to CDW-related order.

cond-mat.supr-con

Topology enables learning-based hydrodynamic prediction of the global river system

Accurate river prediction is essential for water, food and energy security, yet remains challenging across entire river networks. Machine learning has transformed Earth-system modeling, but a system-level advance for river prediction lags for lack of reliable data. Exploiting the connectivity and dissipative dynamics of rivers, we introduce GraphRiverCast, a neural model for global river systems that predicts daily multivariate hydrodynamics at every reach of a 0.25°network with only sparse gauges and no initial state. It shows no intrinsic skill decay with lead time, outperforms leading global river models by 25% in accuracy, and robustly generalizes to ungauged reaches and finer resolutions. GraphRiverCast lifts learning-based river prediction from isolated basins to a unified global system and offers insights for machine learning in data-scarce Earth systems.

cs.LG

A Study of Bluetooth Access Control Based on NFT Soft Pairing

This paper proposes a Non-Fungible Token (NFT) soft pairing framework for Bluetooth service access control. Unlike conventional Bluetooth systems where pairing implicitly grants persistent service access, the proposed approach decouples native Bluetooth pairing from authorization without modifying the underlying protocol stack. The framework introduces a three-layer architecture consisting of a Bluetooth layer for connectivity, a blockchain layer for trusted execution and on-chain state verification, and an application layer where NFT soft pairing defines the authorization logic. In this design, Non-Fungible Bluetooth Tokens (NFBTs) represent user-side access credentials, while Non-Fungible Device Tokens (NFDTs) represent device identities. Their bidirectional on-chain binding forms a revocable and verifiable NFT soft pairing relationship. During access, users prove ownership of valid NFBTs through challenge-response signatures, and devices verify the corresponding on-chain state before granting service access. A prototype implemented with MetaMask and Ethereum demonstrates secure authentication, dynamic revocation, acceptable latency, and gas-efficient credential issuance based on ERC1155.

cs.CR

PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation

Bimanual manipulation in cluttered scenes requires policies that remain stable under occlusions, viewpoint changes and scene variations. Existing vision-language-action models often lack such robustness because (i) multi-view features are fused via view-agnostic token concatenation, yielding limited cross-view spatial representations, and (ii) language is injected as global conditioning, resulting in coarse instruction grounding. In this paper, we introduce PEAfowl, a perception-enhanced multi-view VLA policy for bimanual manipulation. For spatial perception, PEAfowl predicts per-token depth distributions, performs differentiable 3D lifting, and aggregates local cross-view neighbors to form geometrically grounded, cross-view aligned representations. For language utilization, we propose to replace global conditioning with a Perceiver-style text-aware readout over frozen CLIP visual features, enabling iterative evidence accumulation. To better exploit commodity RGB-D sensing despite noisy and incomplete depth, PEAfowl's depth-distribution lifting naturally supports training-only depth distillation, where a pretrained depth teacher supervises the depth-distribution head to inject refined geometric priors without adding inference overhead. On RoboTwin 2.0 under domain-randomized setting, PEAfowl improves the strongest baseline by 23.0 pp in success rate, and physical experiments further demonstrate improved performance on the evaluated real-robot tasks. Project website: https://peafowlvla.github.io/.

cs.CV

HiTac-WAM: A Hierarchical Tactile World Action Model for Contact-Rich Robot Manipulation

World action models jointly predict future visual observations and actions, whereas existing tactile-aware variants typically represent future touch as an image or latent stream without modeling the physical dependencies that organize tactile states hierarchically. We present HiTac-WAM, a hierarchical tactile world action model that forecasts a sequence of future tactile states for each candidate action chunk before execution. The forecast factorizes into contact state, a 3D deformation field, and slip risk, organized as a directed hierarchy in which each downstream stage is conditioned on stop-gradient signals from preceding stages. A directed attention mask allows tactile queries to attend to the video-action context of each candidate while preventing video and action queries from attending to tactile tokens. For planning, HiTac-WAM ranks candidate action chunks using tactile forecasts and task-progress estimates. For execution, the selected tactile forecast is retained as a reference; persistent discrepancies between predicted and observed tactile states trigger corrective replanning. HiTac-WAM achieves a mean contact F1 of 0.921; under matched training budgets, the directed hierarchy reduces 3D displacement L2 error by 17.6% relative to the deformation-only predictor and improves slip AUPRC by 60.4% relative to the slip-only predictor. Across chip grasping, blackboard erasing, and USB insertion, selection guided by the hierarchical forecasts increases the average real-robot success rate from 31.1% to 61.1%, while the full system attains 72.2%.

cs.RO

Large Field-Free Superconducting Diode Effect with Nonmonotonic Polarity Reversals in NbSe$_2$/CrBr$_3$ Heterostructures

The superconducting diode effect (SDE), characterized by nonreciprocal dissipationless supercurrent, offers a promising route toward ultralow-power superconducting electronics. Yet most realizations require an external magnetic field, and achieving a large, controllable SDE at zero field remains challenging. Here we report a field-free SDE with an efficiency reaching 35.7% in a van der Waals NbSe$_2$/CrBr$_3$ heterostructure, whose polarity is programmable by magnetic-field history. Unlike conventional mechanisms based on finite-momentum pairing, the SDE originates from the leading symmetry-allowed cubic term odd in Cooper-pair momentum, which yields unequal critical currents while leaving the equilibrium condensate at zero momentum. Remarkably, upon sweeping a perpendicular magnetic field, the diode polarity undergoes nonmonotonic and hysteretic reversals that cannot be explained by Meissner screening or conventional ferromagnetic proximity. Combining transport measurements, micromagnetic simulations, and a generalized Ginzbur-Landau theory, we attribute these unconventional behaviors to layered ferrimagnetism in CrBr$_3$, where coexisting ferromagnetic and antiferromagnetic interlayer couplings produce a history-dependent interfacial exchange field acting on NbSe$_2$. Our results reveal a distinct mechanism for nonreciprocal superconductivity and establish layered van der Waals magnetism as a versatile platform for high-efficiency, programmable, field-free superconducting diodes.

cond-mat.supr-con