Search arXivSearch

arXiv subjects

Dong Yang

Publications and source records attributed to Dong Yang.

At least 19 recordsLinked to original sources

K-shell x-ray spectroscopy: A reliable probe for stimulated Raman scattering in inertial confinement fusion---

Accurate characterization of stimulated Raman scattering (SRS) remains a critical challenge in inertial confinement fusion (ICF), as SRS not only scatters laser energy but also generates suprathermal electrons that preheat the fuel and degrade implosion performance. Conventional backscatter diagnostics provide direct measurements of SRS but cannot collect the entire scattered-light signal, limiting the accurate characterization of SRS strength. Using a non-local thermodynamic equilibrium collisional-radiative model with a double-Maxwellian electron distribution, we systematically investigate how suprathermal electrons modify the titanium \textit{K}-shell x-ray emission spectra. The spectra exhibit high sensitivity to the suprathermal-electron fraction at relatively low bulk electron temperatures, making them particularly suitable for diagnosing suprathermal electrons during the early stage of ICF, when even a small suprathermal-electron population can compromise fuel compression. The calculated spectra reproduce experimental measurements from the Nova Laser Facility with good agreement, and the inferred suprathermal-electron fractions show improved consistency with independently measured SRS losses compared with the original spectral analysis by Glenzer~\href{https://doi.org/10.1103/PhysRevLett.81.365} {\text{[S. H. Glenzer \textit{et al}., Phys. Rev. Lett. \textbf{81}, 365(1998)]}}. These results demonstrate that \textit{K}-shell spectroscopy, combined with accurate NLTE collisional-radiative modeling, provides a reliable probe of SRS strength in laser-produced plasmas. Hence, this approach may be extended to spatially resolved diagnosis of SRS strength, offering a promising complement to conventional backscatter diagnostics.

physics.plasm-ph

On the derived Hall algebra of a graded gentle one-cycle algebra I: the triangle structure

Under a mild condition, the perfect derived category and the finite-dimensional derived category of a graded gentle one-cycle algebra are described as twisted root categories of certain infinite quivers of type $\mathbb{A}_\infty^\infty$. As a consequence, it is shown that if $\ct$ is the perfect (respectively, finite-dimensional) derived category of such a graded gentle one-cycle algebra, then its triangle structure is up to triangle equivalence determined by the underlying additive category.

math.RT

Consensus-based Decentralized Distributed Swarm Learning with Heterogeneous Big Data

Artificial intelligence increasingly relies on large-scale, distributed, and heterogeneous data collected by edge devices. However, the practice of edge intelligence remains challenging due to non-convex objectives, data heterogeneity, and complex wireless network topology. To address these issues, this paper proposes a consensus-based decentralized distributed swarm learning (CD-DSL) framework for wireless edge networks. Our CD-DSL integrates consensus optimization with particle swarm optimization (PSO), by reaching the model consensus among neighboring devices while leveraging the PSO exploration and exploitation. The consensus mechanism supports decentralized coordination without raw-data exchange, while PSO-inspired updates utilize historical and neighbor-shared experience to enhance exploration for non-convex optimization, improve robustness to data heterogeneity, and accelerate convergence. We further develop an adaptive neighbor-mixing strategy that learns performance-aware consensus weights, improving decentralized collaboration among heterogeneous edge devices. Theoretical analysis establishes that CD-DSL maintains participant consistency and achieves non-ergodic convergence to a neighborhood of a stationary point under non-convex objectives. Experimental results show that CD-DSL can mitigate the performance degeneration of existing decentralized baselines caused by heterogeneous data.

cs.DC

Self-Trapping Enabled Highly Bright Momentum-Indirect Interlayer Excitons

Interlayer excitons in two dimensional material heterostructures exhibit large exciton binding energies and long lifetimes, making them ideal platforms for studying excitonic devices and many body quantum phenomena. However, the spatially separated electron and hole nature of IXs reduces their oscillator strength by two orders of magnitude compared to intralayer excitons. Achieving high efficiency IX emission remains challenging and requires optimal material selection with appropriate momentum matching and meticulous device fabrication. Here we demonstrate a highly bright momentum indirect IX emission within heterostructures formed between 2D perovskites and monolayer transition metal dichalcogenides. The quantum yield of IX emission reaches 35.2% on average, over 50 times higher than that of the corresponding constituent TMD monolayer, with the highest value exceeding 60%. Notably, the radiative recombination efficiency of this momentum indirect IX exceeds that of momentum direct IXs in monolayer TMD-based heterostructures by two orders of magnitude. We suggest that the remarkably bright IX emission in our heterostructure originates from IX self trapping, induced by strong exciton phonon coupling arising from the soft lattice nature of the 2D perovskite. Our findings provide new insights into achieving high IX emission efficiency and open new avenues for exploring long lifetime excitonic devices.

cond-mat.mes-hall

BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics

Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workflows are multi-stage and data-sensitive, while expert trajectories remain scarce and difficult to share. Structured benchmarks can localize failures through stage-level rubrics, but standard post-training discards these diagnostics before the next training round. We present Benchmark-as-Teacher (BaT), a recursive self-improvement system for agent post-training. BaT contains two linked components: the asynchronous Stage Bank data pipeline and BiCuRL (Bilevel Curriculum Reinforcement Learning), its self-improving post-training method. Stage Bank synthesizes content-isolated training states outside the policy-update loop. BiCuRL uses a fixed held-out evaluation to select the next stage curriculum, verifies rollouts with task rubrics, updates the policy with GRPO, and returns the candidate checkpoint to evaluation. On AutoMedBench-Lite, BaT-4B and BaT-9B more than double the Overall scores of their Qwen Instruct baselines. BaT-9B Agent reaches 79.6 Overall, exceeding Claude Opus 4.6 with Claude Code at 77.5.

cs.AI

Latent Interpolation Learning Using Diffusion Models for Cardiac Volume Reconstruction

Cardiac Magnetic Resonance (CMR) imaging is a critical tool for diagnosing and managing cardiovascular disease, yet its utility is often limited by the sparse acquisition of 2D short-axis slices, resulting in incomplete volumetric information. Accurate 3D reconstruction from these sparse slices is essential for comprehensive cardiac assessment, but existing methods face challenges, including reliance on predefined interpolation schemes (e.g., linear or spherical), computational inefficiency, and dependence on additional semantic inputs such as segmentation labels or motion data. To address these limitations, we propose a novel Cardiac Latent Interpolation Diffusion (CaLID) framework that introduces three key innovations. First, we present a data-driven interpolation scheme based on diffusion models, which can capture complex, non-linear relationships between sparse slices and improves reconstruction accuracy. Second, we design a computationally efficient method that operates in the latent space and speeds up 3D whole-heart upsampling time by a factor of 24, reducing computational overhead compared to previous methods. Third, with only sparse 2D CMR images as input, our method achieves SOTA performance against baseline methods, eliminating the need for auxiliary input such as morphological guidance, thus simplifying workflows. We further extend our method to 2D+T data, enabling the effective modeling of spatiotemporal dynamics and ensuring temporal coherence. Extensive volumetric evaluations and downstream segmentation tasks demonstrate that CaLID achieves superior reconstruction quality and efficiency. By addressing the fundamental limitations of existing approaches, our framework advances the state of the art for spatio and spatiotemporal whole-heart reconstruction, offering a robust and clinically practical solution for cardiovascular imaging.

eess.IV

Token-Based Affordance Grounding with Large Vision-Language Models

Affordance grounding aims to localize image regions that support a specific action, serving as a core capability for physical intelligence and embodied perception. Previous studies have primarily relied on weakly supervised learning with action labels from exocentric images. However, these methods often struggle with visually ambiguous exocentric images containing co-occurring actions; moreover, they fail to distinguish semantically similar actions because existing methods typically rely on brief action phrases that lack rich semantic details for action-specific localization. Although large vision-language models (LVLMs) encode rich action semantics and their action-conditioned textual outputs implicitly contain spatial cues, they do not directly provide action-specific spatial localization. To address these problems, we propose TokAG, a zero-shot affordance grounding framework that exploits the token-level semantic-spatial signals in LVLMs to localize action-relevant regions without external supervision. We observe that attention maps associated with different LVLM output tokens vary significantly, with many attending to irrelevant regions such as the background. Thus, we introduce a spatial-aware token-selection mechanism to systematically evaluate each output token and select the one whose attention maps exhibit dominant activation over the target object, instead of relying on arbitrary attention maps. By extracting these object-focused attention maps, we transform the LVLM's implicit semantic signals into zero-shot affordance heatmaps. Our zero-shot framework consistently outperforms prior weakly supervised approaches across multiple benchmarks, improving NSS by 10.7% on the unseen split of AGD20K and by 29.7% on HICO-IIF. The code and models will be made publicly available.

cs.CV

CraBERT: Efficient Phoneme Encoder Pre-Training via Cascade Fusion of Subword Representations for Text-to-Speech

This paper introduces CraBERT, a pre-trained phoneme encoder (PPEnc) designed for efficient pre-training in text-to-speech (TTS). CraBERT employs a cascade-fusion architecture and a subword-phoneme alignment algorithm to integrate representations from a pre-trained subword-level BERT into a phoneme-level BERT. This design provides prior word- and sentence-level information, reducing the amount of pre-training required by the phoneme encoder. Subjective listening evaluations show that CraBERT achieves MOS values comparable to existing PPEncs after approximately one epoch of pre-training, whereas the baselines in our comparison are pre-trained for approximately ten epochs. These results demonstrate that CraBERT can efficiently learn representations suitable for improving the perceived naturalness and prosody of synthesized speech.

eess.AS

AQA-TTRL: Self-Adaptation in Audio Question Answering with Test-Time Reinforcement Learning

Large Audio Language Models (LALMs) exhibit strong capabilities in general audio understanding but remain static after deployment, limiting their adaptability to real-world data. Since supervised fine-tuning is costly, we propose AQA-TTRL, a novel framework for audio understanding that enables on-the-fly evolution via test-time reinforcement learning using only unlabeled test data. It generates pseudo-labels via majority voting and optimizes the model through reinforcement learning. To address the noise in self-generated labels, we introduce confidence weighting to adjust training signals. Furthermore, multiple-attempt sampling mitigates advantage collapse and stabilizes training. Across MMAU, MMAR, and MMSU, AQA-TTRL achieves significant average improvements of 4.42% for Qwen2.5-Omni 7B and 11.04% for the 3B model. Notably, the adapted 3B model outperforms direct inference of the unadapted 7B model, highlighting the effectiveness of test-time adaptation in audio understanding.

eess.AS

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models

Autonomous agents are increasingly expected to support end-to-end medical-AI research workflows, moving beyond isolated prediction tasks or short-form clinical question answering. However, existing medical agent benchmarks primarily evaluate final outputs, providing limited visibility into agent behavior within the research process. To address this gap, we present AutoMedBench, a workflow-aware benchmark for autonomous medical-AI research across diverse medical imaging and multimodal inference tasks, organizing agent execution into a unified five-stage workflow (S1-S5): Plan, Setup, Validate, Inference, and Submit. It comprises long-horizon tasks with each run averaging 33 agent turns, spanning five research tracks: segmentation, image enhancement, visual question answering (VQA), report generation, and lesion detection. Each task is evaluated under two difficulty tiers, Lite and Standard, which use the same data and metrics but differ in the amount of task-brief scaffolding, and each run is scored using both final task performance and S1-S5 stage scores, enabling stage-level analysis from the initial task brief to the final submitted artifact. Across thousands of recorded runs, stage-level scoring reveals that Validate is the weakest workflow stage on average, whereas Setup is the strongest, suggesting that current agents are better at making pipelines executable than at verifying their reliability. Post-run error analysis further shows that verification and submission failures dominate tagged errors, accounting for 37.7% and 38.1% of fired codes respectively, whereas task-understanding errors are rare at 0.9%, and runs with one fired error code have a 48% lower overall score than runs with no error code on average.

cs.AI

MusaCoder: Native GPU Kernel Generation with Full-Stack Training on Moore Threads GPU

Native GPU kernel generation turns high-level tensor programs into executable, efficient low-level code. Existing Large Language Models (LLMs) struggle with this task, while execution-based reinforcement learning suffers from sparse rewards, reward hacking, and training instability. We present MusaCoder, a full-stack training framework for native GPU kernel generation on CUDA and MUSA backends. MusaCoder combines progressive kernel-oriented data synthesis, diversity-preserving rejection fine-tuning, and execution-feedback Reinforcement Learning (RL) through MooreEval, a distributed verifier and reward environment. To stabilize RL, MusaCoder introduces PrimeEcho for first-turn-anchored multi-turn rewards, Buffered Dynamic Retry for recovering signals from all-failed hard samples, and MirrorPop for off-policy sequence filtering. Experiments on KernelBench and a MUSA-ported variant show that MusaCoder outperforms strong open-source and proprietary baselines in both correctness and empirical speedup, with the 9B model matching or exceeding frontier closed-source models and the 27B model establishing a new state of the art. These results demonstrate not only the effectiveness of full-stack execution-feedback training for native kernel generation, but also the capability of Moore Threads GPUs to support the complete LLM post-training stack, providing a practical foundation for large-model training and optimization on emerging accelerators.

cs.CV

CmIVTP: Cross-modal Interaction-based Vessel Trajectory Prediction for Maritime Intelligence

Maritime intelligent transportation systems (MITS) are essential for ensuring navigation safety and efficiency in busy waterways. However, accurate vessel trajectory prediction remains challenging due to the limitations of single-source data. Automatic identification system (AIS) data is often sparse or unavailable for small vessels, while closed-circuit television (CCTV) data alone cannot fully capture dynamic vessel behavior. To mitigate these challenges, we propose a cross-modal interaction-based vessel trajectory prediction (named CmIVTP) framework to model the intricate interactions between vessel dynamics and environmental constraints. Specifically, we introduce a target-aware scene encoder to extract scene semantic features, effectively capturing vessel-environment interactions and enhancing trajectory prediction accuracy. In addition, we propose a cross-modal interaction transformer, which integrates AIS-derived motion features, CCTV-based environmental features, and scene representations. It leverages cross-modal attention mechanisms to simultaneously capture intra-modal semantics and inter-modal interactions, ensuring dynamically consistent and environmentally feasible predictions. Furthermore, we construct a vessel group trajectory bank by clustering historical AIS trajectories into representative motion patterns, providing an efficient and scalable approach for candidate trajectory generation. Additionally, we introduce the maritime multimodal dataset plus (named Maritime-MmD$^+$), a large-scale dataset that synchronizes AIS data and CCTV video data, providing robust support for multimodal trajectory prediction research. Extensive experiments demonstrate that CmIVTP achieves better performance on multimodal-driven vessel trajectory prediction benchmarks. The code resources for this work can be available at https://github.com/LouisYxLu/CmIVTP.

cs.CV

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech

Metric-induced discrete flow matching (MI-DFM) exploits token-latent geometry for discrete generation, but its practical use is limited by two issues: heuristic schedulers requiring hyperparameter search, and finite-step path-tracking error from its first-order continuous-time Markov chain (CTMC) solver. We address both issues. First, we derive a kinetic-optimal scheduler for prescribed scalar-parameterized probability paths, and instantiate it for MI-DFM as a training-free numerical schedule that traverses the path at constant Fisher-Rao speed. Second, we introduce a finite-step moment correction that adjusts the jump probability while preserving the CTMC jump destination distribution. We validate the resulting method, GibbsTTS, on codec-based zero-shot text-to-speech (TTS). Under controlled comparisons with a unified architecture and large-scale dataset, GibbsTTS achieves the best objective naturalness and is preferred in subjective evaluations over masked discrete generative baselines. Additionally, in comparison with the evaluated state-of-the-art TTS systems, GibbsTTS shows strong speaker similarity, achieving the highest similarity on three of four test sets and ranking second on the fourth. Project page: https://ydqmkkx.github.io/GibbsTTSProject

eess.AS

Improved Evidence Extraction and Metrics for Document Inconsistency Detection with LLMs

Large language models (LLMs) are becoming useful in many domains due to their impressive abilities that arise from large training datasets and large model sizes. However, research on LLM-based approaches to document inconsistency detection is relatively limited. We address this gap by investigating evidence extraction capabilties of LLMs for document inconsistency detection. To this end, we introduce new comprehensive evidence-extraction metrics and a redact-and-retry framework with constrained filtering that substantially improves evidence extraction performance over other prompting methods. We support our approach with strong experimental results and release a new semi-synthetic dataset for evaluating evidence extraction.

cs.CL

Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence

Surgical intelligence has the potential to improve the safety and consistency of surgical care, yet most existing surgical AI frameworks remain task-specific and struggle to generalize across procedures and institutions. Although multimodal foundation models, particularly multimodal large language models, have demonstrated strong cross-task capabilities across various medical domains, their advancement in surgery remains constrained by the lack of large-scale, systematically curated multimodal data. To address this challenge, we introduce Surg$Σ$, a spectrum of large-scale multimodal data and foundation models for surgical intelligence. At the core of this framework lies Surg$Σ$-DB, a large-scale multimodal data foundation designed to support diverse surgical tasks. Surg$Σ$-DB consolidates heterogeneous surgical data sources (including open-source datasets, curated in-house clinical collections and web-source data) into a unified schema, aiming to improve label consistency and data standardization across heterogeneous datasets. Surg$Σ$-DB spans 6 clinical specialties and diverse surgical types, providing rich image- and video-level annotations across 18 practical surgical tasks covering understanding, reasoning, planning, and generation, at an unprecedented scale (over 5.98M conversations). Beyond conventional multimodal conversations, Surg$Σ$-DB incorporates hierarchical reasoning annotations, providing richer semantic cues to support deeper contextual understanding in complex surgical scenarios. We further provide empirical evidence through recently developed surgical foundation models built upon Surg$Σ$-DB, illustrating the practical benefits of large-scale multimodal annotations, unified semantic design, and structured reasoning annotations for improving cross-task generalization and interpretability.

cs.AI

FEASTS and MHONGOOSE: HI Column Density Distribution at $z=0$ for $N_\mathrm{HI}>10^{17.8}\, \mathrm{cm}^{-2}$

We present the first $z=0$ HI column density distribution function, $f(N_\mathrm{HI})$, extending down to $\log (N_\mathrm{HI}/\mathrm{cm}^{-2})=17.8$. This was derived from high-sensitivity 21-cm emission-line imaging at $\sim$1 kpc resolution. At high-column-densities (19.8$< \log (N_\mathrm{HI}/\mathrm{cm}^{-2}) <$21.3), our results align with earlier $z=0$ studies but benefit from 100 times greater sensitivity. Comparisons with $z\sim3$ quasar absorption-line studies reveal that $f(N_\mathrm{HI})$ at $z=0$ is systematically lower by 0.1-0.4 dex for $19.2< \log (N_\mathrm{HI}/\mathrm{cm}^{-2}) <21$. However, the distributions become comparable at $17.8< \log (N_\mathrm{HI}/\mathrm{cm}^{-2}) <19.2$, suggesting weak evolution in this regime. Extrapolating the length incidence ($\mathrm{d}N/\mathrm{d}X$) for $\log (N_\mathrm{HI}/\mathrm{cm}^{-2}) >17.5$ implies a covering fraction ($f_\mathrm{cov}$) of $\sim0.7$ within 1-kpc-scale HI-detected pixels at $z=0$. Notably, for $17.8< \log (N_\mathrm{HI}/\mathrm{cm}^{-2}) <20$, impact parameters at a given $N_\mathrm{HI}$ are significantly lower than previous $z\sim0$ absorption-line results and TNG50 simulation predictions. This discrepancy indicates challenges in identifying galaxy counterparts for absorbers and in recovering low-column-density HI within cosmological simulations. Finally, we derive a covering fraction of 0.006 for $\log (N_\mathrm{HI}/\mathrm{cm}^{-2}) >17.8$ gas within the virial radius around Milky-Way-like galaxies. These findings provide new constraints on the baryonic flows and gaseous dynamics governing galaxy evolution.

astro-ph.GA

EMBERS I: Low redshift post-starburst galaxies are frequently depleted in molecular gas relative to star forming progenitors

The cold gas content of post-starburst galaxies (PSBs) provides important insight into the mechanisms that drive rapid quenching, but a multiphase assessment of both the atomic and molecular gas in PSBs does not yet exist. We introduce the Ensemble of Multiphase Baryons Evolving in Rapidly-quenching Systems, or EMBERS, a homogeneously selected, nearly mass- and redshift-complete survey of the global atomic (HI) and molecular gas (H2) in PSBs, observed with the Five Hundred-metre Aperture Spherical Telescope (FAST) and the Institut de radioastronomie millimetrique (IRAM) 30m telescope. We present new CO(1-0) observations for 52 PSBs with the IRAM 30m, which, combined with 9 archival observations, gives a total H2 sample of 61, of which 58/61 have ancillary HI measurements. We detect CO(1-0) in 34/61 galaxies, corresponding to molecular gas fractions (fH2 = MH2/M*) ranging from two to 250 per cent. By comparing with a stellar-mass matched star-forming (SF) control sample from xCOLD GASS, we find that PSBs on average are 0.3-0.6 dex depleted in H2. However, considering both HI and H2, individual PSBs host diverse gas reservoirs ranging from gas-rich in both phases, elevated in one phase, or gas-poor, the latter of which is common at lower stellar mass. The existence of gas-normal and gas-depleted PSBs in both phases suggests that some PSBs may rejuvenate their star formation, but the rapid shutdown of star formation in others is likely terminal. Despite this diversity, the majority of EMBERS PSBs are gas-poor compared to SF controls, with the typical PSB hosting gas reservoirs intermediate to those found in star-forming and quenched galaxies.

astro-ph.GA