Search arXivSearch

arXiv subjects

Gang Wu

Publications and source records attributed to Gang Wu.

At least 19 recordsLinked to original sources

FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model

While test-time scaling enhances Large Language Model (LLM) agents in long-horizon software engineering (SWE), sparse binary rewards (Pass/Fail) create a severe credit assignment crisis and waste failed exploratory trajectories. Current trajectory optimization and scaling methods are costly and structurally limited, relying on heuristic state reuse without causal diagnosis or delayed scalar scoring without actionable online guidance. We propose FLARE (Full-Lifecycle Alignment and Reward Engine), a novel dense supervision paradigm driven by a lightweight Generative Reward Model (GRM). First, RADAR, an offline causal-aware diagnostic framework, extracts high-fidelity, hindsight-free supervision through causal-chain backtracking to distill a GRM providing real-time, step-level risk feedback. Second, FLARE uses this GRM to continuously optimize the agent across its entire lifecycle. During inference, FLARE acts as an Active Scaffold, autonomously intercepting high-risk generation steps for localized breakpoint re-execution, drastically reducing compute overhead. During post-training, the GRM's structured signals serve as process-supervised reranking scores for Supervised Fine-Tuning (SFT) and step-level dense rewards for Reinforcement Learning (RL), mitigating policy collapse in sparse environments. Extensive evaluations show that FLARE establishes a new Pareto frontier across the agent lifecycle: FLARE (N=1) outperforms Global Rollout (N=5) with a 5x reduction in token consumption. Extending FLARE to training overcomes the sparse reward problem in long-horizon interactive tasks, delivering relative performance gains of 19.13% in SFT through process-aware data curation and a consistent 9.19% improvement in RL.

cs.CL

On Convergence Behavior of Randomized Kaczmarz-type Methods for Solving Doubly Noisy Linear Systems

In this paper, we investigate the limiting behavior of the RK algorithm for solving doubly noisy inconsistent linear systems without imposing any additional initial assumptions. The proposed bounds effectively characterize the convergence behavior of these algorithms when applied to doubly noisy linear systems. Furthermore, to the best of our knowledge, this work provides the first convergence analysis of the randomized extended Kaczmarz (REK), randomized block Kaczmarz (RBK), and randomized double block Kaczmarz (RDBK) algorithms for doubly noisy linear systems. We prove that these algorithms converge to a neighborhood of the least-squares solution of the underlying noiseless system. These results show that RK and RBK, designed for consistent systems, outperform REK and RDBK, designed for inconsistent systems, in both the convergence rate and the convergence horizon. What's more, the analytical skills to the sketch-and-project method completely overcome the limitation when analyzing RK and RBK. Finally, numerical experiments are conducted to validate the theoretical results.

math.NA

On space-time derivative estimates for the fractional Navier-Stokes equations

In this paper, we are concerned with space-time derivative estimates of solutions to the fractional Navier-Stokes equations. It is shown that $Λ^{nα}u^{(m)}_{t} \in L^{\frac{2(6α-5)}{4mα+2nα+4 α-5}}~~~~~~~(0,T; L^{2}(\mathbb{R}^{3}))$ and $ Λ^{n }u^{(m)}_{t} \in L^{\frac{2(6α-5)}{4mα+2n +4 α-5}}~~~~~(0,T; L^{2}(\mathbb{R}^{3}))$. This generalizes a priori bounds for the classical Navier-Stokes system by Duff in [7, Acta Math. 164, 1990] and Boutros and Gibbon's spatial derivative estimates in [1, Nonlinearity 37, 2024]. In addition, we derive that $ u \in L^{\frac{q}{q-3}}~~(0,T;L^{q} (\mathbb{R}^{3}))$ with $ 6\leq q\leq\infty $ and $Λ^{k}u \in L^{\frac{q}{ q(k+1)-3}}~~~(0,T;L^{q} (\mathbb{R}^{3})) $ with $k\geq1, 2\leq q\leq\infty $ in the standard Navier-Stokes equations.

math.AP

On space-time derivative estimates for the magnetohydrodynamic equations

In this paper, we present the space-time derivative estimates of solutions to the MHD equations. It is a generalization of a priori bounds for the classical Navier-Stokes system by Duff in \cite[Acta Math. 164, 1990]{[Duff]} and gives an affirmative answer to a question proposed by Zheligovsky in \cite[Mathematics. 9, 2021]{[Zheligovsky]}. In addition, we show that $u, B \in L^{\f{q}{q-3}}(0,T;L^{q} (\mathbb{R}^{3}))$ with $6\leq q\leq\infty $ and $ D ^{k}u, D ^{k}B \in L^{\f{q}{ q(k+1)-3}}(0,T;L^{q} (\mathbb{R}^{3}))$ for $k\geq1, 2\leq q\leq\infty$ in this system.

math.AP

MedGEN-Bench: A Contextually Entangled Benchmark for Open-ended Multimodal Medical Generation

Medical vision-language models (VLMs) are increasingly expected to support clinical workflows through diagnostic text and relevant medical images. However, current medical visual benchmarks have three recurring limitations: query-image misalignment from queries weakly grounded in specific image instances, closed-ended formats that narrow answer space and encourage shortcut-based prediction, and text-centric output paradigms that limit evaluation of image-generation and image-editing capabilities. We introduce MedGEN-Bench, a benchmark for open-ended multimodal medical generation. The evaluation snapshot reported in this manuscript comprises 6,422 image-text pairs reviewed by clinical experts and models, spanning 6 canonical imaging modalities, 15 clinical tasks, and 27 named subtasks. It includes 1,100 Visual Question Answering (VQA) pairs, 3,872 Image Editing pairs, and 1,450 Contextual Multimodal Generation pairs. MedGEN-Bench centers on contextual entanglement: dependence of an instruction's intended output on the particular image instance rather than on task wording alone. The benchmark operationalizes this concept through image-grounded instructions and extends evaluation to open-ended multimodal outputs. Its tiered evaluation protocol combines reproducible reference-based fidelity and similarity measures with a structured, checklist-guided assessment by a medical VLM judge. We evaluate 10 compositional frameworks, 2 dedicated image-editing models, 3 unified models, and 5 VLMs. The results show image-output tasks remain unsaturated. Contextual augmentation increases mean image-instruction similarity from 0.273 to 0.372, while a 1,000-case medical-expert audit shows moderate agreement between judge scores and clinician ratings. Source code and dataset are available at https://yangjj007.github.io/medgen.

cs.CV

ALMA Reveals an Explosive Outflow Candidate in IRAS 16119--5048

We present a multiwavelength study of the massive star formation region IRAS 16119-5048 (I16119) using ALMA ATOMS Band 3 and QUARKS Band 6 observations, complemented by archival ATCA radio continuum and Spitzer mid-infrared data. The CO (2-1) emission reveals a system of high-velocity streamer-like structures around the central region. Using dendrogram analysis of velocity-channel maps followed by linking in position-position-velocity space, we identify 16 approximately radially distributed streamers whose projected trajectories converge toward a common central region. The kinetic energy of the outflows is at least an order of magnitude lower than those of most known explosive outflows, while their mass entrainment and momentum rates are high compared with typical protostellar outflows, suggesting that I16119 may represent a low-energy explosive outflow candidate. Dense-gas and photodissociation-region tracers reveal shell-like structures associated with the 8 $μ$m emission, indicating that feedback from the H II region may influence the streamer morphology. The 1.3 mm continuum resolves 27 dense cores along a fragmented filamentary structure. Their separations are consistent with thermal Jeans or cylindrical fragmentation, while the collect-and-collapse scenario is unsupported. The dense cores also show evidence of mass segregation, with the most massive cores concentrated near the inferred explosive centre. We suggest that I16119 is a plausible low-energy explosive outflow candidate, possibly triggered by dynamical interactions among centrally concentrated massive cores. However, the complex velocity structure and possible contamination from individual core-driven outflows prevent a definitive classification. More sensitive, higher angular-resolution observations are required to confirm the nature of the outflow.

astro-ph.GA

ALOHA IRDCs Molecular Line Follow-up: I. Gas properties and kinematics

Infrared Dark Clouds are ideal sites for investigating the initial conditions of massive star and cluster formation. The A Lei Of the Habitat and Assembly of Infrared Dark Clouds (ALOHA IRDCs), a James Clerk Maxwell Telescope (JCMT) Large Program, has mapped nearby IRDCs with SCUBA-2. Complementary molecular line observations are needed to characterise the physical, kinematic, and chemical properties of the dense gas. We aim to determine the thermal, kinematic, and chemical properties of clumps identified in the ALOHA IRDCs, and to assess their evolutionary status and level of star-forming activity. We performed single-pointing K-band and W-band observations towards 56 ALOHA IRDCs clumps using the Effelsberg 100-m and Yebes 40-m telescopes, respectively. We derived NH3 kinetic temperatures using the hyperfine group ratio (HFGR) method and identified infall and shock signatures from HCO+, H13CO+, SiO, and HNCO profiles. Water masers and NH2D emission were used as complementary tracers of chemical evolution and star formation. The clumps exhibit kinetic temperatures of 15-29 K. We detect NH2D emission towards 18 sources, with NH2D centroid velocities consistent with NH3, indicating both species trace the same dense gas component. More than half of the clumps display blue-asymmetric HCO+ profiles, identifying them as infall candidates. Water masers are detected in 22 sources, with prominent velocity ranges and variability. Broad SiO emission (>~20 km/s) indicates strong shocks, while narrower extents (<~6km/s) likely trace large-scale interactions or low-velocity shocks. The widespread infall signatures, shock tracers, masers, and NH2D emission suggest that relatively quiescent, chemically young material can coexist with dynamically active gas affected by early protostellar feedback, providing insight into the coupled physical and chemical evolution of massive IRDC clumps.

astro-ph.GA

The TOP-SCOPE Survey of Planck Galactic Cold Clumps: Molecular gas properties

We surveyed 2008 Planck Galactic Cold Clumps (PGCCs) in $^{12}\mathrm{CO}$ and $^{13}\mathrm{CO}$ $J=1$--0 lines using the Taeduk Radio Astronomy Observatory (TRAO) 14 m telescope's multi-beam receiver. We detected 2784 ($^{12}\mathrm{CO}$) and 2291 ($^{13}\mathrm{CO}$) velocity components, their closely correlated centroid velocities suggest that $^{12}$CO and $^{13}$CO generally trace kinematically associated gas. PGCCs have low excitation temperatures (mean $\sim$10 K), mean $^{13}\mathrm{CO}$ optical depth $\sim$0.5, and mean $^{13}\mathrm{CO}$-derived H$_2$ column density $4.3\times10^{21}$~cm$^{-2}$. Gas--dust correlations are moderate, with $N_{^{13}\mathrm{CO}}$ more tightly correlated with the dust-derived H$_2$ column density from the PGCC catalog than $I_{^{12}\mathrm{CO}}$. Colder PGCCs tend to have higher CO-to-H$_2$ conversion factor ($X_{\mathrm{CO}}$) and $[\mathrm{H_{2}}]/[^{13}\mathrm{CO}]$ ratio. $X_{\mathrm{CO}}$ increases clearly with the dust-derived H$_2$ column density, consistent with enhanced CO freeze-out in high-column-density gas. Supersonic non-thermal motions are widespread: the Mach number derived from $^{13}\mathrm{CO}$ has a mean of 4.3 and a median of 3.6, increasing slightly with dust-derived H$_2$ column density. Overall, PGCCs are cold but dynamically active, serving as a valuable laboratory for studying the initial conditions of star formation.

astro-ph.GA

Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations

Large language model agents operate in partially observable, long-horizon settings where obtaining supervision remains a major bottleneck. We address this by utilizing a source of supervision overlooked in existing post-training methods: unintended yet successful goals embedded within agent rollouts. Specifically, we introduce Hindsight Supervised Learning (HSL), where an auxiliary LLM reviews each completed trajectory and relabels it with all of the natural-language goals the agent actually achieved. HSL then pairs the trajectory with its relabeled goals and uses these pairs for additional fine-tuning. To mitigate suboptimality in the relabeled data, we propose two learning techniques for HSL, irrelevant-action masking and sample reweighting. Our experiments show that HSL is flexible and compatible with existing post-training pipelines. It improves both SFT and DPO, with larger gains on long-horizon tasks with more diverse goal spaces. Moreover, HSL is sample-efficient: on ALFWorld, it surpasses baselines trained on the full dataset while using only one quarter of the ground-truth demonstrations.

cs.CL

Evolving Intelligent Complex Systems via Intellicise Networks: Architecture, Technologies, and Pathways

Future engineering infrastructures are evolving into large-scale, open, heterogeneous, and wirelessly interconnected complex systems. These systems present significant challenges in optimizing network resource utilization, managing high-dimensional information spaces, and accommodating diverse business requirements. Intellicise networks, characterized by Intent-driven operation, semantic-native capability, and distributed intelligence, offer a promising paradigm for enabling such intelligent complex systems. We provide a systematic exploration of future intelligent complex systems from the perspective of intellicise networks. Specifically, we propose a cross-domain intelligent communication network architecture based on intellicise networks, grounded in information theory, systems theory, game theory, and cybernetics. The architecture comprises a cross-layer organizational framework, multi-functional planes, and novel information flows. The cross-layer framework defines the vertical evolution from perception and cognition to decision, while the control, user, data, computation, intelligence, and security planes deliver horizontal intellicise capabilities. Moreover, data, knowledge, model, and task flows interconnect the various layers and planes, forming a closed-loop process that derives simplicity from high-level intelligene while concurrently pursuing enhanced. Building on this architecture, we review key enabling technologies, tracing their evolution from semantic extraction to intent understanding, from heterogeneous resource integration to self-configuration and self-optimization, from generative artificial intelligence (AI) to agentic AI, and from embodied AI to symbodied AI. Additionally, we present a case study on intellicise networks for embodied agent communications and discuss representative applications and services for intelligent complex systems.

eess.SP

Surface code logical operations on a superconducting quantum processor

Fault-tolerant quantum computation requires logical operations that manipulate encoded information while preserving quantum error-correction protection. In planar surface-code architectures, code deformation and lattice surgery provide a local, measurement-based route to such operations. Here we experimentally realize key elements of patch-based surface-code logical processing on a 107-qubit superconducting quantum processor. We first implement a reusable primitive layer comprising merge and split, patch expansion and shrinkage, and deformations mediated by domain walls and twist defects. We then compose these primitives to realize logical state routing, the logical controlled-NOT gate, and the single-qubit Hadamard and phase gates, which together form a Clifford-generating set. All operations are implemented on distance-three rotated surface-code patches with multi-round syndrome extraction and neural-network decoding, without post-selection. Our results advance superconducting surface-code experiments from protected logical memory to active, patch-based fault-tolerant logical operations.

quant-ph

Effective Depth in Joint Source-Channel Coding: An Implicit Equilibrium Analysis

A fundamental design question in deep joint source-channel coding (Deep JSCC) remains insufficiently explored: given a channel signal-to-noise ratio (SNR), what effective computation depth is required for semantic reconstruction? Existing Deep JSCC systems typically employ fixed-depth neural architectures selected through empirical hyperparameter tuning, which may lead to unnecessary computation under favorable channel conditions and insufficient refinement under severe channel noise. This paper proposes \emph{Implicit-JSCC}, an implicit equilibrium framework in which semantic encoding and decoding are formulated as fixed-point equilibrium processes. The effective encoder and decoder depths are determined by residual-based solver convergence rather than manually predefined layer numbers, while parameter sharing across equilibrium iterations enables depth-independent parameter complexity. To analyze the resulting effective-depth behavior, we develop a Gaussian-process-inspired kernel evolution framework that models equilibrium iterations as an effective-depth propagation process. Since channel noise is injected between the encoder and decoder, the analysis tracks channel-induced representation perturbations across receiver-side equilibrium iterations and derives a theory-guided depth--SNR relationship. After offline calibration of the system-specific parameters, the resulting model characterizes the required receiver-side refinement depth under different SNRs. Extensive experiments show that Implicit-JSCC achieves competitive reconstruction performance while enabling residual-based adaptive inference and controllable computation--quality tradeoffs. The depth--SNR model further provides a characterization of the SNR-dependent refinement depth required to reach a prescribed perturbation tolerance.

eess.SP

A 1.3 cm spectral line study of the W33 region

At a distance of 2.4kpc, W33 is one of the most prolific sources of molecular line emission, and it is an excellent research target for a centimeter spectral line search. We carried out a 1.3cm spectral line survey in the frequency range 18-26GHz. The lines we identified include 44 radio recombination lines (RRLs) and 24 molecular lines, excluding transitions from the main isotopolog of NH3. The RRLs are associated with the ionized gas from W33Main. Intensity ratios between RRL pairs with varying differences in the principal quantum number $n$ (i.e., $Δn$) from the same element at adjacent frequencies agree with ratios expected under conditions of local thermodynamical equilibrium. In spite of a resulting helium-to-hydrogen abundance ratio (equal emitting volumes assumed) of (10.7$\pm$1.8)\%, which is consistent with expectations, helium shows broader turbulent line widths than hydrogen. The difference amounts to a few kilometers per second, hinting that the spatial distributions are slightly different. The molecular lines are attributed to nine different species (CH3OH, HC3N, SiS, c-C3H2, CH3CN, NH2D, HNCO, H2O and CCS). Rotation temperatures and column densities were derived from CH3OH transitions using rotational temperature diagram analysis. Maser emission produced by water vapor and methanol have been observed in W33Main, W33A, and W33B. Our survey discovered a CH3OH(10$_{2,8}$-10$_{1,9}$E) maser in W33Main. Toward W33B1, the fractionated deuterium-to-hydrogen ratio (D/H) deduced from para-NH2D/NH3 is estimated to be $\lesssim$(1.0$\pm$0.2)$\times$10$^{-3}$. For the other molecular W33-hotspots, 3$σ$ upper limits are (5.0$\pm$0.4)$\times$10$^{-3}$. At linear scales of (0.5pc), fractional abundances and excitation temperatures do not reach values close to those in well-established hot cores, but higher-resolution measurements may alter this picture.

astro-ph.GA

The ALMA-QUARKS survey: Investigating Thermal Feedback of Massive Protostars in Hot Molecular Cores

We identify a sample of 83 spatially resolved hot molecular cores (HMCs) in the QUARKS survey, aiming at investigating thermal feedback from massive stars. Using CH$_3$CN\,(12--11) line emission together with 1.3\,mm continuum data we derive the radial temperature, volume density and \ch3cn{} abundance profiles for the 83 HMCs. Based on the envelope temperature and density profiles, we compute the luminosities of the embedded massive protostars with \radmc{} radiation transfer model. The derived luminosities are comparable (within $\sim1$ dex) to the bolometric luminosities of their natal clumps and show strong correlations with several core-scale properties, including the HMC mass ($Log[ M_\mathrm{env}] = 1.01\,Log [L_\star] - 4.80$), the inner core radius (the flat radius of Plummer-like volume density profile) ($Log[a] = 0.46\,Log[L_\star] + 0.52$) and the central density $ (Log[n_c] = -0.55 Log[L_\star] +10.47) $. These empirical relations provide useful observational constraints for physical models of protostellar objects. Importantly, we find a strong positive correlation between the massive protostellar luminosity and the local thermal Jeans mass. The derived Jeans masses, $M_\mathrm{Jeans}$, exceed the HMC masses $M_\mathrm{env}$, with the average $M_\mathrm{Jeans}$ being two times larger than the average $M_\mathrm{env}$. This provides observational evidence that thermal feedback from massive protostars can effectively suppress further fragmentation of HMCs, thereby promoting massive star formation. In addition, the positive correlation between massive protostellar luminosity and natal clump mass suggests that more massive clumps preferentially host more luminous protostars, leading to stronger thermal feedback.

astro-ph.GA

Divide and Conquer: Reliable Multi-View Evidential Learning for Deepfake Detection

With the evolution of generative models, deepfakes have achieved near-perfect semantic realism, leaving forensic traces only in subtle structural anomalies. However, existing single-view paradigms often fail to generalize, as dominant semantic features overwhelm subtle artifact cues within entangled representations. This imbalance leads to overconfident yet brittle predictions -- a phenomenon we term the Semantic Masking Effect. To address this challenge, we propose a reliable framework called Divide-and-Conquer Multi-View Evidential Learning (DiCoME) for Deepfake Detection. In the "Divide" phase, we employ Geometric View Purification to decompose the entangled representation space through principled geometric projection. This process suppresses semantic interference within artifact-sensitive representations, forming the foundation for decorrelated yet complementary semantic and artifact views. In the "Conquer" phase, we leverage Uncertainty-Aware Evidential Learning to synthesize these distinct views. By explicitly modeling the "epistemic conflict" between semantic and artifact cues, this mechanism provides calibrated uncertainty estimates instead of forcing rigid deterministic decisions. Extensive experiments across multiple benchmarks demonstrate that our method consistently outperforms existing approaches in generalization performance, while providing reliable uncertainty estimation for trustworthy deepfake detection. Code is available at https://github.com/kxl0825/DiCoME.git.

cs.CV

AMARIS: A Memory-Augmented Rubric Improvement System for Rubric-Based Reinforcement Learning

Rubric-based reward shaping provides interpretable and editable reward signals for fine-tuning LLMs via reinforcement learning (RL), but existing adaptive rubric methods typically update criteria from local evidence such as the current batch or instance-level comparisons. This local view discards diagnostic information produced during training, making it difficult to track recurring failures, evaluate previous rubric edits, or raise standards once earlier criteria become saturated. We introduce AMARIS, A Memory-Augmented Rubric Improvement System that grounds rubric updates in longitudinal training evidence. AMARIS stores rollout analyses, step-level summaries, and rubric update records in a persistent evaluation memory, then retrieves recent and semantically relevant history to revise rubrics. We evaluate AMARIS across science, medicine, instruction following, and creative writing under both global and instance-specific rubric settings. AMARIS improves over static, local-adaptive, and memory-ablated baselines, such as +2.8 points on GPQA-Diamond and +2.2 points on IFBench over the strongest baselines, while analysis shows that memory reduces oscillatory rubric edits and supports a progression from early failure correction to later curriculum advancement. AMARIS runs asynchronously alongside the normal RL loop, reducing blocking latency relative to synchronous rubric updates.

cs.LG

Single-Sample Black-Box Membership Inference Attack against Vision-Language Models via Cross-modal Semantic Alignment

Vision-Language Models (VLMs) have achieved remarkable success, yet their reliance on massive datasets and unintended memorization of training data raise significant data security risk. Membership Inference Attacks (MIAs) aim to assess these risks by determining whether a data sample was included in a model's training set. However, existing MIA methods against VLMs face critical bottlenecks: gray-box method relies on internal logits that are typically restricted in real-world Application Programming Interfaces (APIs), while black-box method depends on large-scale statistical distributions, which struggle in single-sample scenarios. To this end, we investigate MIAs from the perspective of cross-modal semantic alignment, and observe that member images exhibit significantly stronger image-caption alignment due to training memorization, whereas generated captions for non-members may deviate from the original visual content. Leveraging this insight, we propose a novel MIA framework designed for strict black-box and single-sample setting that quantifies such alignment within a joint embedding space, thereby bypassing these unrealistic assumptions. We conducted extensive experiments on three open-source and two closed-source VLMs. On the VL-MIA/Flicker dataset, our method achieves an AUC of 0.821 against LLaVA-1.5, significantly outperforming existing baselines. Furthermore, it remains robust under diverse image perturbations, highlighting its practicality.

cs.CV

Generative design of inorganic materials

Materials discovery is fundamental to advance next-generation technologies as well as for sustainable and circular economy. Beyond computational screening, generative models are efficient at finding materials with desired properties, via multi-modal learning using multiscale data. This perspective examines the landscape of generative design for inorganic materials and discusses the integration of multi-modal learning with high-throughput experimental validation. We contextualize these challenges through the lens of a generative design framework as a unified approach to address the data-driven inverse design of functional materials. The central idea of the framework is constructed around a foundation AI model for inorganic materials interlinked deeply with various property databases and high-throughput experiments via a machine learning driven closed loop, which enables the framework to solve key challenges in functional materials. We argue that domain-specific implementations of such integrated workflows represent a promising pathway toward the unresolved challenge of data-driven inverse design for atom-engineered inorganic functional materials.

cond-mat.mtrl-sci