Search arXivSearch

arXiv subjects

Zheng Yu

Publications and source records attributed to Zheng Yu.

At least 19 recordsLinked to original sources

An All-Sky Catalog of 6.5 Million Primary Red Clump Stars from Gaia DR3 XP Spectra

Red clump (RC) stars are excellent standard candles for mapping the three-dimensional structure of the Milky Way. We construct an all-sky catalog of 6.5 million primary RC stars using spectroscopic estimates of asteroseismic parameters inferred from Gaia DR3 low-resolution XP spectra. We train a mixture density network (MDN) on a cross-matched sample. The network maps each 343-dimensional corrected XP spectrum to estimates of $\Delta\Pi_{1}$, the asymptotic period spacing of dipole gravity modes, and $\Delta\nu$, the large frequency separation. We select primary RC stars using these two estimates. We release two complementary catalogs defined by different parameter-space cuts and thresholds: a high-purity Tier 1 sample of 532,189 stars with 97% purity and a high-completeness Tier 2 sample of 6,534,931 stars with 86% completeness. Both catalogs cover the full celestial sphere. We provide distances and extinction values for all stars. We obtain a median photometric distance precision of 6% for Tier 1 RC stars. Using the resultant distances, we independently calibrate the Gaia DR3 parallax zero-point offset. The three-dimensional density distribution traced by the Tier 2 sample extends continuously to $R\approx25$ kpc. Both catalogs are publicly available. These catalogs provide a valuable resource for Galactic archaeology, delivering a homogeneous dataset to trace the chemo-dynamical evolution of the Milky Way and to calibrate models of stellar populations.

astro-ph.SR

Temperature-Induced Reorganization of Supported Zn$_3$ Clusters on Cu(111): From Minimum-Energy Structures to Finite-Temperature Ensembles

Understanding the nature of catalytic active sites under reaction conditions remains a central challenge in heterogeneous catalysis. In industrial copper/zinc oxide/alumina catalysts for methanol synthesis, small Zn-based species at the Cu interface have long been proposed as active-site candidates, yet their atomic-scale structure and stability remain controversial. Computational studies typically identify such species from optimized 0 K structures, assuming that minimum-energy configurations remain representative under reaction conditions. Here, we combine machine-learning-interatomic-potential-accelerated global optimization, molecular dynamics, and enhanced-sampling free-energy calculations to investigate supported Zn$_3$(OH)$_3$ and Zn$_3$(OH)$_2$CHOO clusters on Cu(111)-based surfaces from 0 to 450 K. While compact triangular configurations are generally favored among minimum-energy structures at 0 K, finite-temperature free-energy calculations reveal a pronounced shift toward extended linear configurations with increasing temperature. This transition is driven primarily by entropic stabilization and cannot be inferred from potential energies alone. Molecular dynamics further shows substantial cluster mobility on pristine Cu(111), indicating that long-term persistence depends not only on configurational stability but also on surface mobility. Surface Zn alloying strongly suppresses diffusion, thereby stabilizing isolated interfacial Zn species. Together, these results show that thermodynamically relevant structures of supported Zn-based clusters can differ fundamentally from static 0 K predictions because of competing enthalpic and entropic effects. Our findings highlight the limitations of identifying catalytic active sites solely from 0 K structures and underscore the importance of explicit finite-temperature sampling in catalyst modeling.

cond-mat.mtrl-sci

On the Limitations of Non-GPU AI Accelerators for Large-Model Inference: A Field Study of MoE and Multimodal Serving on Huawei Ascend

Non-GPU AI accelerators are increasingly adopted as alternatives to general-purpose GPUs for large-model inference, but the real engineering cost of migrating demanding workloads beyond CUDA remains poorly documented. We present a field study of deploying two large inference workloads on a 16-device Huawei Ascend 910 system using CANN and vLLM-Ascend: an LLM-as-a-judge safety and alignment evaluation pipeline based on a W8A8 MoE judge model, DeepSeek-V4-Flash, and a multimodal medical vision--language benchmark based on DeepSeek-V4-Flash-Vision for MMMU and MMMU-Pro. Making these workloads reliable required twelve source-level patches to the vendor inference plugin, disabling several high-throughput features to preserve numerical correctness, and adding operational safeguards for recurring device-level failures. We summarize the main platform limitations in eight categories: incomplete operator and feature support, fragile parallelism, numerical faults in low-level kernels, immature graph compilation, unstable advanced features, limited scalability, weak observability, and ecosystem fragmentation. For each category, we report the symptoms, evidence, and likely causes. We also quantify the integration effort, concurrency behavior, and benchmark quality to show that both workloads were served correctly. Our study provides a reproducible reference for teams evaluating or operating non-GPU accelerators for large-model inference.

cs.DC

From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation

Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external APIs, including adaptation to platform-specific data distributions, policy-specific objectives, and product-level safety constraints. As a result, platforms must undertake internal model development, naturally turning to shared public research for guidance. However, existing multimodal foundation-model studies primarily report architectures, training recipes, data scaling strategies, and benchmark results, but provide less systematic guidance on how failures should be localized and translated into targeted model-development interventions. Interventions are essential because deployment failures are rarely self-explanatory. Similar failures can originate from different causes. Without targeted interventions, improvement reduces to heuristic trial-and-error, where benchmark improvements are weakly attributable, and failures are difficult to trace to their underlying causes. To address this gap, we present a diagnostic methodology for industry-scale Audio-Visual-Language Models AVLM development. The methodology maps model failures into a taxonomy of observable failure signatures and links each class of failure to an intervention space. We instantiate this methodology across the development and alignment lifecycle of an AVLM foundation model for a large-scale video and live-streaming platform. The resulting system supports over 100 regions and is designed for noisy, ambiguous, and highly diverse content drawn from global platform traffic.

cs.LG

From Bulk to Surface: Structure and Dynamics of Amorphous Alumina from Deep Potential Molecular Dynamics

Understanding the atomic-scale structure and dynamics of amorphous oxide surfaces is essential for interpreting their chemical reactivity, mechanical stability, and interfacial behavior, yet direct experimental characterization remains challenging. We employ Deep Potential (DP) molecular dynamics to generate large-scale, ab initio-quality models of amorphous Al$_2$O$_3$ bulk glasses and melt-quenched free surfaces, enabling a quantitative analysis of both structure and relaxation dynamics with statistical confidence inaccessible to direct ab initio simulation. The trained DP model reproduces experimental liquid and glass structure, captures the cooling-rate dependence of the bulk glass transition, and corrects systematic biases in the polyhedral populations predicted by widely used classical force fields. At the free surface, mass density recovers to bulk values over ~10 $\unicode{x212B}$, while local coordination requires a slightly wider subsurface region to fully converge. The outermost layer is oxygen-enriched, exhibits altered polyhedral connectivity with contracted Al-O bonds, and hosts a broad population of under-coordinated motifs (notably AlO$_3$ and OAl$_2$) whose abundances are governed by glass stability. These reactive Lewis acid and Br$\unicode{x00F8}$nsted base sites are locally paired in a manner consistent with bond-valence compensation, yet remain spatially dispersed rather than aggregating into extended clusters. Despite this pronounced structural heterogeneity, the surface relaxes on the same timescale as the bulk and exhibits a comparable glass transition temperature, suggesting that the disordered surface is kinetically stable once formed. Together, these results establish a molecular-level picture of amorphous alumina surfaces and demonstrate the capability of machine-learned potentials to resolve structure-property relationships in disordered oxide interfaces.

cond-mat.mtrl-sci

Hindsight Preference Optimization for Financial Time Series Advisory

Time series models predict numbers; decision-makers need advisory -- directional signals with reasoning, actionable suggestions, and risk management. Training language models for such predictive advisory faces a fundamental challenge: quality depends on outcomes unknown at prediction time. We bridge two ideas from reinforcement learning -- using information unavailable during execution to retrospectively generate training signal, and preference alignment -- and propose Hindsight Preference Optimization: observed outcomes let an LLM judge rank candidate advisories on dimensions that scalar metrics cannot capture, producing preference pairs for DPO without human annotation. We apply this to Vision-Language-Model-based predictive advisories on S&P 500 equity time series, demonstrated by a 4B model outperforming its 235B teacher on both accuracy and advisory quality.

cs.LG

Patch Validation in Automated Vulnerability Repair

Automated Vulnerability Repair (AVR) systems, especially those leveraging large language models (LLMs), have demonstrated promising results in patching vulnerabilities -- that is, if we trust their patch validation methodology. Ground-truth patches from human developers often come with new tests that not only ensure mitigation of the vulnerability but also encode extra semantics such as root cause location, optimal fix strategy, or subtle coding styles or conventions. And yet, none of the recent AVR systems verify that the auto-generated patches additionally pass these new tests (termed as $\text{PoC}^+$ tests). This is a subtle yet critical omission. To fill this gap, we constructed a benchmark, $\textrm{PVBench}$, with 209 cases spanning 20 projects. Each case includes basic tests (functional tests before the patch and the PoC exploit) as well as the associated $\text{PoC}^+$ tests. Evaluated on three state-of-the-art AVR systems, we find that over 40\% of patches validated as correct by basic tests fail under $\text{PoC}^+$ testing, revealing substantial overestimation on patch success rates. Analyzing these patches that are falsely labeled as correct, we suggest that AVR tools should improve in three critical areas: root cause analysis, adherence to program specifications, and capturing developer intention.

cs.SE

From LLMs to Agents: A Comparative Evaluation of LLMs and LLM-based Agents in Security Patch Detection

The widespread adoption of open-source software (OSS) has accelerated software innovation but also increased security risks due to the rapid propagation of vulnerabilities and silent patch releases. In recent years, large language models (LLMs) and LLM-based agents have demonstrated remarkable capabilities in various software engineering (SE) tasks, enabling them to effectively address software security challenges such as vulnerability detection. However, systematic evaluation of the capabilities of LLMs and LLM-based agents in security patch detection remains limited. To bridge this gap, we conduct a comprehensive evaluation of the performance of LLMs and LLM-based agents for security patch detection. Specifically, we investigate three methods: Plain LLM (a single LLM with a system prompt), Data-Aug LLM (data augmentation based on the Plain LLM), and the ReAct Agent (leveraging the thought-action-observation mechanism). We also evaluate the performance of both commercial and open-source LLMs under these methods and compare these results with those of existing baselines. Furthermore, we analyze the detection performance of these methods across various vulnerability types, and examine the impact of different prompting strategies and context window sizes on the results. Our findings reveal that the Data-Aug LLM achieves the best overall performance, whereas the ReAct Agent demonstrates the lowest false positive rate (FPR). Although baseline methods exhibit strong accuracy, their false positive rates are significantly higher. In contrast, our evaluated methods achieve comparable accuracy while substantially reducing the FPR. These findings provide valuable insights into the practical applications of LLMs and LLM-based agents in security patch detection, highlighting their advantage in maintaining robust performance while minimizing false positive rates.

cs.CR

PortGPT: Towards Automated Backporting Using Large Language Models

Patch backporting, the process of migrating mainline security patches to older branches, is an essential task in maintaining popular open-source projects (e.g., Linux kernel). However, manual backporting can be labor-intensive, while existing automated methods, which heavily rely on predefined syntax or semantic rules, often lack agility for complex patches. In this paper, we introduce PORTGPT, an LLM-agent for end-to-end automation of patch backporting in real-world scenarios. PORTGPT enhances an LLM with tools to access code on-demand, summarize Git history, and revise patches autonomously based on feedback (e.g., from compilers), hence, simulating human-like reasoning and verification. PORTGPT achieved an 89.15% success rate on existing datasets (1815 cases), and 62.33% on our own dataset of 146 complex cases, both outperforms state-of-the-art of backporting tools. We contributed 9 backported patches from PORTGPT to the Linux kernel community and all patches are now merged.

cs.CR

RAG Meets Temporal Graphs: Time-Sensitive Modeling and Retrieval for Evolving Knowledge

Knowledge is inherently time-sensitive and continuously evolves over time. Although current Retrieval-Augmented Generation (RAG) systems enrich LLMs with external knowledge, they largely ignore this temporal nature. This raises two challenges for RAG. First, current RAG methods lack effective time-aware representations. Same facts of different time are difficult to distinguish with vector embeddings or conventional knowledge graphs. Second, most RAG evaluations assume a static corpus, leaving a blind spot regarding update costs and retrieval stability as knowledge evolves. To make RAG time-aware, we propose Temporal GraphRAG (TG-RAG), which models external corpora as a bi-level temporal graph consisting of a temporal knowledge graph with timestamped relations and a hierarchical time graph. Multi-granularity temporal summaries are generated for each time node to capture both key events and broader trends at that time. The design supports incremental updates by extracting new temporal facts from the incoming corpus and merging them into the existing graph. The temporal graph explicitly represents identical facts at different times as distinct edges to avoid ambiguity, and the time hierarchy graph allows only generating reports for new leaf time nodes and their ancestors, ensuring effective and efficient updates. During inference, TG-RAG dynamically retrieves a subgraph within the temporal and semantic scope of the query, enabling precise evidence gathering. Moreover, we introduce ECT-QA, a time-sensitive question-answering dataset featuring both specific and abstract queries, along with a comprehensive evaluation protocol designed to assess incremental update capabilities of RAG systems. Extensive experiments show that TG-RAG significantly outperforms existing baselines, demonstrating the effectiveness of our method in handling temporal knowledge and incremental updates.

cs.IR

Predicting Drug-Drug Interactions Using Heterogeneous Graph Neural Networks: HGNN-DDI

Drug-drug interactions (DDIs) are a major concern in clinical practice, as they can lead to reduced therapeutic efficacy or severe adverse effects. Traditional computational approaches often struggle to capture the complex relationships among drugs, targets, and biological entities. In this work, we propose HGNN-DDI, a heterogeneous graph neural network model designed to predict potential DDIs by integrating multiple drug-related data sources. HGNN-DDI leverages graph representation learning to model heterogeneous biomedical networks, enabling effective information propagation across diverse node and edge types. Experimental results on benchmark DDI datasets demonstrate that HGNN-DDI outperforms state-of-the-art baselines in prediction accuracy and robustness, highlighting its potential to support safer drug development and precision medicine.

cs.LG

Empirically Predicted Absolute Magnitudes for Red Clump Stars in Mephisto and CSST Filters

Red clump (RC) stars are reliable standard candles for studying the structure and evolution of the Milky Way. In this study, we present empirical calibrations of RC absolute magnitudes in the Mephisto (v, g, r, i) and CSST (g, r, i) photometric systems using a high-purity sample of 25,059 RC stars cross-matched between APOGEE and Gaia DR3 XP spectra. Through synthetic photometry and polynomial fitting, we find that RC absolute magnitudes exhibit strong dependencies on effective temperature and metallicity, with the strongest variations observed in bluer bands and progressively decreasing towards redder wavelengths. In particular, the Mephisto v band exhibits the highest sensitivity, with variations reaching up to 2.0 mag across the metallicity range (-1.0 dex <[Fe/H] <0.5 dex) and the temperature range (4500-5200 K). The calibrations achieve high precision for all bands, enabling accurate determination of RC absolute magnitudes and distances. Furthermore, we evaluate the metallicity estimation capabilities of both systems using a Random Forest-based method, achieving a precision of 0.12 dex for Mephisto and 0.14 dex for CSST under typical photometric uncertainties (< 0.01 mag). These results provide robust tools for distance and metallicity determinations, supporting future Galactic structure studies with Mephisto and CSST data.

astro-ph.SR

Stellar parameters, extinction, and distances for stars in SMSS DR2 by SPar method

The availability of large datasets containing stellar parameters, distances, and extinctions for stars in the Milky Way, particularly within the Galactic disk, is essential for advancing our understanding of the Galaxy's stellar populations, structure, kinematics, and chemical evolution. In this study, we present a catalog of stellar parameters, including effective temperature (\teff), metallicity (\feh), absolute magnitudes ($M_{G}$), distances ($d$), and reddening values (\ebr), for a sample of 141 million stars from the SkyMapper Southern Survey (SMSS). These parameters are derived using the SPar algorithm, which employs a fitting procedure to match multi-band photometric observations and estimate the stellar properties (\teff, \feh, $M_G$, $d$, and \ebr) on an individual star basis, following the methodology outlined in our previous work. This study successfully determines stellar parameters and extinction values simultaneously for stars located in high and low Galactic latitudes. The resulting stellar parameters and extinction values are in good agreement with results from other studies, demonstrating the robustness of our method. We derive a temperature dispersion of 195\,K and a metallicity dispersion of 0.31\,dex when comparing our results with spectroscopic data. The catalog produced in this work provides a valuable resource for future research on the Galactic metallicity distribution function, the structure of the Galaxy, three-dimensional extinction mapping, and other related astrophysical topics.

astro-ph.GA

The Stellar Disk Structure Rrevealed by the Mono-age Populations of the LAMOST Red Clump Sample

Understanding the structure of the Galactic disk is crucial for understanding the formation and evolutionary history of the Milky Way. This study examines the structure of the Galactic disk by analyzing a sample of 138,667 primary red clump (RC) stars from the LAMOST and Gaia datasets. We have categorized these RC stars into mono-age populations and investigated their spatial distributions within the R - Z plane, estimating scale heights and lengths through the fitting of their vertical and radial density profiles. Our analysis indicates that the vertical profiles of these mono-age populations fit a dual-component disk model, where both components exhibit significant flaring, particularly in the outer disk regions. Within a constant Galactocentric radius R, the scale heights of the first component, representing the morphologically thin disk, rise with age. In contrast, the scale heights of the second component, corresponding to the morphologically thick disk, remain comparatively stable across different age groups. Additionally, the radial density profiles of both disk components predominantly peak within a radial range of 7.5-8.5 kpc. These findings underscore the importance of age as a crucial factor in shaping the spatial distribution and structural evolution of the Galactic disk, offering valuable insights into its complex dynamics and history.

astro-ph.GA

FuseAnyPart: Diffusion-Driven Facial Parts Swapping via Multiple Reference Images

Facial parts swapping aims to selectively transfer regions of interest from the source image onto the target image while maintaining the rest of the target image unchanged. Most studies on face swapping designed specifically for full-face swapping, are either unable or significantly limited when it comes to swapping individual facial parts, which hinders fine-grained and customized character designs. However, designing such an approach specifically for facial parts swapping is challenged by a reasonable multiple reference feature fusion, which needs to be both efficient and effective. To overcome this challenge, FuseAnyPart is proposed to facilitate the seamless "fuse-any-part" customization of the face. In FuseAnyPart, facial parts from different people are assembled into a complete face in latent space within the Mask-based Fusion Module. Subsequently, the consolidated feature is dispatched to the Addition-based Injection Module for fusion within the UNet of the diffusion model to create novel characters. Extensive experiments qualitatively and quantitatively validate the superiority and robustness of FuseAnyPart. Source codes are available at https://github.com/Thomas-wyh/FuseAnyPart.

cs.CV

ShadowBound: Efficient Heap Memory Protection Through Advanced Metadata Management and Customized Compiler Optimization

In software development, the prevalence of unsafe languages such as C and C++ introduces potential vulnerabilities, especially within the heap, a pivotal component for dynamic memory allocation. Despite its significance, heap management complexities have made heap corruption pervasive, posing severe threats to system security. While prior solutions aiming for temporal and spatial memory safety exhibit overheads deemed impractical, we present ShadowBound, a unique heap memory protection design. At its core, ShadowBound is an efficient out-of-bounds defense that can work with various use-after-free defenses (e.g. MarkUs, FFMalloc, PUMM) without compatibility constraints. We harness a shadow memory-based metadata management mechanism to store heap chunk boundaries and apply customized compiler optimizations tailored for boundary checking. We implemented ShadowBound atop the LLVM framework and integrated three state-of-the-art use-after-free defenses. Our evaluations show that ShadowBound provides robust heap protection with minimal time and memory overhead, suggesting its effectiveness and efficiency in safeguarding real-world programs against prevalent heap vulnerabilities.

cs.CR

CAMP: Compiler and Allocator-based Heap Memory Protection

The heap is a critical and widely used component of many applications. Due to its dynamic nature, combined with the complexity of heap management algorithms, it is also a frequent target for security exploits. To enhance the heap's security, various heap protection techniques have been introduced, but they either introduce significant runtime overhead or have limited protection. We present CAMP, a new sanitizer for detecting and capturing heap memory corruption. CAMP leverages a compiler and a customized memory allocator. The compiler adds boundary-checking and escape-tracking instructions to the target program, while the memory allocator tracks memory ranges, coordinates with the instrumentation, and neutralizes dangling pointers. With the novel error detection scheme, CAMP enables various compiler optimization strategies and thus eliminates redundant and unnecessary check instrumentation. This design minimizes runtime overhead without sacrificing security guarantees. Our evaluation and comparison of CAMP with existing tools, using both real-world applications and SPEC CPU benchmarks, show that it provides even better heap corruption detection capability with lower runtime overhead.

cs.CR

Shortest paths govern fracture nucleation in thermoset networks

In this work, we introduce a predictive approach for determining fracture nucleation in thermosets based on shortest paths (SPs) of the network topology. This method enumerates SP sets in networks with periodic boundary conditions, with applications to both all-atom and coarse-grained simulations. We find that network fracture is most likely to nucleate on the first (shortest) SP and the strain at nucleation is linearly correlated with the topological path length. Subsequent fracture events are dictated by the instantaneous SP of partially broken networks. We further quantify the length scale dependence of SP distributions, introducing a means of bridging simulated and experimental fracture profiles.

cond-mat.soft