Search arXivSearch

arXiv subjects

Zehao Lin

Publications and source records attributed to Zehao Lin.

At least 19 recordsLinked to original sources

Star Formation in the H II Region Sh 2-205: 3D Morphology and Kinematics from Young Stars and Molecular Gas

Using Gaia astrometry of young stars combined with CO observations, we present the first systematic three-dimensional (3D) analysis of the structure, kinematics, and evolutionary history of the star-forming regions in the environs of the H II region Sh 2-205 (S205). S205 exhibits a complex morphology and coherent expansion on both global and subregional scales. We identify several O9-B1 stars and a 0.56 Myr old pulsar that are likely associated with the region. A momentum estimate suggests that feedback from these objects may account for the observed overall expansion. Trace-back analysis of the expansion, combined with color-magnitude diagram fitting for young star clusters, indicates at least two episodes of star formation. These results reveal a complex star-formation history of S205 and provide new insights into its 3D evolution.

astro-ph.GA

200 mm Wafer-Scale Monolithic 3D Integration of Atomic Layer-Deposited Oxide Semiconductors

Monolithic 3D (M3D) integration offers a pathway to overcome the scaling limits of conventional silicon complementary metal-oxide-semiconductor (CMOS) technology by extending dense vertical stacking of multifunctional logic and memory devices. Here, we demonstrate wafer-scale M3D integration of three tiers of atomic-layer-deposited (ALD) indium oxide (InOx)-based devices (>100,000 fabricated), including ferroelectric, enhancement-mode, and depletion-mode field-effect transistors, on 200 mm silicon wafers. We achieve threshold voltage standard deviation as low as 0.04 V, average electron mobility up to 91.6 cm2V-1s-1, and fully functional cross-tier circuits. A four-tier 3D computing-in-memory (CIM) accelerator targeting large language model workloads is developed using a custom InOx process design kit, delivering 1.4x to 2.9x speedup and comparable energy-delay product improvements over 2D baselines. These results establish ALD InOx M3D integration as a scalable and CMOS-compatible platform for next-generation artificial intelligence hardware and advanced electronics.

physics.app-ph

Metis: Memory Foundation Model

Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily implemented through external modules, leaving the native memory capability largely unexplored. In this paper, we take a first step toward this direction by introducing memory foundation models, which empower foundation models with native memory capabilities. We formalize native memory from two perspectives: a persistent and dynamically evolving memory state within the backbone, and native memory procedures that autonomously store and utilize information through model computation. We show that native memory offers advantages in architecture, end-to-end optimization, and efficiency. Based on this formulation, we propose Metis, the first prototype of memory foundation models. Metis introduces a new architecture that equips a foundation model with a native memory state, allowing historical information to be compressed into the model and accessed through memory attention. We construct large-scale memory-specific training data and introduce multiple optimization objectives to acquire these native memory procedures through mid-training. The online memory maintenance of Metis is gradient-free, and the memory update requires only a forward pass. At inference time, all learned model weights remain frozen, while the native memory states are autonomously transformed through standard forward computation. Through extensive experiments, we show that Metis exhibits native memory capabilities and further provide a detailed analysis of its strengths, limitations, and behaviors. To facilitate future research on memory foundation models, we release our project and model checkpoints.

cs.CL

MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations

Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions. Existing benchmarks, however, evaluate such memory almost exclusively through downstream question answering, scoring only the correctness of a final answer. This black-box formulation conflates the heterogeneous causes of memory failure, such as missing the introduction of a relevant fact, binding an operation to the wrong target, or relying on stale values after a correction. As a result, it can credit correct answers despite their reliance on inconsistent or unsafe memory states. In this paper, we argue that, in dynamic long-horizon interactions, memory is not a static collection of facts but a lifecycle of explicit operations, including remembering, forgetting, updating, reflecting, and their compositions. We introduce MemOps, a benchmark that reformulates conversational memory as a sequence of lifecycle operations and represents each memory event with a structured trace specifying its trigger, target, scope, state transition, and supporting evidence. A controllable generation pipeline embeds these operations into long, task-oriented conversations and produces gold operation traces together with six categories of operation-level probes, evaluated under both adjacent-evidence and long-context settings. Across long-context, retrieval-based, parametric and managed-memory systems, MemOps disentangles failure modes that final-answer accuracy alone conceals, revealing that current systems remain far from uniformly reliable. For instance, session-level retrieval outperforms turn-level retrieval, and long-context models remain notably weak at reconstructing ordered memory-state trajectories. These results move long-term memory evaluation from final-answer scoring toward interpretable, operation-level diagnosis.

cs.AI

Mapping the Milky Way with Masers

SKA-VLBI is poised to revolutionize our understanding of the Galactic structure through its unprecedented astrometric precision and sensitivity. As a next-generation facility, it will answer long-standing questions about the Galactic structure by mapping its entire spiral structure in detail, spanning from the solar neighborhood, through the Galactic Center, to the far side of the Milky Way. Its access to the Southern sky will allow us to obtain more precise 3D parameters of the Galactic bar, reveal the nature of the 3-kpc Arm, and clarify the dynamical coupling between the bar and the spiral arms. By leveraging high-precision astrometry of numerous celestial objects with SKA-VLBI, the Galactic fundamental parameters such as the Solar motion and the Galactic rotation curve can be constrained more precisely. These advancements will not only elucidate the structure of our Milky Way, but also provide benchmarks for understanding barred spiral galaxies in general. Furthermore, they are important for advancing our knowledge of cosmological structure formation. The capabilities of SKA-VLBI will open a new era of high precision Galactic astrometry.

astro-ph.GA

SparseX: Efficient Segment-Level KV Cache Sharing for Interleaved LLM Serving

In long-context LLM serving, the prefill stage often dominates time-to-first-token and computational cost. Although Prefix Cache in vLLM/PagedAttention has been widely used to reuse identical prompt prefixes, repeated content in practical applications frequently appears as non-prefix, cross-request, cross-turn, and cross-agent segments, which makes conventional cache mechanisms insufficient. This paper presents SparseX, a segment-level KV Cache sharing method for common serving scenarios. SparseX uses contiguous token segments as reuse units and exploits Sparse-Q indices that naturally arise in KV Cache reuse workloads to estimate the key tokens that require correction. Based on this estimate, SparseX performs Sparse-KV Recomputation within a single forward pass, thereby restoring cross-segment contextual interactions under complex interleaved reuse patterns while avoiding additional models or separate preprocessing stages for token selection. SparseX further implements a full+sparse hybrid attention mode based on a layer-specific threshold: early layers retain full attention to obtain a more stable token-importance signal, and later layers switch to sparse recomputation to improve reuse quality on complex long-context tasks. We implement SparseX-vLLM on top of vLLM, integrating segment-level cache lookup, PagedAttention management, RoPE alignment, Sparse-Q token selection, and FlashAttention backends into a unified execution path. SparseX is model-agnostic, training-free, and compatible with Prefix Cache, and it provides unified support for common online serving scenarios including multi-round chat, retrieval-augmented generation (RAG), and agent workflows.

cs.PF

East Asian VLBI Network astrometry toward the star-forming region G040.96+02.48 in the Extreme Outer Galaxy

Accurate astrometric measurements for star-forming regions located on the far side of the Milky Way remain scarce. In this work, we present the astrometric results for a 22\,GHz water maser associated with star-forming region G040.96+02.48 located on the far side of the Milky Way, using the East Asian VLBI Network. The target water maser's proper motion was determined to be ($\mu_{\alpha}\cos\delta, \mu_{\delta}$) = ($-2.06_{-0.51}^{+0.53}$, $-2.95_{-0.44}^{+0.45}$)~mas~yr$^{-1}$. The derived three-dimensional kinematic distance to the star-forming region is 20.2$\pm$3.2\,kpc, placing it slightly outside the Outer Scutum$-$Centaurus Arm. The corresponding vertical height of 872$\pm$139\,pc indicates a significant warp of the outer Galactic disk, which is in good agreement with the latest precessing warp model. Moreover, the resulting peculiar motions reveal a complex kinematic pattern, characterized by a large outward radial velocity of $-32\pm$18\,km~s$^{-1}$. Our observations substantially expand the valuable sample of star-forming regions with accurate astrometric measurements in the Extreme Outer Galaxy.

astro-ph.GA

A Survey on Long-Term Memory Security in LLM Agents: Attacks, Defenses, and Governance Across the Memory Lifecycle

The emergence of writable, cross-session persistent memory in LLM agents introduces a qualitatively different threat landscape from conventional input-centric security concerns, characterized by three properties: persistence, statefulness, and propagation. To systematically characterize this landscape, we propose a Memory Lifecycle Framework that organizes attacks, defenses, and their cross-phase dependencies along two axes: six lifecycle phases (Write, Store, Retrieve, Execute, Share & Propagate, Forget & Rollback) and four security objectives (Integrity, Confidentiality, Availability, Governance). This analysis in turn exposes the need for formal security guarantees at the system level, motivating Verifiable Memory Governance(VMG), a framework of five architectural primitives that specifies what verifiable mechanisms a long-term-memory system must provide to maintain auditable, recoverable control over its memory state. Our analysis indicates that robust Long-Term Memory (LTM) security cannot be retrofitted at retrieval or execution time alone, but must be anchored in storage-time provenance, versioning, and policy-aware retention from the outset.

cs.CR

Breakdown of Ohm's Law by Disorders in Low-Dimensional Transistors

Ohm's law provides a fundamental framework for understanding charge transport in conductors and underpins the concept of electrical scaling that has enabled the continuous advancement of modern CMOS technologies. As transistors are scaled to even smaller dimensions, device channels inevitably enter low-dimensional regimes to achieve higher performance. Low-dimensional materials such as atomically thin oxide semiconductors, 2D van der Waals semiconductors, and 1D carbon nanotubes, have thus emerged as key candidates for extending Moore's law. Here, we reveal the fundamental distinction between three-dimensional and low-dimensional conductors arising from disorder-induced electron localization, which leads to the breakdown of Ohm's law and lateral linear scaling. We develop a quantitative model that captures the role of the disordered region, a unique characteristic intrinsically to low-dimensional transistors. Furthermore, the disorder-induced localization framework consistently explains experimental observations in atomically thin In2O3 field-effect transistors across variations in channel length, temperature, thickness, and post-annealing conditions. This work establishes a unified physical picture for understanding and optimizing disorder-driven electronic transport in low-dimensional transistors.

cond-mat.mes-hall

TAdaRAG: Task Adaptive Retrieval-Augmented Generation via On-the-Fly Knowledge Graph Construction

Retrieval-Augmented Generation (RAG) improves large language models by retrieving external knowledge, often truncated into smaller chunks due to the input context window, which leads to information loss, resulting in response hallucinations and broken reasoning chains. Moreover, traditional RAG retrieves unstructured knowledge, introducing irrelevant details that hinder accurate reasoning. To address these issues, we propose TAdaRAG, a novel RAG framework for on-the-fly task-adaptive knowledge graph construction from external sources. Specifically, we design an intent-driven routing mechanism to a domain-specific extraction template, followed by supervised fine-tuning and a reinforcement learning-based implicit extraction mechanism, ensuring concise, coherent, and non-redundant knowledge integration. Evaluations on six public benchmarks and a real-world business benchmark (NowNewsQA) across three backbone models demonstrate that TAdaRAG outperforms existing methods across diverse domains and long-text tasks, highlighting its strong generalization and practical effectiveness.

cs.CL

How Does Environmental Information Disclosure Affect Corporate Environmental Performance? Evidence from Chinese A-Share Listed Companies

Global climate warming and air pollution pose severe threats to economic development and public safety, presenting significant challenges to sustainable development worldwide. Corporations, as key players in resource utilization and emissions, have drawn increasing attention from policymakers, researchers, and the public regarding their environmental strategies and practices. This study employs a two-way fixed effects panel model to examine the impact of environmental information disclosure on corporate environmental performance, its regional heterogeneity, and the underlying mechanisms. The results demonstrate that environmental information disclosure significantly improves corporate environmental performance, with the effect being more pronounced in areas of high population density and limited green space. These findings provide empirical evidence supporting the role of environmental information disclosure as a critical tool for improving corporate environmental practices. The study highlights the importance of targeted, region-specific policies to maximize the effectiveness of disclosure, offering valuable insights for promoting sustainable development through enhanced corporate transparency.

econ.GN

TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models

Many robotic manipulation tasks require sensing and responding to force signals such as torque to assess whether the task has been successfully completed and to enable closed-loop control. However, current Vision-Language-Action (VLA) models lack the ability to integrate such subtle physical feedback. In this work, we explore Torque-aware VLA models, aiming to bridge this gap by systematically studying the design space for incorporating torque signals into existing VLA architectures. We identify and evaluate several strategies, leading to three key findings. First, introducing torque adapters into the decoder consistently outperforms inserting them into the encoder.Third, inspired by joint prediction and planning paradigms in autonomous driving, we propose predicting torque as an auxiliary output, which further improves performance. This strategy encourages the model to build a physically grounded internal representation of interaction dynamics. Extensive quantitative and qualitative experiments across contact-rich manipulation benchmarks validate our findings.

cs.RO

Innovator: Scientific Continued Pretraining with Fine-grained MoE Upcycling

A large language model (LLM) with knowledge in both scientific and general tasks is the foundation of science general intelligence. However, directly continued pretraining an LLM using science data usually leads to catastrophic forgetting, which indicates severe degradation in general ability. In this report, we present Innovator, which solves this problem by upcycling a pre-trained dense LLM into a fine-grained Mixtures-of-Experts model during continued pretraining, where different experts are expected to learn science knowledge in different disciplines, and a shared expert is utilized for general tasks. Innovator introduces a four-stage upcycle training paradigm: (1) Scientific Expert Induction on discipline-specific data, (2) Fine-grained Expert Splitting via FFN dimension decomposition, (3) Science-Aware Routing warmup, and (4) Generalist-Scientist Integration training on hybrid datasets. Such a paradigm enables knowledge in the general domain, and different scientific disciplines can be decoupled, avoiding the negative influence among knowledge in different domains. With 53.3B total parameters and 13.3B activated, Innovator extends Qwen2.5-7B using a shared general expert and 64 specialized scientific experts with 8 activated. Trained on 300B tokens with tri-level quality-controlled data, Innovator achieves 25% average improvement across 30 scientific tasks with a win rate as 70%, while retaining 99% performance in general tasks. Furthermore, Innovator-Reason, which is post-trained from Innovator for reasoning boosting, exhibits excellent reasoning performance in solving complex scientific problems with improvements over 30%.

cs.LG

MemOS: A Memory OS for AI System

Large Language Models (LLMs) have become an essential infrastructure for Artificial General Intelligence (AGI), yet their lack of well-defined memory management systems hinders the development of long-context reasoning, continual personalization, and knowledge consistency.Existing models mainly rely on static parameters and short-lived contextual states, limiting their ability to track user preferences or update knowledge over extended periods.While Retrieval-Augmented Generation (RAG) introduces external knowledge in plain text, it remains a stateless workaround without lifecycle control or integration with persistent representations.Recent work has modeled the training and inference cost of LLMs from a memory hierarchy perspective, showing that introducing an explicit memory layer between parameter memory and external retrieval can substantially reduce these costs by externalizing specific knowledge. Beyond computational efficiency, LLMs face broader challenges arising from how information is distributed over time and context, requiring systems capable of managing heterogeneous knowledge spanning different temporal scales and sources. To address this challenge, we propose MemOS, a memory operating system that treats memory as a manageable system resource. It unifies the representation, scheduling, and evolution of plaintext, activation-based, and parameter-level memories, enabling cost-efficient storage and retrieval. As the basic unit, a MemCube encapsulates both memory content and metadata such as provenance and versioning. MemCubes can be composed, migrated, and fused over time, enabling flexible transitions between memory types and bridging retrieval with parameter-based learning. MemOS establishes a memory-centric system framework that brings controllability, plasticity, and evolvability to LLMs, laying the foundation for continual learning and personalized modeling.

cs.CL

Scalable Complexity Control Facilitates Reasoning Ability of LLMs

The reasoning ability of large language models (LLMs) has been rapidly advancing in recent years, attracting interest in more fundamental approaches that can reliably enhance their generalizability. This work demonstrates that model complexity control, conveniently implementable by adjusting the initialization rate and weight decay coefficient, improves the scaling law of LLMs consistently over varying model sizes and data sizes. This gain is further illustrated by comparing the benchmark performance of 2.4B models pretrained on 1T tokens with different complexity hyperparameters. Instead of fixing the initialization std, we found that a constant initialization rate (the exponent of std) enables the scaling law to descend faster in both model and data sizes. These results indicate that complexity control is a promising direction for the continual advancement of LLMs.

cs.LG

Xinyu AI Search: Enhanced Relevance and Comprehensive Results with Rich Answer Presentations

Traditional search engines struggle to synthesize fragmented information for complex queries, while generative AI search engines face challenges in relevance, comprehensiveness, and presentation. To address these limitations, we introduce Xinyu AI Search, a novel system that incorporates a query-decomposition graph to dynamically break down complex queries into sub-queries, enabling stepwise retrieval and generation. Our retrieval pipeline enhances diversity through multi-source aggregation and query expansion, while filtering and re-ranking strategies optimize passage relevance. Additionally, Xinyu AI Search introduces a novel approach for fine-grained, precise built-in citation and innovates in result presentation by integrating timeline visualization and textual-visual choreography. Evaluated on recent real-world queries, Xinyu AI Search outperforms eight existing technologies in human assessments, excelling in relevance, comprehensiveness, and insightfulness. Ablation studies validate the necessity of its key sub-modules. Our work presents the first comprehensive framework for generative AI search engines, bridging retrieval, generation, and user-centric presentation.

cs.IR

MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models

Large Language Models (LLMs) have emerged as foundational infrastructure in the pursuit of Artificial General Intelligence (AGI). Despite their remarkable capabilities in language perception and generation, current LLMs fundamentally lack a unified and structured architecture for handling memory. They primarily rely on parametric memory (knowledge encoded in model weights) and ephemeral activation memory (context-limited runtime states). While emerging methods like Retrieval-Augmented Generation (RAG) incorporate plaintext memory, they lack lifecycle management and multi-modal integration, limiting their capacity for long-term knowledge evolution. To address this, we introduce MemOS, a memory operating system designed for LLMs that, for the first time, elevates memory to a first-class operational resource. It builds unified mechanisms for representation, organization, and governance across three core memory types: parametric, activation, and plaintext. At its core is the MemCube, a standardized memory abstraction that enables tracking, fusion, and migration of heterogeneous memory, while offering structured, traceable access across tasks and contexts. MemOS establishes a memory-centric execution framework with strong controllability, adaptability, and evolvability. It fills a critical gap in current LLM infrastructure and lays the groundwork for continual adaptation, personalized intelligence, and cross-platform coordination in next-generation intelligent systems.

cs.CL

Calibrating the Color-Magnitude Relation of M Giants by Using Open Clusters

M giants, with their distinctive properties such as high luminosity, serve as excellent indicators for mapping the structure of the Milky Way. The distance to distant M giants can be determined by using the color-magnitude relation (CMR), which is derived from color-magnitude diagrams of specific systems in previous studies. In this work, we aimed to achieve more accurate distance determination for M giants by focusing on open clusters (OCs) with a large number of member stars and thus improve the CMR. For the first time, we compiled a census of OCs harboring M giants using Gaia Data Release 3 (DR3) and Large Sky Area Multi-Object Fiber Spectroscopic Telescope Data Release 9. We identified 58 M giants associated with 43 OCs and obtained their astrometric and photometric parameters from Gaia DR3. Using the distances of these OCs, we derived the CMR for M giants as a linear correlation, expressed as $M_{Ks}=3.85-8.26(J-K_s$). This linear relation proved superior to the empirical distance relation in characterizing the CMR of M giants. The photometric distances of M giants derived from the CMR are consistent with the parallax distances from Gaia and known spectroscopic distances, with median deviations of 1.5% and 2.3%, respectively. Using the distances of M giants derived from the CMR, we computed their radial velocity ($V_R$), azimuthal velocity ($V{\phi}$), and vertical velocity ($V_Z$), respectively. The distributions of these velocities revealed key features of the Galactic disk, including oscillation, north-south rotational asymmetry, and warp. These findings are consistent with previous studies and further validate the reliability of the derived CMR.

astro-ph.SR