Search arXivSearch

arXiv subjects

Rui He

Publications and source records attributed to Rui He.

At least 19 recordsLinked to original sources

Distinct dynamics of conceptual and referential disruptions in human reading and large language model processing

Linguistic meaning is grounded in conceptual content, from which reference to particular entities emerges as words enter discourse. To examine the processing dynamics associated with these two dimensions of meaning, we selectively disrupted conceptual or referential information in short narratives and traced the resulting effects in human self-paced reading and in the predictive and representational processing of large language models. In human reading, conceptual disruptions produced a strong but localized processing cost, emerging immediately after the distorted word, reaching an early maximum, and then declining rapidly. Referential disruptions produced weaker effects, which decreased more gradually across subsequent words, and were more strongly modulated by sentence boundaries. In the language model, both disruptions emerged immediately at the manipulated word. Contextual model surprisal showed a pattern closely paralleling human reading: conceptual disruption produced a larger, more locally concentrated effect that decayed rapidly, whereas referential disruption produced a smaller and more gradual downstream effect. Output-layer representations showed a different pattern: referential disruption produced a larger initial displacement, while both distortions were subsequently characterized by power-law decay. Together, these results provide convergent evidence for distinguishable processing dynamics of two types of meaning: conceptual information imposes a more locally concentrated integration cost, whereas referential information engages a more distributed process of maintaining discourse-level identity.

cs.CL

The Changing Geometry of Grammar: Dimensionality and Neighborhood Reorganization across Transformer Layers

Transformer representations describe trajectories through high-dimensional vector spaces, which are shaped dynamically as tokens incorporate relational context across layers. Such data tend to concentrate on lower-dimensional sub-manifolds, a form of compression quantified by the Intrinsic Dimensionality (ID), the minimum number of independent variables needed to represent them without significant information loss. In this work, we ask whether the grammatical role of tokens, as marked by their part-of-speech (PoS) tag, shapes the local geometry of this manifold. To this end: (1) We investigate the layer-wise evolution of ID, finding that closed-class items expand earlier and collapse sooner than open-class ones; (2) We show its expansion and contraction to be explained by changes in the neighborhood structure, and hence in the relations between words within a sentence; (3) We compare encoders (ModernBERT, bigbird-roberta-large) and decoders (gemma-2-2B, Llama-3.2-3B), finding that the two families evolve differently across layers, consistently with how each integrates context;(4) We show that geometric features alone recover a token's grammatical role, and use them to interpret how the semantic content of each PoS evolves across layers in a downstream classification task.

cs.CL

TDD-Agent: Test-Driven Reasoning for Code Generation

Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level tasks remains challenging. Existing approaches often use generated tests as static post-hoc validators, which limits their ability to guide implementation and may introduce misleading feedback when the tests themselves are incomplete or incorrect. In this paper, we introduce TDD-Agent, which operationalizes the test-driven development paradigm for code generation. TDD-Agent first prompts the model to generate executable tests, encouraging it to clarify expected behaviors before implementation, and then performs iterative dual-track refinement over both the generated code and tests using execution feedback. We first isolate the effect of test-first reasoning through a prompt variant TDD-prompt on LiveCodeBench, where it consistently improves upon reasoning-based prompting baselines. Building on this finding, we evaluate the full TDD-Agent framework on RepoEval, a repository-level benchmark, and show that it consistently outperforms retrieval-based and agent-based baselines. Additional analyses show that iterative refinement improves not only code correctness but also the effectiveness of the generated tests, yielding higher pass rates, coverage, and mutation scores, suggesting that tests can serve as evolving reasoning artifacts rather than fixed validators. Our source code is available at https://anonymous.4open.science/r/TDD-Agent-Framework-6370/.

cs.SE

Activation-Guided Neuron Intervention to Induce Alzheimer's-Related Computational Language Phenotypes in a Large Language Model

Changes in spontaneous speech provide an early signal of cognitive dysfunction in Alzheimer's disease (AD) that large language models (LLMs) can detect. However, detection alone cannot establish whether the underlying model representations contribute functionally to behavior. We introduce an activation-guided intervention framework using Qwen3-8B. The framework identifies feed-forward neurons with higher activation rates for AD than control transcripts and modulates their output contributions during generation by scaling the corresponding down-projection weights. This yielded nine edited variants differing in intervention direction, magnitude, and scope. The original and edited models completed the same 12-turn neuropsychological battery, assessed through blinded human ratings and computational linguistic measures. Amplifying AD-associated neurons produced graded impairments in story recall, verbal fluency, working memory, procedural discourse, scene construction, and coreference resolution. Attenuation largely preserved performance and selectively improved several outcomes. Amplification also reduced lexical surprisal, idea density, syntactic complexity, and discourse quantity, broadly paralleling changes reported in human AD speech. These findings show that neurons identified solely from clinical language differences can influence behavior across multiple cognitive domains, providing proof of concept for an AD-related computational phenotype and a controlled framework for experimentally examining links between language and broader cognitive dysfunction.

cs.CL

Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion Segmentation

White-light imaging (WLI) and narrow-band imaging (NBI) provide complementary views of endoscopic lesions, but their paired observations are often spatially misaligned due to viewpoint changes, tissue deformation, and sequential handheld acquisition. This makes direct WLI/NBI fusion prone to mixing non-corresponding regions and may even degrade segmentation around lesion boundaries. To address this problem, we propose a reliability-aware complex-domain fusion framework for paired-but-unregistered WLI/NBI lesion segmentation. The framework first establishes topology-regularized feature correspondence and further estimates where the cross-modal correspondence is reliable. Guided by this reliability, the model selectively fuses WLI and NBI features in a learnable complex representation. In this representation, WLI-derived cues mainly provide appearance-related magnitude responses, while NBI-derived cues provide structure-sensitive phase responses. Unlike conventional real-valued or symmetric multimodal fusion, the proposed method explicitly models the different roles of WLI and NBI and suppresses unreliable cross-modal interaction in locally mismatched regions. Experiments on paired WLI/NBI endoscopic datasets show that the proposed reliability-aware registration grounding and complex-domain fusion consistently improve lesion segmentation performance. Role-reversal and module ablation studies further validate the necessity of both the modality-role design and reliability-guided cross-modal interaction.

cs.CV

A refined blow-up analysis of the Brezis-Nirenberg equation and its application: The one-bubble case for $N\geq4$

In this paper, we consider the famous Brezis-Nirenberg equation \begin{eqnarray*} \left\{ \aligned &-\Delta u=\lambda u+|u|^{\frac{4}{N-2}}u,\quad&\mbox{in}\,\, \Omega,\\ &u=0,\quad&\mbox{on}\,\, \partial\Omega, \endaligned \right. \end{eqnarray*} where $N\geq3$ is the dimension, $\Omega\subset\mathbb{R}^N$ is a bounded domain with smooth boundary $\partial\Omega$ and $\lambda>0$ is a parameter. By developing a refined blow-up analysis based on the inverse reduction argument developed in \cite{WW2019,WW2019-2}, we classify, for the fist time, the Struwe decomposition of the Brezis-Nirenberg equation in the one-bubble case as the parameter $\lambda$ varies for $N\geq4$. As applications, we prove that the $4d$ Brezis-Nirenberg equation has a nontrivial solution (least energy solution) for $\lambda\in\sigma(-\Delta)$ in general bounded domains, where $\sigma(-\Delta)$ is the spectrum of $-\Delta$ in $H^1_0(\Omega)$. Our result completes the existence theory of the Brezis-Nirenberg equation for $N\geq4$ in \cite{AP2025,CFP1985,CFS,CSS1986,CW2005,CSZ2012,SWW2009,TYZ2022} since 1984.

math.AP

DiffusionVS: A Generative Framework for Robust Visual Servoing Based on Diffusion Policy

Visual servoing is a fundamental technique in robotic manipulation and navigation. Regression-based visual servoing frequently experiences trajectory jitter as a result of noise-sensitive single-step mappings and the accumulation of errors during distribution shifts. In contrast, Diffusion Policy maintains temporal consistency by predicting action sequences and improves robustness through implicit data augmentation. This paper presents a novel diffusion-based servoing method. Based on Diffusion Policy, the proposed approach uses normalized image coordinates of observed tag corners as input and generates camera velocity through conditional denoising. To overcome the generalization limitations of models trained on static datasets, an online training paradigm is adopted, continuously expanding the diversity of training data through interactive experience collection. This strategy substantially enhances both the performance and generalization capability of the model. Comprehensive simulations and real-world experiments demonstrate the effectiveness of the proposed method, achieving success rates of nearly 100\% in simulation and 93\% in physical experiments. Beyond the specific pipeline, we further validate the generality of the diffusion mechanism. Experiments show that existing visual servoing networks consistently achieve improved performance when integrated with our diffusion-based module. These results indicate that the proposed strategy possesses broad applicability and can enhance various visual servoing systems beyond the specific architecture presented here.

cs.RO

Cross-lingual brain-language model alignment is robust but challenges hierarchical and computational accounts

Brain-language model alignment is often interpreted as evidence that transformer models implement computations similar to those of the human brain. This assumes that neural predictivity reflects internal computational properties of large language models (LLMs), such as hierarchical contextual processing, predictive coding, or representational compression. An alternative possibility is that brain scores primarily reflect stable lexical-semantic correspondences shared by language models and the brain. Here we tested these interpretations using whole-brain encoding models across Mandarin, English, and French. Across all three languages, transformer representations significantly predicted activity in a distributed network spanning classical language regions, transmodal cortical systems, and subcortical structures. These spatial patterns showed substantial cross-linguistic overlap and remained remarkably stable across layers, providing little evidence that model depth systematically maps onto cortical processing hierarchies. Likewise, contextual transformer embeddings did not consistently outperform static lexical embeddings, despite providing some unique predictive variance. Finally, neither surprisal nor intrinsic dimensionality reproduced the layer-wise profile of brain scores, arguing against prediction and information compression as primary explanations for brain-LLM alignment. Together, these findings suggest that brain-LLM alignment is more robust across languages, transformer depth, and model architectures than previously appreciated, but less informative about shared computational mechanisms. Our results are more consistent with neural predictivity reflecting stable representational structure preserved across model transformations than with a one-to-one correspondence between their underlying computations.

cs.CL

The grip of grammar on meaning uncertainty: cross-linguistic evidence, neural correlates, and clinical relevance

Isolated word meanings are inherently uncertain. This uncertainty reduces when they are combined and anchored in context. We propose that grammar compresses meaning uncertainty cross-linguistically, which is reflected in brain and selectively disrupted in disorders. Compression was operationalized as the relative difference between non-contextual surprisal estimated from lexical frequency, and contextual surprisal from grammar-sensitive models. In narratives from 20 languages, contextual surprisal reduced frequency-based surprisal. This reduction closely tracked the surprisal cost of reversing word order, and scaled with richer, non-redundant lexis as organized by more complex but optimal dependency structure. During fMRI, surprisal and its reduction explained BOLD activity for comprehension and production in overlapping but distinct regions. Uncertainty reduction was significantly attenuated in aphasia, dementia, and schizophrenia, but remained intact where primary deficit is not language. These findings position uncertainty reduction via grammar as a foundational concept that illuminates principles, brain basis, and disruptions of language.

cs.CL

Cross Fusion and Correlation Beamformer for Row-Column Array Based 3D Ultrasound Imaging

Row column addressed (RCA) transducers present a promising solution for ultrafast volumetric imaging with a reduced channel count and a large field of view. However, RCA-based 3D imaging is fundamentally limited by severe sidelobe artifacts and a low signal-to-noise ratio (SNR), primarily due to weak transmit focusing inherent in RCA based ultrafast imaging strategies. To overcome these challenges, we propose a cross fusion and correlation (CFAC) method that leverages the incoherence of sidelobe artifacts and noise across datasets acquired using orthogonal apertures and multiple steering angle sets. The performance of the proposed method was validated through simulations, in vitro imaging of a multi-purpose ultrasound phantom, and in vivo experiments, and benchmarked against four established techniques: orthogonal plane wave (OPW) imaging, XDoppler method, row-column-specific frame-multiply-and-sum beamforming (RC-FMAS), and coherent factor (CF) imaging. Simulation results demonstrated that CFAC reduced sidelobe levels by 42.0 dB, 38.9 dB, 28.3 dB, and 25.5 dB compared to OPW, XDoppler, RC-FMAS, and CF, respectively. In phantom experiments, CFAC improved the CNR by up to 17.5 dB. Furthermore, in vivo imaging of a rat kidney showed that CFAC enables visualization of a significantly more detailed microvascular network, achieving a CNR improvement of over 25 dB against all benchmarked methods. In conclusion, the proposed CFAC method effectively suppresses sidelobe artifacts and noise in RCA-based imaging under low-SNR conditions, enabling high-contrast 3D visualization while preserving the high frame rate capabilities of ultrafast ultrasound imaging.

physics.med-ph

A collaborative agent with two lightweight synergistic models for autonomous crystal materials research

Current large language models require hundreds of billions of parameters yet struggle with domain-specific reasoning and tool coordination in materials science. Here, we present MatBrain, a lightweight collaborative agent system with two synergistic models specialization for crystal materials research. MatBrain employs a dual-model architecture: Mat-R1 (30B parameters) as the analytical model providing expert-level domain reasoning, and Mat-T1 (14B parameters) as the executive model orchestrating tool-based actions. Entropy analysis confirms that this architecture resolves the conflict between tool planning and analytical reasoning by decoupling their distinct entropy dynamics. Enabled by this dual-model architecture and structural efficiency, MatBrain significantly outperforms larger general-purpose models while reducing the hardware deployment barrier by over 95%. MatBrain exhibits versatility across structure generation, property prediction, and synthesis planning tasks. Applied to catalyst design, MatBrain generated 30,000 candidate structures and identified 38 promising materials within 48 hours, achieving approximately 100-fold acceleration over traditional approaches. These results demonstrate the potential of lightweight collaborative intelligence for advancing materials research capabilities.

cs.AI

Exceptional Optical Phonon Coherence in Enriched Cubic Boron Arsenide via Suppression of Three-Phonon Scattering

Cubic boron arsenide (BAs) is a promising semiconductor for next-generation electronics due to its outstanding ambipolar mobility and thermal conductivity, the latter of which is attributed to the suppression of three-phonon scattering. However, precisely accounting for different high-order anharmonic scattering processes is challenging from both theory and experiment, so that questions remain open regarding the ultimate limit of phonon lifetime and thermal conductivity in BAs. Here we show that this gap nearly eliminates three-phonon scattering for zone-center optical phonons in a wide temperature range, leading to a record-high, isotope purity-limited phonon coherence with a quality factor above $3.7\times 10^3$ for >98% enriched $^{11}$BAs below 100 K. We discriminate three decoherence mechanisms by their temperature-dependent contribution to the damping rate using high-resolution Raman and Fourier transform infrared spectroscopy. For the as-synthesized crystals, we find that defect scattering has negligible contributions to the linewidth of optical phonons in comparison to isotope scattering. These results provide critical insights into the intrinsic and extrinsic scattering mechanisms of optical phonons in BAs, motivating further studies to quantify anharmonic effects and realize superior phonon transport.

cond-mat.mtrl-sci

Polarization-Differential Loss Enabled High Polarization Extinction in Hollow-Core Fibers

Delivering a well defined state of polarization over hollow core fibres (HCFs) is pivotal for next generation ultra stable photonic systems. Yet in all existing HCFs, whether birefringent or not, their polarization extinction ratio (PER) rapidly deteriorates during propagation or under mechanical disturbance, leaving no practical high and stable PER solution. Here, we break this impasse by embedding a polarization differential loss (PDL) mechanism directly into the cladding architecture.

physics.optics

Diffusion-Guided Mask-Consistent Paired Mixing for Endoscopic Image Segmentation

Augmentation for dense prediction typically relies on either sample mixing or generative synthesis. Mixing improves robustness but misaligned masks yield soft label ambiguity. Diffusion synthesis increases apparent diversity but, when trained as common samples, overlooks the structural benefit of mask conditioning and introduces synthetic-real domain shift. We propose a paired, diffusion-guided paradigm that fuses the strengths of both. For each real image, a synthetic counterpart is generated under the same mask and the pair is used as a controllable input for Mask-Consistent Paired Mixing (MCPMix), which mixes only image appearance while supervision always uses the original hard mask. This produces a continuous family of intermediate samples that smoothly bridges synthetic and real appearances under shared geometry, enlarging diversity without compromising pixel-level semantics. To keep learning aligned with real data, Real-Anchored Learnable Annealing (RLA) adaptively adjusts the mixing strength and the loss weight of mixed samples over training, gradually re-anchoring optimization to real data and mitigating distributional bias. Across Kvasir-SEG, PICCOLO, CVC-ClinicDB, a private NPC-LES cohort, and ISIC 2017, the approach achieves state-of-the-art segmentation performance and consistent gains over baselines. The results show that combining label-preserving mixing with diffusion-driven diversity, together with adaptive re-anchoring, yields robust and generalizable endoscopic segmentation.

cs.CV

Tunable symmetry breaking in a hexagonal-stacked moir\'e magnet

Symmetry plays a central role in defining magnetic phases, making tunable symmetry breaking across magnetic transitions highly desirable for discovering non-trivial magnetism. Magnetic moir\'e superlattices, formed by twisting two-dimensional (2D) magnetic crystals, have been theoretically proposed and experimentally explored as platforms for unconventional magnetic states. However, despite recent advances, tuning symmetry breaking in moir\'e magnetism remains limited, as twisted 2D magnets, such as rhombohedral (R)-stacked twisted CrI_3, largely inherit the magnetic properties and symmetries of their constituent layers. Here, in hexagonal-stacked twisted double bilayer (H-tDB) CrI_3, we demonstrate clear symmetry evolution as the twist angle increases from 180^{\circ} to 190^{\circ}. While the net magnetization remains zero across this twist angle range, the magnetic phase breaks only the three-fold rotational symmetry at 180^{\circ}, but it breaks all of the rotational, mirror, and time-reversal symmetries at intermediate twist angles between 181^{\circ} and 185^{\circ}, and all broken symmetries are recovered at 190^{\circ}. These pronounced symmetry breakings at intermediate twist angles are accompanied by metamagnetic behaviors, evidenced by symmetric double hysteresis loops around zero magnetic field. Together, these results reveal that H-tDB CrI_3 at intermediate twist angles host a distinct moir\'e magnetic phase, featuring periodic in-plane spin textures with broken rotational, mirror, and time-reversal symmetries, which is markedly different from the out-of-plane layered antiferromagnetism in bilayer CrI_3 and the predominantly out-of-plane moir\'e magnetism in R-tDB CrI_3. Our work establishes H-stacked CrI_3 moir\'e magnets as a versatile platform for engineering magnetic properties, including and likely beyond complex spin textures.

cond-mat.mtrl-sci

SPAC: A Python Package for Spatial Single-Cell Analysis of Multiplex Imaging

Multiplexed immunofluorescence microscopy captures detailed measurements of spatially resolved, multiple biomarkers simultaneously, revealing tissue composition and cellular interactions in situ among single cells. The growing scale and dimensional complexity of these datasets demand reproducible, comprehensive and user-friendly computational tools. To address this need, we developed SPAC (SPAtial single-Cell analysis), a Python-based package and a corresponding shiny application within an integrated, modular SPAC ecosystem (Liu et al., 2025) designed specifically for biologists without extensive coding expertise. Following image segmentation and extraction of spatially resolved single-cell data, SPAC streamlines downstream phenotyping and spatial analysis, facilitating characterization of cellular heterogeneity and spatial organization within tissues. Through scalable performance, specialized spatial statistics, highly customizable visualizations, and seamless workflows from dataset to insights, SPAC significantly lowers barriers to sophisticated spatial analyses.

cs.SE

Melting of Charge Density Waves in Low Dimensions

Charge density waves (CDWs) are collective electronic states that can reshape and melt, even while confined within a rigid atomic crystal. In two dimensions, melting is predicted to be distinct, proceeding through partially ordered nematic and hexatic states that are neither liquid nor crystal. Here we measure and explain how continuous, hexatic melting of incommensurate CDWs occurs in low-dimensional materials. As a CDW is thermally excited, disorder emerges progressively$\unicode{x2013}$initially through smooth elastic deformations that modulate the local wavelength, and subsequently via the nucleation of topological defects. Experimentally, we track three hallmark signatures of CDW melting$\unicode{x2013}$azimuthal superlattice peak broadening, wavevector contraction, and integrated intensity decay.

cond-mat.mtrl-sci

Long-lived Zone-boundary Magnons in an Antiferromagnet

Antiferromagnetic (AFM) insulators exhibit many desirable features for spintronic applications such as fast dynamics in the THz range and robustness to fluctuating external fields. However, large damping typically associated with THz magnons presents a serious challenge for THz magnonic applications. Here, we report long-lived short-wavelength zone boundary magnons in the honeycomb AFM insulator CoTiO3, recently found to host topological magnons. We find that its zone-boundary THz magnons exhibit longer lifetimes than its zone-center magnons. This unusual momentum-dependent long magnon lifetime originates from several factors including the antiferromagnetic order, exchange anisotropy, a finite magnon gap, and magnon band dispersion. Our work suggests that magnon-magnon interaction may not be detrimental to magnon lifetimes and should be included in future searches for topological magnons.

cond-mat.mes-hall