Search arXivSearch

arXiv subjects

Yan Zhu

Publications and source records attributed to Yan Zhu.

At least 19 recordsLinked to original sources

Gaze responses to false-positive computer-aided detection prompts during colonoscopy: a paired-video and real-time eye-tracking study

False-positive computer-aided detection (CADe) prompts may divert endoscopists' attention during colonoscopy, yet the attentional impact of individual prompts remains unclear. We used event-locked eye tracking to quantify gaze attraction and attention occupation in complementary retrospective and prospective studies. In a retrospective paired-video experiment, 3 senior and 2 novice endoscopists viewed 60 colonoscopy videos with and without CADe. The prospective study recorded gaze during 42 real-time CADe-assisted colonoscopies performed by 9 senior endoscopists. Screened CADe prompts outside expert-annotated lesion windows were classified as false-positive artifact events. False-positive prompts attracted gaze in 48.6% (68/140) of retrospective observations and 65.2% (533/817) of prospective events. Among attraction events with complete recovery, median attention occupation lasted 1000 ms in the retrospective study and 1100 ms in the prospective study. Corresponding median prompt durations were 33 ms and 267 ms, with median time amplifications of 17.55-fold and 5.15-fold, respectively. In paired retrospective comparisons, visible artifact prompts drew gaze closer to the prompted region than did the same-coordinate unassisted reference. Secondary retrospective analyses showed high lesion gaze recognition without and with CADe (98.0% versus 99.0%). First gaze entry into lesion regions occurred 147.8 ms earlier with CADe. Across controlled and real-time clinical settings, false-positive CADe prompts frequently captured gaze, with attention persisting beyond prompt visibility. These findings support considering prompt-related attentional burden in CADe evaluation and design.

cs.AI

Extraordinary Lifetime Enhancement of Coherent Phonon-Amplitude Modes in the Excitonic Insulator Phase of Ta2Pd3Te5

The excitonic insulator (EI) is an electronic phase of condensed excitons. However, in many prototypical materials the presence of a concurrent structural transition complicates the identification of a purely electronic origin of the ordered state. Here, we investigate Ta2Pd3Te5, in which an EI phase develops below TC ~ 100 K in the absence of a detectable structural transition. We characterize the excitonic condensate via its coherent phonon-amplitude response. Most strikingly, an extraordinarily strong lifetime enhancement of these modes sets in below TC which we establish as a new and robust fingerprint linked to exciton condensation. We discuss possibilities to capture the enhancement when taking into account a coupling of excitonic and lattice effects in line with the coupled phonon-amplitude response. Within an order-parameter polaron picture coupling to the excitonic state may dress the phonon mode and thereby suppress its relaxation and a phenomenological model combines the anharmonic phonon response and a reduced electronic scattering due to the gap opening in the EI phase.

cond-mat.str-el

How Calibration Content Shapes Attention-Based Reranking

Attention-based rerankers score documents by aggregating query-to-document attention and subtracting a null-query calibration pass to remove positional and structural bias. Although widely used, this calibration assumes that the null pass removes irrelevant signal from each document. We show that modern prompt content, e.g. constraints, instructions, personas, and demonstrations can violate this assumption when it enters the scoring readout, making the null pass relevance-aware rather than null. We find that calibration is especially harmful when applied to prompts containing longer, more detailed instructions as the null-pass step removes relevant signal. Based on these findings, we propose interpolated null calibration, a training-free modification that controls how much of the instruction content enters the null baseline. It recovers attention-based reranking performance on instruction-heavy tasks where standard calibration fails, while preserving calibration's benefits when the null pass remains relevance-agnostic. On instruction heavy tasks, the recovered rankings surpass generative rerankers. We also show that in-context demonstrations improve attention-based reranking with little calibration interference, since demonstrations act only through the query pass and leave the null pass unchanged.

cs.CL

ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement

Query-conditioned vision--language models enable fine-grained interpretation by revealing how visual evidence changes with textual queries. However, evidence conditioned on complete descriptions does not necessarily resolve into object-specific evidence, nor does an exposed evidence map necessarily identify the evidence that constitutes the model's prediction. Across multiple VLM architectures and independent benchmarks, we find that object-level queries often retain evidence from co-occurring objects and shared context. In this paper, we introduce \textbf{ProtoLIP}, a lightweight prototype-mediated evidence layer that organizes reusable visual prototypes into text-derived semantic families and uses query-dependent family routing to constrain which prototypes may provide evidence. Without spatial annotations or backbone retraining, ProtoLIP improves evidence localization and separation across query granularities, with localization gains transferring to independently pretrained VLMs with well-aligned patch--text representations. Despite using only text-derived weak supervision, ProtoLIP remains competitive with a spatially supervised grounding model while maintaining strong matching and competitive image--text retrieval. Crucially, ProtoLIP constructs its matching score directly from localized prototype evidence, enabling the score to be exactly decomposed into semantic-family and prototype contributions.

cs.CV

From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM

Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record. General-purpose vision-language models (VLMs) offer a unified interface for image understanding and report generation. Existing specialization strategies, however, typically rely on task-specific models or model-weight adaptation, leaving unresolved how to introduce reliable specialist knowledge while preserving both this unified interface and the VLM's pretrained capabilities. We introduce a context-fusion framework that specializes a frozen general-purpose VLM through both implicit instruction context and explicit transduction context without modifying its pretrained weights. Specifically, a self-supervised polyp encoder retrieves related image-report pairs as explicit, query-specific evidence, while learned continuous specialist tokens provide implicit instruction context shared across cases. Experiments were conducted on 2,056 expert-annotated public endoscopic images. We compared the framework with general-purpose VLMs, task-specific predictors, and weight-adaptation methods to assess specialist performance, unified reporting, and adaptation efficiency. Across numerical, categorical, and report-generation metrics, the proposed framework substantially improved direct frozen-VLM inference and achieved the strongest overall performance among the evaluated methods. It added trainable parameters equal to only 0.006% of the frozen VLM's parameter count. When the top-1 retrieved case carried the correct target category, our framework corrected 70.5% of the errors made by a weight-adaptation baseline. These findings support the context-fusion framework as a lightweight and effective strategy for specialist adaptation of a frozen VLM.

cs.AI

$P$-polynomial coherent configurations

Suda introduced the notion of a $Q$-polynomial coherent configuration, which provides a natural and important concept. Subsequently, Lato introduced a notion of a $P$-polynomial coherent configuration and proved that every such configuration satisfying the definition has at most two fibers. Although Lato's definition is interesting, particularly because it characterizes distance-biregular graphs, we argue that an alternative definition is desirable. In this paper, we propose an alternative notion of $P$-polynomial coherent configurations that is naturally aligned with Suda's $Q$-polynomial framework. We show that every two-fiber coherent configuration that is $P$-polynomial in Lato's sense is also $P$-polynomial in our sense, whereas the converse does not hold. We further prove that every coherent configuration of type $(2,2;3)$, $(3,2;3)$ or $(3,3;3)$ is $P$-polynomial in our sense. In addition, we present three families of $P$-polynomial coherent configurations with an arbitrary number of fibers: those arising from tight Euclidean $t$-designs in $\mathbb R^2$, the Terwilliger algebra of $H(n,2)$, and the set of all subspaces of $\mathbb F_q^n$. Finally, we give an equivalent condition for the cross-block intersection matrices to be tridiagonal and verify that all three families satisfy this condition.

math.CO

Understanding and Correcting Low-Frequency Bias in EEG Foundation Model

Increasing EEG pretraining data scale or model capacity does not consistently improve downstream performance. We identify a persistent low-frequency bias in representations learned by diverse EEG foundation models, which remains across dataset scales, model capacities, and pretraining objectives. Our analysis links this bias to the interaction between EEG's $1/f^α$-like spectral structure and neural networks' tendency to preferentially learn low-frequency components. In masked autoencoders, the $\ell_2$ reconstruction objective further amplifies this imbalance: under comparable relative reconstruction errors, high-power low-frequency components contribute disproportionately to the loss. To address this issue, we introduce FAME, a frequency-balanced masked autoencoding framework that reconstructs time--frequency activity in predefined EEG bands from masked EEG inputs. FAME independently standardizes the reconstruction targets within each band and assigns equal weight to all band-specific losses, thereby balancing supervision across the EEG spectrum. Evaluated on 41 downstream tasks in OmniEEG-Bench, FAME learns more spectrally balanced representations and achieves state-of-the-art performance on 24 of them. These results underscore the importance of balanced spectral supervision for learning transferable EEG representations.

cs.LG

Stabilization of zigzag order in NiPS$_3$ via positive biquadratic interaction

Despite extensive research, the precise spin Hamiltonian of the van der Waals antiferromagnet NiPS$_3$ -- which hosts a zigzag-ordered ground state -- remains debated. While consensus has emerged on ferromagnetic nearest-neighbor ($J_1$) and antiferromagnetic third-nearest-neighbor ($J_3$) Heisenberg interactions, recent studies suggest a biquadratic ($B$) exchange term may also play a role, though its estimated magnitude varies widely. To address this controversy, we perform density functional theory calculations and extract a positive biquadratic interaction with $B/J_3 \approx 0.44$. Within the minimal $J_1$-$J_3$-$B$ model, we show that these parameters naturally stabilize zigzag ordering using minimally augmented spin-wave theory. Density-matrix renormalization group calculations further validate our extracted parameters as a reasonable description of the ground state. Although fully resolving the spin Hamiltonian of NiPS$_3$ requires further investigation, our findings provide new insights into its biquadratic interaction.

cond-mat.str-el

A report-grounded vision-language foundation model for colonoscopy from 280000 routine reports

Vision-language models remain underused in colonoscopy despite the rich expert descriptions recorded in routine reports. These reports document lesion appearance, size and location but summarise entire procedures rather than caption individual frames, leaving clinical findings only weakly linked to the corresponding images. Here we develop EndoCLIP, a colonoscopy vision-language foundation model trained on 125,756 lesion-level image-text pairs progressively recovered from 280,476 routine colonoscopy records. Across lesion-level image-text retrieval, structured report generation and six multi-centre clinical classification tasks, EndoCLIP outperforms general-purpose and biomedical vision-language encoders in both zero-shot and linear-probe settings. On benign-versus-malignant classification, its linear probe approaches the performance of expert readers in a blinded study involving 12 endoscopists. These results suggest that recovering finding-to-frame correspondence can transform routine documentation into scalable supervision, enabling clinical targets to be specified in language rather than separately annotated for each task.

cs.AI

KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback

In large-scale text analysis tasks, pre-trained language models are often used to embed text corpora for downstream analysis. However, such models may struggle to capture domain-specific semantics and adapting them typically requires large amounts of labeled data and technical expertise to implement training pipelines. Recent approaches have demonstrated how visual interactions in document projections can capture human feedback as training signals for model tuning. However, these methods operate on document-level feedback, which requires users to open and assess individual documents in order to provide effective feedback. In this paper, we propose KeySI, an interaction framework that enables feature-level feedback through keyword-based concept specification. Users specify feedback by organizing extracted keywords into groups representing concepts, which KeySI translates into document-level supervision for subsequent tuning. By operating on keywords as the primary interaction medium, KeySI reduces the need for manual document inspection and labeling and lowers the barrier to adapting embedding models. We present a prototype implementation that, given a corpus, curates representative keywords, visualizes keywords and document embeddings via dimensionality reduction, allows interactive specification of keyword groups, and supports iterative refinement through system feedback. We evaluate KeySI through a user study, usage scenarios, and quantitative experiments demonstrating its effectiveness in capturing user intent and improving embedding alignment.

cs.AI

Learning to Reconstruct Wigner Functions in Phase Space

Wigner function learning is a central tool for characterizing continuous variable quantum systems. A fundamental challenge in this setting is to infer a continuous phase-space function from sparse pointwise measurement data, a task that becomes increasingly demanding as the effective dimension enlarges. Here, we develop a general machine learning framework to reconstruct Wigner functions directly as continuous functions from sparse phase-space data. For states with sparse Fock-space or coherent-state representations, such as binomial code states and cat states, we devise provably efficient regression models whose measurement complexity scales only logarithmically with the effective Hilbert-space dimension. For more general states, such as the Gottesman-Kitaev-Preskill (GKP) states, we design a deep learning model that reconstructs the Wigner function from sparse measurements and generalizes to arbitrary phase-space resolution. We demonstrate the broad applicability of our framework on both simulated data and experimental data from a circuit quantum electrodynamic (circuit-QED) system. Interestingly, on experimental data, we find that our model reconstructs Wigner functions of GKP code states across multiple rounds of quantum error correction and identifies the dominant error process using significantly fewer measurements than conventional estimation techniques.

quant-ph

Debugging as Evidence-Driven Reasoning: Visualization Opportunities in Data-Intensive Programming

Visualization has been recognized as a valuable means of supporting debugging by externalizing runtime behavior that would otherwise remain hidden or scattered. However, most visual debugging research has focused on traditional software development settings, leaving the distinct challenges of data-intensive workflows largely uncharacterized. To build visual debugging support for these settings, we first need to characterize how practitioners debug in these settings and translate their challenges into concrete visualization opportunities. To this end, we conducted semi-structured interviews with nine participants from diverse data-intensive domains and analyzed the data using thematic analysis. Our analysis reveals three cross-cutting challenge: assembling fragmented evidence, detecting expected-observed discrepancies, and tracing state evolution across workflow components. We distill these challenges into three concrete requirements that current debuggers support only partially but that visualization is well suited to address: cross-artifact evidence alignment, expectation-grounded comparison, and traceable state evolution. Together, these requirements begin to characterize a design space for future visual debugging research in data-intensive programming.

cs.HC

Phase-slip residual-order spin state in FeSe

In unconventional superconductors, the microscopic form of magnetic correlations is crucial for identifying the origin of spin fluctuations and the associated pairing interaction. FeSe superconducts without chemical doping and shows no static long-range magnetic order, yet inelastic neutron scattering reveals a strong stripe response, finite linewidths, and reproducible Neel-side spectral weight. Here we propose a phase-slip residual-order spin state (ROSS). Stripe, Neel, pair-checkerboard, and staggered trimer antiferromagnetic states can be unified as symmetric phase-slip derivatives of a stripe background, while more general asymmetric phase slips form lower-energy configurations and reconstruct the spin structure factor S(q) within a finite coherence length. The ROSS therefore reconciles the absence of static magnetic order with strong spin excitations, provides a microscopic picture for the origin of spin fluctuations in FeSe, and establishes a magnetic basis for understanding pairing in unconventional superconducting systems with similar magnetic fingerprints.

cond-mat.supr-con

HiFAST: An HI data calibration and imaging pipeline for the FAST IV: The stray-radiation correction

Stray radiation is a considerable challenge for radio telescopes, requiring careful assessment due to its effects. This is crucial when the strong background flux from side lobes significantly affects the total flux, especially for extended sources. In this study, we introduced the beam pattern of the L-band receiver on the Five-hundred-meter Aperture Spherical Telescope (FAST), covering various frequencies based on recent observations. We discovered that the main beam efficiency of all beams exceeds 90\% throughout the L band frequencies, with efficiency decreasing slowly as frequency increases. Subsequently, we developed a module to mitigate stray radiation effects, incorporating it into FAST's standard \HI data reduction process, referred to as \texttt{HiFAST}. Our analysis shows that side lobe flux's influence, particularly for extended sources with significant surface density gradients, necessitates detailed evaluation. Corrections for the extended M33 galaxy can reach up to 20\%. Moreover, the pattern data presented here is vital for studying HI intensity maps at high redshift. The module, along with HiFAST and beam pattern data across 15 frequency bins, can be accessed at \textrm{https://hifast.readthedocs.io}. The datasets of beam pattern presented in this paper, are openly available at \textrm{https://doi.org/10.57760/sciencedb.j00113.00266} (https://www.scidb.cn/s/bqQRNv).

astro-ph.IM

VIP: Visual-guided Prompt Evolution for Efficient Dense Vision-Language Inference

Pursuing training-free open-vocabulary semantic segmentation in an efficient and generalizable manner remains challenging due to the deep-seated spatial bias in CLIP. To overcome the limitations of existing solutions, this work moves beyond the CLIP-based paradigm and harnesses the recent spatially-aware dino$.$txt framework to facilitate more efficient and high-quality dense prediction. While dino$.$txt exhibits robust spatial awareness, we find that the semantic ambiguity of text queries gives rise to severe mismatch within its dense cross-modal interactions. To address this, we introduce Visual-guided Prompt evolution (VIP) to rectify the semantic expressiveness of text queries in dino$.$txt, unleashing its potential for fine-grained object perception. Towards this end, VIP integrates alias expansion with a visual-guided distillation mechanism to mine valuable semantic cues, which are robustly aggregated in a saliency-aware manner to yield a high-fidelity prediction. Extensive evaluations demonstrate that VIP: 1. surpasses the top-leading methods by 1.4%-8.4% average mIoU, 2. generalizes well to diverse challenging domains, and 3. requires marginal inference time and memory overhead.

cs.CV

Pretraining Induces a Reusable Spectral Basis for Downstream Task Adaptation

Finetuning pretrained models occurs in a low-dimensional subspace of the full parameter space. Prior work has focused on characterizing this optimization subspace, but largely ignored the complementary question: why do certain directions remain unexplored during finetuning? Are these stable directions irrelevant to downstream tasks, or do they already encode task-relevant structure that requires no further adjustment? Answering this question is central to understanding how pretrained knowledge transfers. Through systematic spectral analysis across vision and language models, we show that the leading singular vectors of pretrained weight matrices remain highly stable under finetuning and are shared across unrelated downstream tasks, revealing that pretraining establishes a reusable spectral coordinate system. Models pretrained on larger datasets exhibit greater spectral stability under distribution shift or task change, directly linking pretraining scale to geometric transferability. Motivated by these findings, we propose a parameter-efficient method that freezes pretrained singular vectors and optimizes only leading spectral coefficients, achieving competitive performance on GLUE with 0.2% trainable parameters. Our results reveal that the stable directions encode transferable structure rather than irrelevant noise: successful pretraining discovers spectral bases that downstream tasks inherit and operate within.

cs.LG

Bridging Textual Profiles and Latent User Embeddings for Personalization

Personalized systems rely on user representations to connect behavioral history with downstream recommendation applications. Existing methods typically employ either supervised latent user embeddings, which are effective for retrieval but difficult to interpret, or textual user profiles, which are interpretable but challenging to optimize for downstream utility due to lack of direct supervision. To bridge this gap, we present BLUE, a reinforcement learning framework that unifies these two forms of user representation by aligning language-based user profiles with embedding-based recommendation objectives. Given a user interaction history, BLUE leverages a profiler Large Language Model (LLM) to generate textual profiles, while an embedding model provides reward signals. This encourages the resulting textual representations to move closer to positive items and farther from negative ones in the embedding space. We further introduce a text-space supervision signal based on next-item prediction, ensuring the learned profiles remain both semantically meaningful and highly effective for downstream retrieval. Experiments on Amazon Reviews 2023 and Google Local Reviews in zero-shot sequential recommendation settings demonstrate that BLUE consistently outperforms strong baselines under both frozen and trainable embedding conditions. Notably, BLUE achieves clear gains in cross-domain transfer, highlighting the strong generalization ability of the learned user profiles. Furthermore, these generated profiles provide superior personalized context for question answering compared to raw user histories or alternative profile optimization methods. Overall, these results show that BLUE provides an effective way to unify interpretable textual profiling with discriminative latent embeddings for personalization.

cs.IR

Video-based Heart Rate Estimation with Angle-guided ROI Optimization and Graph Signal Denoising

Remote photoplethysmography (rPPG) enables non-contact heart rate measurement from facial videos, but its performance is significantly degraded by facial motions such as speaking and head shaking. To address this issue, we propose two plug-and-play modules. The Angle-guided ROI Adaptive Optimization module quantifies ROI-Camera angles to refine motion-affected signals and capture global motion, while the Multi-region Joint Graph Signal Denoising module jointly models intra- and inter-regional ROI signals using graph signal processing to suppress motion artifacts. The modules are compatible with reflection model-based rPPG methods and validated on three public datasets. Results show that jointly use markedly reduces MAE, with an average decrease of 20.38\% over the baseline, while ablation studies confirm the effectiveness of each module. The work demonstrates the potential of angle-guided optimization and graph-based denoising to enhance rPPG performance in motion scenarios.

cs.CV