Search arXivSearch

arXiv subjects

Yu Xie

Publications and source records attributed to Yu Xie.

At least 19 recordsLinked to original sources

Benchmarking MLLMs via Cognitive Expected Scene Graph for Safety-Critical Visual Negation Understanding

True machine intelligence requires transcending passive pixel registration to master top-down functional reasoning over absent information via visual negation understanding. However, unconstrained visual negation paradigms remain overly open-ended, and pervasive affirmation bias causes both existing Multi-Modal Large Language Models (MLLMs) and evaluation metrics to fail under negative semantics. To solve these intertwined challenges systematically, we first anchor the boundaries of negation reasoning within specific cognitive goals. Specifically, by focusing on safety as a highly pragmatic and critical cognitive dimension, we define the task of \textbf{S}cene \textbf{N}egation \textbf{U}nderstanding under \textbf{S}afety Cognition (\textbf{SNUS}). Under this framework, we construct a high-fidelity negative caption dataset mapping dense assertions of localized hazards. Concurrently, we propose the Cognitive Expected Scene Graph (CESG) Score, a structure-grounded, polarity-aware evaluation metric. Extensive experiments demonstrate that while current models struggle on the task, traditional metrics completely collapse under semantic reversals. Conversely, our framework delivers a solid benchmark for SNUS, providing a rigorous foundation to advance risk-aware situational comprehension and counterfactual cognition.

cs.CV

Absence is Presence: Understanding Visual Scene Negative Events Under Safety Cognitive Constraint

Traditional scene understanding focuses on affirmative information objectively present in images. However, in safety-critical domains, comprehending key information that should exist but is actually absent is vital for risk mitigation. To bridge this gap, we focus on visual scene negative captioning with safety as the cognitive constraint. The core challenge is to convert physical absence into semantic negative events. Existing vision-language models (VLMs) struggle with this process because affirmation bias suppresses negative reasoning, while limited mental filling capability and representation bias further hinder the inference of absent information. To address these challenges, we propose a negative captioning framework based on counterfactual reconstruction and contrastive decoding (CRCD). Inspired by human cognition, CRCD reformulates the task as counterfactual latent change captioning to bypass affirmation bias. It contrasts a synthesized safe expectation with reality to identify semantic omissions. To address limited mental filling, we design a dual-branch counterfactual reconstruction architecture. The amodal completion branch restores defective objects, while the functional association branch infers completely absent safety objects. Concurrently, a multi-condition representation learning mechanism is integrated to mitigate representation bias by projecting universal features onto predefined safety criteria subspaces, thereby capturing information across more dimensions. By decoding feature-level semantic residuals between the reconstructed scene prototype and raw input, CRCD bounds the non-existence search space and activates the decoder's negative logic. Extensive experiments validate the effectiveness of CRCD, establishing a high-performance baseline for this pioneering task.

cs.CV

Measuring Human Contribution in AI-Assisted Content Generation

With the growing prevalence of generative artificial intelligence (AI), an increasing amount of content is no longer exclusively generated by humans but by generative AI models with human guidance. This shift presents notable challenges for the delineation of originality due to the varying degrees of human contribution in AI-assisted works. This study raises the research question of measuring human contribution in AI-assisted content generation and introduces a framework to address this question that is grounded in information theory. By calculating mutual information between human input and AI-assisted output relative to self-information of AI-assisted output, we quantify the proportional information contribution of humans in content generation. Our experimental results demonstrate that the proposed measure effectively discriminates between varying degrees of human contribution across multiple creative domains. We hope that this work lays a foundation for measuring human contributions in AI-assisted content generation in the era of generative AI.

cs.CY

GSO-Net: Visual State Machines for Hazardous Freight Transfer Compliance at Petrochemical Logistics Nodes

Hazardous-freight operations at petrochemical logistics nodes are safety-critical for intelligent transportation systems, yet existing vision benchmarks rarely address procedural compliance under realistic deployment constraints. In large infrastructure networks, cameras often operate under sparse round-robin polling, so transfer status must be inferred from incomplete observations and localized evidence. We present GSO-Net, a large-scale benchmark for visual understanding of standard operating procedures (SOPs) in petrochemical unloading scenarios. To our knowledge, GSO-Net is the first public benchmark dataset dedicated to visual SOP understanding in petrochemical hazardous-freight transfer scenarios. It contains over 50,000 independently sampled frames from 64 real expressway petrochemical logistics nodes and adopts an SOP-derived hierarchy linking 9 macroscopic procedural steps with 15 microscopic operational states. Two tasks are defined: joint detection of microscopic states and macroscopic steps as the core benchmark, and frame-level step classification as a diagnostic reference. Experiments with lightweight, transformer-based, open-vocabulary, and holistic models reveal a clear gap between object perception and transfer-stage understanding. Current models remain weak on contact-level state grounding, transient step recognition, and stage consistency, especially under sparse polling, tiny critical targets, and long-tailed operational evidence. GSO-Net provides a practical benchmark for fine-grained state perception and vision-based safety monitoring in hazardous freight transportation. The dataset is publicly available at https://github.com/yuxieHarrison/GSO-Net

cs.CV

Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation

However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unified three-component benchmark for diagnosing and mitigating stylistic bias in LLM-based idea evaluation: (i) First, SciStyleStage, a three-stage evaluation environment that applies controlled stylistic perturbations to fixed scientific content across three settings no context, fixed-domain context, and open-domain retrieval context, covering 600 scientific ideas and 15 style variants, with 9,000 evaluation instances per setting; (ii) Second, SciStyleMetrics, a set of quantitative measures, including Style Bias Index (SBI), Substance Recognition Rate (SRR), and Adversarial Win Rate (AWR), to characterize how stylistic variation affects scoring stability, substance discrimination, and ranking robustness; (iii) Third, SciStyleExtractor, a plug-and-play evaluation module that separates presentation style from scientific content by predicting style type and deviation before style-conditioned evaluation, enabling us to assess whether style awareness reduces stylistic bias. Experiments on SciStyleBench show that direct LLM judges remain sensitive to writing style and struggle to distinguish scientific substance. In contrast, SciStyleExtractor reduces SBI from 0.566 to 0.501 while increasing SRR and AWR from 0.504 and 0.554 to 0.759 and 0.899, respectively. These results suggest that robust idea evaluation requires invariance to stylistic variation without sacrificing sensitivity to scientific substance. Overall, SciStyleBench provides a systematic framework for identifying, quantifying, and mitigating stylistic bias in scientific idea evaluation.

cs.CL

A Versatile Foundation Model for AI-enabled Mammogram Interpretation

Breast cancer is the most commonly diagnosed cancer and the leading cause of cancer-related mortality in women globally. Mammography is essential for the early detection and diagnosis of breast lesions. Despite recent progress in foundation models (FMs) for mammogram analysis, their clinical translation remains constrained by several fundamental limitations, including insufficient diversity in training data, limited model generalizability, and a lack of comprehensive evaluation across clinically relevant tasks. Here, we introduce VersaMammo, a versatile foundation model for mammograms, designed to overcome these limitations. We curated the largest multi-institutional mammogram dataset to date, comprising 706,239 images from 21 sources. To improve generalization, we propose a two-stage pre-training strategy to develop VersaMammo, a mammogram foundation model. First, a teacher model is trained via self-supervised learning to extract transferable features from unlabeled mammograms. Then, supervised learning combined with knowledge distillation transfers both features and clinical knowledge into VersaMammo. To ensure a comprehensive evaluation, we established a benchmark comprising 92 specific tasks, including 68 internal tasks and 24 external validation tasks, spanning 5 major clinical task categories: lesion detection, segmentation, classification, image retrieval, and visual question answering. VersaMammo achieves state-of-the-art performance, ranking first in 50 out of 68 specific internal tasks and 20 out of 24 external validation tasks, with average ranks of 1.5 and 1.2, respectively. These results demonstrate its superior generalization and clinical utility, offering a substantial advancement toward reliable and scalable breast cancer screening and diagnosis.

cs.CV

On regression with estimated covariates and conditional effects given the propensity score

Motivated by the study of heterogeneous returns to education in Brand & Xie 2010, which considers how the effect of completing college on earnings varies with the (unknown) probability of completing college, we analyze the problem of estimating a nonparametric regression function when certain covariates are estimated in a first step. Plug-in estimators that treat the estimated covariates as known generally suffer from first-stage estimation error. To mitigate this issue, we analyze two debiasing approaches within a framework that is agnostic to the choice of the first-stage estimation method and relies on either local-smoothing or sieve-based methods for the second-stage regression. In particular, we consider: (i) influence function-based estimators of pathwise differentiable parameters that approximate the target estimand, and (ii) a variant of plug-in estimators that directly aims to correct their bias. For each method, we upper bound the estimation error and characterize conditions under which oracle rates can be approached, highlighting the possible gains in terms of convergence rates relative to the plug-ins. Simulation studies illustrate the finite-sample behavior of the methods. We apply our methodology to data from the National Longitudinal Survey of Youth 1997 and find evidence that completing college yields the largest reductions in unemployment for individuals least likely to do so, consistent with earlier findings in the literature (Brand & Xie 2010; Brand 2023).

stat.ME

MemeBench: What LVLMs Miss When Interpreting Culture-Dependent Memes

Large vision-language models have improved at describing visual content, but accurate descriptions do not ensure interpretation when meaning depends on knowledge beyond the pixels. Memes expose this gap because they rely on cultural entities, background knowledge, and community conventions. Most meme benchmarks reduce interpretation to labels or holistic scores, obscuring where an explanation breaks down. We introduce MemeBench, a diagnostic benchmark of 1,253 Chinese and English memes with human-written references and quality-controlled VIKR annotations, centered on anime, comics, games, and adjacent online subcultures. Its VIKR schema decomposes explanations into Visual clues, Identity links, Knowledge units, and Reasoning mechanisms. Across 26 LVLMs, every model covers visible content more reliably than the knowledge needed to interpret it, and even the strongest retains a 22.6% Visual-Knowledge gap. To test whether this diagnosis can guide improvement, we introduce KAR, an entity-guided retrieval baseline built on CultureBase. Across four controlled models, KAR raises VIKR Success by 3.6-7.4% and, compared with generic retrieval, repairs more answers and breaks fewer. Yet both retrieval conditions improve Identity and Knowledge while reducing Visual coverage in every comparison. MemeBench reveals whether an interpretation succeeds, what is missing, and whether targeted evidence fills the diagnosed gap.

cs.AI

Learning behavior accounts for background-related advantage in AI-assisted education

Generative AI has been found, and will likely be found increasingly, useful in education. However, existing AI-for-education studies provide inconsistent evidence on its average effects. More broadly, research on prior educational technologies shows that average effects often mask substantial heterogeneity across student populations. Motivated by this evidence, this study examines heterogeneity in students' learning behavior with AI, which students benefit from AI assistance, and how learner profiles and learning behavior shape these patterns. To this end, we recruited 318 university students to participate in structured learning experiments lasting up to 125 minutes. Our findings indicate that students' learning behavior is strongly associated with learning outcomes, with behaviors characterized by proactive and critical engagement, rather than limited engagement, associated with significantly better performance. These behavioral differences are related to learner profiles, with students from higher-ranking universities and those with greater prior knowledge tending to benefit more, consistent with their greater likelihood of adopting proactive interaction strategies. Accounting for learning behavior substantially weakens or eliminates the associations between learner profiles and learning outcomes, suggesting that how students use AI is a key pathway through which background differences are linked to learning gains. Overall, this work provides a deeper understanding of AI assistance in education by showing how differences in learner profiles and learning behavior shape who benefits from AI-supported learning. These insights can help educators and students better navigate and integrate AI into educational practices.

cs.HC

Enhanced Diffusion Sampling: Efficient Rare Event Sampling and Free Energy Calculation with Diffusion Models

The rare-event sampling problem has long been the central limiting factor in molecular dynamics (MD), especially in biomolecular simulation. Recently, diffusion models such as BioEmu have emerged as powerful equilibrium samplers that generate independent samples from complex molecular distributions, eliminating the cost of sampling rare transition events. However, a sampling problem remains when computing observables that rely on states which are rare in equilibrium, for example folding free energies. Here, we introduce enhanced diffusion sampling, enabling efficient exploration of rare-event regions while preserving unbiased thermodynamic estimators. The key idea is to perform quantitatively accurate steering protocols to generate biased ensembles and subsequently recover equilibrium statistics via exact reweighting. We instantiate our framework in three algorithms: UmbrellaDiff (umbrella sampling with diffusion models), MetaDiff (a batchwise analogue for metadynamics), and $Δ$G-Diff (free-energy differences via tilted ensembles). Across toy systems, protein folding landscapes and folding free energies, our methods achieve fast, accurate, and scalable estimation of equilibrium properties within GPU-minutes to hours per system-closing the rare-event sampling gap that remained after the advent of diffusion-model equilibrium samplers.

stat.ML

Topological phase transition driven by in-plane spin rotation

The intrinsic coupling between magnetism and nontrivial band topology in magnetic topological insulators makes external magnetic fields a powerful tool for manipulating topological states. However, conventional magnetic control mechanisms, such as driving magnetic phase transitions or fully reversing magnetization, typically demand large magnetic fields and lack continuous tunability. Here, we establish a symmetry framework for the reversible switching of topological states via continuous in-plane spin rotation, governed by magnetic point group constraints on the Berry curvature distribution. Using a two-dimensional kagome ferromagnetic Chern insulator as a prototype, we demonstrate that a 60°in-plane magnetization rotation reverses the sign of the Chern number, transitioning through a topologically trivial state. Crucially, micromagnetic simulations confirm that this spin-reorientation-driven switching operates under exceptionally small magnetic fields and on ultrafast timescales. This work provides a highly efficient, low-energy paradigm for the manipulation of topological states.

cond-mat.mtrl-sci

Manifold partitioning induced sequential optical reasoning and decision framework for photonic computing

Real-world data are intrinsically embedded in highly entangled manifolds, making the extraction of separable representations a central challenge for artificial intelligent (AI) systems. While optical neural networks (ONNs) offer ultrafast and energy-efficient data processing, their capacity is constrained by limited physical depth. Here, we introduce a sequential optical reasoning and decision (SORD) framework, an architecture that performs time-sequenced hierarchical inference by decomposing global tasks into coarse-to-fine steps via geometry-guided data partitioning. At each step, SORD executes small reasoning via dynamic operator selection, effectively reducing the overall task complexity without scaling up physical architecture. Experimentally, SORD enables a single-layer diffractive ONN to achieve otherwise intractable 100-class optical fiber speckle classification with 94% accuracy and a system energy efficiency of 23.3 TOPS/W. This high-fidelity recognition is further examined in a human-machine interface, featuring real-time interactive all-optical sensing. Overall, our work establishes a scalable and hardware-efficient approach to expanding the effective expressivity of compact photonic AI systems, and may advance their deployment in applications requiring real-time sensing, inference, and control.

physics.optics

Cross-Modal Clinical Knowledge Integration for Mammography Report Generation

Breast cancer is a major global health concern, and mammography screening plays a central role in early detection. The large volume of screening examinations creates a substantial workload for radiologists, making accurate and consistent report generation a critical clinical challenge. Existing automated mammography report generation methods primarily focus on direct visual-to-text mapping, while overlooking the structured clinical reasoning process followed by radiologists in real-world practice. To address this limitation, we propose MammoRG, a mammography report generation framework that explicitly simulates the clinical reporting workflow by following the BI-RADS guideline and incorporating prior clinical knowledge to produce diagnostic reports. Specifically, MammoRG adopts a two-stage training framework. In the first stage, the model learns to integrate clinically relevant prior knowledge from a patient's four-view mammograms through classification-based supervision. In the second stage, a terminology-aware supervised fine-tuning strategy is introduced to model mammography-specific clinical terms as atomic semantic units, enabling the generation of high-quality reports with improved clinical consistency. To facilitate clinical efficacy evaluation of generated reports, we further develop MammoRGTool, a dedicated mammography report parsing tool that extracts structured clinical information from free-text reports. Extensive experiments demonstrate that MammoRG consistently outperforms existing methods across multiple clinical efficacy metrics, particularly in diagnosis-related BI-RADS F1, where it surpasses the second-best model by 2.73%, 2.04%, 1.90%, and 3.27% on the internal, external 1, external 2, and VinDr-Mammo datasets, respectively.

cs.CV

Domain-Wall Mediated Polarization Switching in Ferroelectric AlScN: Strain Relief and Field-Dependent Dynamics

While scandium-doped aluminum nitride (AlScN) exhibits robust ferroelectricity and excellent thermal stability, its utility is limited by an exceptionally high coercive field ($E_c$) for polarization switching. Unraveling the atomistic switching dynamics is therefore critical for tailoring $E_c$. Here, we combine density functional theory and machine-learning molecular dynamics to elucidate the polarization switching mechanisms in AlScN over various Sc concentrations and applied electric fields. We find that excessive lattice strain strictly prohibits collective polarization switching, but the pre-existing domain walls relieve strain and lead to a distinct switching dynamics -- dictating a field-dependent switching mechanism. At low electric fields, switching occurs via gradual domain-wall propagation consistent with the Kolmogorov-Avrami-Ishibashi model. In contrast, high fields stimulate additional nucleation, driving a rapid, homogeneous reversal process described by the simultaneous non-linear nucleation and growth model. These findings highlight the critical role of domain-wall dynamics and suggest domain engineering as a viable strategy to tailor coercive fields in AlScN and related ferroelectrics.

cond-mat.mtrl-sci

Hierarchical Multi-Fidelity Learning for Predicting Three-Dimensional Flame Wrinkling and Turbulent Burning Velocity

High-fidelity experimental characterization of turbulent premixed flames remains limited by the cost and complexity of advanced diagnostics, particularly under elevated pressures and intense turbulence where measurements of coupled flame morphology and burning dynamics are sparse. Here, we develop a hierarchical multi-fidelity neural network framework (MuFiNNs) to address this challenge by integrating sparse high-fidelity experimental data with structured low-fidelity representations encoding dominant physical trends. The framework combines hierarchical low-fidelity construction with nonlinear multi-fidelity correction to learn coupled geometric and reactive flame behavior while recovering discrepancies that simplified models alone cannot capture. The methodology is applied to expanding turbulent premixed flames to predict three-dimensional flame wrinkling dynamics and turbulent mass burning velocity across varying fuels, pressures, and turbulence intensities. Using experimentally informed low-fidelity trend models with sparse high-fidelity measurements, MuFiNNs accurately reconstruct observed flame behavior, enable interpolation across unseen operating conditions, and demonstrate robust extrapolation beyond the training domain. Importantly, the framework remains effective in noisy, weakly structured, or experimentally inaccessible regimes where conventional data-driven approaches often fail. These results show that hierarchical multi-fidelity learning provides a scalable and physically grounded strategy for predictive combustion modeling in data-limited regimes. More broadly, this work establishes multi-fidelity scientific machine learning as a practical framework for extracting physically meaningful predictive models from sparse experiments, particularly for instability-dominated and turbulence-sensitive reactive flows where high-fidelity data acquisition is demanding.

cs.LG

Effective phonon models based on symmetry-adapted multipole basis -- Hidden chiral phonon angular momentum splitting in ferroaxial systems

We propose a symmetry-based framework for constructing effective harmonic phonon models using a symmetry-adapted multipole basis. By decomposing the force-constant matrix into bond-centered electric multipoles, we identify the minimal microscopic ingredients responsible for phonon angular-momentum splitting. Applying this framework to a minimal zigzag-chain model, we show that ferroaxial order gives rise to a hidden sublattice-resolved chiral phonons, while an additional polar contribution leads to finite global chirality. Our results provide a unified symmetry-based description of hidden and emergent phonon phenomena and suggest a route to control phonon properties via electronic orderings and external fields.

cond-mat.mtrl-sci

SafeCtrl: Region-Aware Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress

The widespread deployment of text-to-image diffusion models is significantly challenged by the generation of visually harmful content, such as sexually explicit content, violence, and horror imagery. Common safety interventions, ranging from input filtering to model concept erasure, often suffer from two critical limitations: (1) a severe trade-off between safety and context preservation, where removing unsafe concepts degrades the fidelity of the safe content, and (2) vulnerability to adversarial attacks, where safety mechanisms are easily bypassed. To address these challenges, we propose SafeCtrl, a Region-Aware safety control framework operating on a Detect-Then-Suppress paradigm. Unlike global safety interventions, SafeCtrl first employs an attention-guided Detect module to precisely localize specific risk regions. Subsequently, a localized Suppress module, optimized via image-level Direct Preference Optimization (DPO), neutralizes harmful semantics only within the detected areas, effectively transforming unsafe objects into safe alternatives while leaving the surrounding context intact. Extensive experiments across multiple risk categories demonstrate that SafeCtrl achieves a superior trade-off between safety and fidelity compared to state-of-the-art methods. Crucially, our approach exhibits improved resilience against adversarial prompt attacks, offering a precise and robust solution for responsible generation.

cs.CV

From Passive Feeds to Guided Discovery: AI-Initiated Interaction for Vague Intent in Content Exploration

Recommendation feeds work well when people are simply browsing, and search works well when they can formulate a query. Between these two cases is a common but poorly supported state: users feel that their feed has become repetitive, yet cannot clearly specify what they want instead. We refer to this state as vague intent. We present Red-Rec, an AI-supported exploration interface for this middle ground. After a period of browsing, the system summarizes patterns in the current feed (e.g., dominant content categories and possible latent interests), offers clickable exploration options, asks at most one follow-up question, and then gradually blends new content into the feed. The design is motivated by a formative study which found that users often recognize feed staleness but struggle to articulate alternatives, suggesting the need for proactive and low-effort interaction.We evaluated Red-Rec in a mixed-design lab study against three comparison conditions: a passive feed, search, and a user-initiated chat interface. Compared with user-initiated chat, Red-Rec led to broader exploration, higher serendipity ratings, and lower interaction effort. Participants in the AI-initiated condition typed very little , relying mainly on option selection, whereas participants in the user-initiated chat condition typed substantially more . We discuss how proactive, option-based AI support can help users move beyond repetitive feeds without undermining their sense of control, and we outline design implications for recommendation interfaces that support open-ended exploration.

cs.HC