Search arXiv⌕ Search

arXiv subjects

Zhuo Wang

Publications and source records attributed to Zhuo Wang.

At least 19 recordsLinked to original sources

Signals of AI Hallucination: Designing Hallucination-Aware Cues for Embodied Conversational Agents in VR

LLM-powered conversational agents (CAs) often present uncertainty and provenance cues alongside their responses to help users assess response reliability and identify potential hallucinations. In immersive environments such as Virtual Reality (VR), CAs often take the form of speech-based embodied conversational agents (ECAs), where uncertainty and provenance cues cannot rely on persistent inline text and may be missed or disrupt comprehension when delivered through speech. We conducted a within-subjects study (N = 24) to compare three designs for presenting the hallucination-awareness information (uncertainty and provenance) in ECAs in VR against a no-cue baseline: embodied cues using gestures and posture, icon cues using visual indicators, and text cues using color-coded text with inline citations. We evaluated how these designs affect users' ability to identify hallucination-related information, trust in the ECA, and interaction experience (immersion and task load). Our results show that all three designs support users in identifying hallucinations. Embodied cues were associated with higher trust and immersion, text cues offered clearer interpretability, and icon cues preserved relatively good interpretability while causing less disruption to immersion compared with embodied cues and text cues. This work contributes to the VR and AI research community by comparing different designs of hallucination cues in immersive ECA settings and examining how they affect users' ability and experiences to identify hallucinations. It also offers practical insights and design implications for developing future hallucination-awareness interfaces for ECA.

cs.HC↗

Feature-Adaptive Fusion in Hybrid Quantum-Classical Neural Networks for Robust Biomedical Image Classification

Hybrid quantum-classical neural networks provide a promising approach for incorporating quantum circuits into machine learning in the noisy intermediate-scale quantum regime. However, existing hybrid models often rely on fixed or globally shared fusion strategies, which may limit their ability to exploit complementary information carried by quantum branches, especially under distribution shifts. In this work, we propose a Feature-Adaptive Fusion Hybrid Quantum-Classical Neural Network (FAF-HQNN) for biomedical image classification. The model combines a classical deep feature encoder with a variational quantum circuit (VQC) and introduces a feature-adaptive fusion mechanism to dynamically weight classical and quantum predictions. We evaluate FAF-HQNN on two MedMNIST benchmarks, PathMNIST and BloodMNIST, under clean and corrupted test conditions. FAF-HQNN achieves the strongest overall performance on clean data among the compared methods and shows improved robustness under Gaussian, salt-and-pepper, and Poisson corruptions. Further analysis of circuit layout, measurement basis, and depth shows that even shallow variational quantum circuits can provide useful complementary information. These results demonstrate that feature-adaptive fusion is an effective strategy for improving accuracy and robustness in hybrid quantum-classical models for biomedical image classification.

quant-ph↗

$^{139}$La nuclear quadrupole resonance studies of pressurized La$_4$Ni$_3$O$_{10}$

Density-wave (DW) orders are considered as competing orders to unconventional superconductivity and are commonly seen in a variety of superconductors including but not limited to the recently discovered Ruddlesden-Popper-phase nickelates. By utilizing $^{139}$La nuclear quadrupole resonance, we systematically investigate into the nature of DW orders and their evolution under pressure in La$_4$Ni$_3$O$_{10}$. Spin and charge DW orders are found to be intertwined in this material, which is in stark contrast to those in La$_3$Ni$_2$O$_7$. Short-range DW orders are observed near 150 K, well above the development of long-range DW orders at around 139 K. Upon applying a hydrostatic pressure of 2.3 GPa, the transition temperatures of the short-range and long-range orders decrease at rates of 1 K/GPa and 10 K/GPa, respectively. Our results thus affirm that both spin density wave and charge density wave as competing orders with the superconducting state in La$_4$Ni$_3$O$_{10}$, and provide new insights into the interplay between DW orders and unconventional superconductivity.

cond-mat.supr-con↗

Anomalous magnetocaloric effects in the quasi-one-dimensional antiferromagnet BaCo$_2$V$_2$O$_8$

We investigate the transverse-field thermodynamics of the quasi-one-dimensional Ising-like antiferromagnet BaCo$_2$V$_2$O$_8$, whose tilted screw-chain geometry and anisotropic Landé $g$ tensor generate spatially modulated Zeeman couplings. Angle-resolved magnetocaloric-effect (MCE) measurements reveal a high-field temperature minimum near the transverse-field Ising critical field for $H\parallel[110]$ that persists and shifts only weakly upon field rotation. Tensor-network calculations show that the rotation-induced staggered transverse field rapidly lowers the Ising critical field and that the magnetic Grüneisen ratio changes sign near the high-field temperature minimum, consistent with experiment. Our results establish that a dominant MCE response can persist away from the Ising critical region, suggesting a route to magnetic cooling by tailoring anisotropic Zeeman-coupling configurations in quantum magnets.

cond-mat.str-el↗

Multiple superconducting phases and order-parameter evolution in pressurized UTe$_2$

The recently discovered heavy-fermion spin-triplet superconductor candidate UTe$_2$ provides a rich platform for unconventional pairing and topological phenomena. However, limited has been known about its superconducting order parameters and their evolution with control parameters, largely due to the lack of appropriate symmetry-sensitive detections. Here, we report comprehensive point-contact spectroscopy measurements of pressurized UTe$_2$ on the (0~0~1) surface. The observation of Andreev bound states strongly suggests the presence of a $p_z$ component in the superconducting order parameters. Quantitative analysis based on an extended Blonder-Tinkham-Klapwijk model unveils the superconducting order parameters with a finite odd-$k_z$ component (e.g. $B_{2u}$ or $B_{3u}$) for both ambient and pressurized UTe$_2$. Remarkably, the multiple superconducting phases can be distinguished by a single parameter $\langle Δ_{z}\rangle/\langleΔ_{x(y)}\rangle$, the relative weight between the $p_z$-wave and $p_{x(y)}$-wave pairings. These findings place stringent constraints on the pairing symmetry and provide essential spectroscopic signatures for distinguishing pressure-induced multiple superconducting phases in UTe$_2$.

cond-mat.str-el↗

Navigating Through Turbulence: Charting Early Careers in Weather and Climate Science

The field of weather and climate science is at a pivotal moment, defined by simultaneous forces of institutional disruption and unprecedented technological advancements. While a shifting research and employment landscape has created career uncertainty, prompting many scientists to consider or pursue opportunities in the private sector, it has simultaneously spurred an expansion of the ecosystem through the emergence of new computational tools and the growing role of industry innovators and stakeholders. This perspective paper argues that this new, expanded ecosystem presents extraordinary opportunities for students and early-career professionals. We outline the emerging scientific frontiers powered by high-resolution simulations and artificial intelligence, suggest a practical path for navigating a more fluid career landscape, and propose how education and training must evolve. We argue that these changes expand rather than diminish the reader's capacity to do what drew most of us to this field: helping people prepare for what the atmosphere is about to do.

physics.ao-ph↗

Tilted $p$-wave magnet candidate CeNiAsO

The unexpectedly small ordered moments of CeNiAsO, a candidate for correlated $p$-wave magnet, have posed a serious challenge to the precise determination of its magnetic structure, hindering the understanding of its fundamental properties. By leveraging the high sensitivity to local internal fields, our $^{75}$As nuclear quadrupole / magnetic resonance experiments reveal a commensurate antiferromagnetic order with a small out-of-plane moment $m_z\approx0.05$ $μ_{\mathrm{B}}$. This tilted magnetic configuration not only rotates the spin polarization axis away from the crystallographic $\mathbf{c}$-axis, but also enhances the non-relativistic spin splitting. We refer to this rare paradigm as a \textit{tilted $p$-wave magnet}.

cond-mat.str-el↗

QWRF-Net: A Quantum-Wavelet Framework with Rectified Flow for Short-Term Precipitation Nowcasting

Short-term precipitation nowcasting is important for hydrometeorological early warning, especially when intense convective rainfall may trigger urban flooding, flash floods, and other high-impact hazards. A key challenge in warning-oriented nowcasting is that radar precipitation fields contain strongly coupled multi-scale structures, while forecast quality often degrades at later lead times, making it difficult to preserve intense precipitation cores and their spatial organization over the full warning-relevant horizon. To address this problem, we propose QWRF-Net, a quantum-wavelet framework with rectified flow for short-term precipitation nowcasting. The core idea is to improve the conditional representation of precipitation by explicitly decomposing latent features into wavelet sub-bands and then performing differentiated quantum-inspired modulation in the decomposed latent space, before generating future sequences through a rectified-flow-based non-autoregressive decoder. Experiments on the KNMI radar and SEVIR benchmarks under a unified evaluation protocol show that QWRF-Net achieves favorable overall performance, with relatively consistent gains at medium-to-high precipitation thresholds, on an extreme-event subset, and in preserving intense precipitation cores and fine-scale structures. Ablation results further indicate that wavelet-based scale disentanglement, differentiated sub-band modulation, and flow-based generation provide complementary benefits within the proposed framework. Overall, these results suggest that jointly enhancing multi-scale precipitation representation and stable multi-step generation is a promising direction for warning-oriented short-term precipitation nowcasting. The observed improvements may also provide a more useful precipitation basis for downstream hydrological and warning-related applications.

cs.LG↗

Seasonality of hail outbreaks in the United States and its link to weather regimes

Hailstorms in the United States produce immense economic losses and have an occurrence frequency that appears to have a long-term positive trend. Here we contribute an updated examination of such hail activity and the relevant environmental parameters over the period 1990--2024. We define a hail outbreak as any day with $\gt 6$ significant hail reports ($hail\geq 2$'' diameter) and find a positive trend in hail outbreaks since 1990 ($slope=0.8$ days $yr^{-1}$) with the largest increases occurring in May and June and in the Southern Great Plains. Analysis of environmental parameters shows that the long-term increase in hail outbreaks correlates with a long-term increase in the significant hail parameter (SHIP, $cc=0.67$, $p\lt 0.01$). The interannual variability and long-term trend of SHIP are driven both by thermodynamic and kinematic processes, though kinematic processes play a more important role than thermodynamic processes in the interannual variability of hail outbreaks. Using warm-season (April--July) weather regimes (WRs), we find that large-scale circulation modulates the interannual variability of hail outbreak frequency. An empirical model using WR frequency captures the interannual variability in warm-season hail outbreaks reasonably well ($cc=0.38$, $p=0.03$). Our study is the first to relate hail activity to WRs and presents a better understanding of the trend and year-to-year variability of hail outbreaks.

physics.ao-ph↗

WP-MIP: An Artificial Intelligence, Hybrid, and Physically Based Model Intercomparison Project for Weather Prediction

Rapid progress in the field of machine-learning for weather prediction has led to the emergence of algorithms whose forecasting skill can exceed that of traditional physically based models. This development represents an opportunity to improve the quality of forecasting services provided by operational centers, particularly given the speed at which machine-learning based models generate predictions. Despite the clear promise of these systems, questions remain about the ability of the current generation of machine-learning models to generate physically consistent predictions of the full suite of required forecast fields under all conditions. Answering these questions will require careful comparisons between the well-understood physically based models, current state-of-the-art machine-learning models, and the hybrid models that combine elements of these two archetypes. The Weather Prediction Model Intercomparison Project (WP-MIP) is a World Meteorological Organization-supported initiative whose initial goal is to create a centralized database of physically based, machine-learning, and hybrid model forecasts to enable a distributed assessment and evaluation effort. The first instance of WP-MIP focuses on global deterministic predictions using both center-specific and common initializations to facilitate sensitivity studies. Forecasts contributed by institutions across six continents will be used to develop AI-ready verification techniques that highlight the strengths and weaknesses of each class of prediction system, with the goal of establishing best-practice guidance to model developers and national weather centers. The broad engagement of the operational and forecast-evaluation communities in WP-MIP will ensure that the project results are highly relevant to the development and deployment of next-generation weather prediction systems.

physics.ao-ph↗

Possible inverse magnetic melting effect in vdW-like Kondo lattice CeSn$_{0.75}$Sb$_2$

Given the intimate connection between magnetic orders and the interplay among multiple degrees of freedom in heavy-fermion systems, controlling and understanding the associated inverse melting effect is crucial for unveiling novel condensed-matter states and their potential applications. Here, we report the growth of single crystalline quasi-two-dimensional van-der-Waals-like (vdW-like) Kondo lattice CeSn$_{0.75}$Sb$_2$, and its physical properties by a combination of transport / magnetic / thermodynamic measurements. We find that it hosts a fragile antiferromagnetic (AFM) order and a cluster glass (CG) ground state, both of which are highly sensitive to external fields. Upon cooling under low in-plane magnetic fields, the AFM phase evolves into a polarized paramagnetic phase, either directly or indirectly through the intermediate CG phase. This process constitutes a possible inverse magnetic melting effect that restores the broken translational / rotational symmetries. Our work provides a rare paradigm of inverse magnetic melting effect in vdW-like heavy-fermion materials, and enriches the physics in conventional Kondo-lattice models.

cond-mat.str-el↗

Contrastive Spectral Rectification: Test-Time Defense towards Zero-shot Adversarial Robustness of CLIP

Vision-language models (VLMs) such as CLIP have demonstrated remarkable zero-shot generalization, yet remain highly vulnerable to adversarial examples (AEs). While test-time defenses are promising, existing methods fail to provide sufficient robustness against strong attacks and are often hampered by high inference latency and task-specific applicability. To address these limitations, we start by investigating the intrinsic properties of AEs, which reveals that AEs exhibit severe feature inconsistency under progressive frequency attenuation. We further attribute this to the model's inherent spectral bias. Leveraging this insight, we propose an efficient test-time defense named Contrastive Spectral Rectification (CSR). CSR optimizes a rectification perturbation to realign the input with the natural manifold under a spectral-guided contrastive objective, which is applied input-adaptively. Extensive experiments across 16 classification benchmarks demonstrate that CSR outperforms the SOTA by an average of 18.1% against strong APGD with modest inference overhead. Furthermore, CSR exhibits broad applicability across diverse visual tasks. Code is available at https://github.com/Summu77/CSR.

cs.CV↗

Impacts of Histories and Models on LLM Grading: A Study in Advanced Software Engineering Courses

Graduate-level research reading report assessment creates a substantial labor burden for educators. While large language models (LLMs) hold great potential for automating academic grading, their reliability for this specialized task remains understudied, particularly regarding grading consistency, the lack of which represents a primary obstacle to educational fairness. This paper proposes a human-aligned LLM-assisted grading workflow and presents a case study based on 180 student submissions from a graduate advanced software engineering course. We evaluate two mainstream LLMs, Grok and GPT, in terms of grading consistency and alignment with human scores. We find LLMs exhibit distinct levels of intra-model consistency and significant inter-model grading inconsistencies, while simple ensemble approaches cannot improve alignment with human evaluation. Critically, continuous interaction history drives systematic drift in models' grading standards away from human expert scores. Our findings demonstrate LLMs' potential in reducing grading workload for educators in graduate education, while highlighting that indiscriminate LLM grading may introduce systemic unfairness, suggesting that specific operational practices are required to mitigate such disparities.

cs.SE↗

FT-Dojo: Towards Autonomous LLM Fine-Tuning with Language Agents

Fine-tuning large language models for vertical domains remains labor-intensive, requiring practitioners to curate data, configure training, and iteratively diagnose model behavior. Despite growing interest in autonomous machine learning and language agents, end-to-end LLM fine-tuning has not been systematically studied as an interactive agent task. We introduce FT-Dojo, an interactive benchmark environment for autonomous LLM fine-tuning, comprising 13 tasks across 5 domains. Rather than a new collection of static datasets, FT-Dojo standardizes a task interface, shared raw-data repository, sandboxed execution environment, structured feedback protocol, and held-out evaluation procedure. We further develop FT-Agent, a fine-tuning-oriented autonomous framework that uses structured iteration planning, fail-fast validation, and multi-level feedback analysis to refine data and training strategies. Experiments show that FT-Agent provides a strong initial baseline, achieving the best performance on 10 out of 13 tasks, with additional controlled comparisons against frontier agents, open-source planning backbones, and multi-run statistics supporting the main findings. Case studies show that agents can recover from failures through cumulative learning, while still exposing limitations in causal diagnosis and long-horizon planning. The implementation is available at https://github.com/microsoft/rd-agent.

cs.AI↗

Agent^2 RL-Bench: Can LLM Agents Engineer Agentic RL Post-Training?

We introduce Agent2 RL-Bench, a compact diagnostic benchmark for evaluating agentic RL post-training, which tests whether LLM agents can autonomously design, implement, debug, and execute post-training pipelines that improve foundation models. RL post-training increasingly drives model alignment and specialization, yet existing benchmarks are largely static, rewarding supervised fine-tuning or script generation without assessing an agent's ability to close an interactive RL loop. Agent2 RL-Bench provides a unified agent-facing interface: each run starts from an isolated workspace containing a base model, task data, instructions, and a grading API, and agents must iterate within a fixed budget by training models and submitting artifacts for evaluation. The benchmark spans six tasks across three levels, from static rule-based training to judge-based optimization and closed-loop online RL with trajectory collection. Two diagnostic skills, namely runtime recording and post-hoc summarization, enable structured analysis of agent behavior, facilitating smooth and effective iteration of the benchmark's evaluation framework. Across five agent systems and six driver LLMs, agents show intelligent behavior but clear limitations: one RL-oriented run improves ALFWorld from 4.85 to 93.28 via SFT warm-up and GRPO with online rollouts, yet DeepSearchQA remains difficult, most successful routes rely on supervised pipelines, and interactive outcomes show large single-run differences across agent stacks. Overall, Agent2 RL-Bench shows that current agents can sometimes engineer online RL, but stable agent-driven RL post-training remains rare under fixed budgets. It also demonstrates that our benchmark provides a strong and effective evaluation framework for future research in this direction. Code is available at https://github.com/microsoft/RD-Agent/blob/main/rdagent/scenarios/rl/autorl_bench/README.md

cs.AI↗

Remember the Decision, Not the Description: A Rate-Distortion Framework for Agent Memory

Long-horizon language agents must operate under limited runtime memory, yet existing memory mechanisms often organize experience around descriptive criteria such as relevance, salience, or summary quality. For an agent, however, memory is valuable not because it faithfully describes the past, but because it preserves the distinctions between histories that must remain separated under a fixed budget to support good decisions. We cast this as a decision-centric rate-distortion problem, measuring memory quality by the loss in achievable decision quality induced by compression. This yields an exact forgetting boundary for what can be safely forgotten, and a memory-distortion frontier characterizing the optimal tradeoff between memory budget and decision quality. Motivated by this decision-centric view of memory, we propose DeMem, an online memory learner that refines its partition only when data certify that a shared state would induce decision conflict, and prove near-minimax regret guarantees. On both controlled synthetic diagnostics and long-horizon conversational benchmarks, DeMem yields consistent gains under the same runtime budget, supporting the principle that memory should preserve the distinctions that matter for decisions, not descriptions.

cs.AI↗

RobustSora: De-Watermarked Benchmark for Robust AI-Generated Video Detection

The proliferation of AI-generated video models poses new challenges to information integrity and digital trust. A key confound, however, remains unaddressed: commercial generators embed visible overlay watermarks for provenance tracking, yet no existing benchmark controls for this variable, leaving open whether detectors learn genuine generation artefacts or merely associate watermark patterns with AI-generated labels. We present RobustSora, a benchmark of 6,500 manually verified videos in four categories: Authentic-Clean (A-C), Generated-Watermarked (G-W), Generated-DeWatermarked (G-DeW), and Authentic-Spoofed (A-S), sourced from Vript, DVF, and UltraVideo (authentic) and from Sora, Sora 2, Pika, Open-Sora 2, and KLing (generated). Two evaluation tasks isolate watermark effects: Task-I (Watermark Erasure Robustness) tests detection on watermark-removed AI videos; Task-II (Watermark Spoofing Robustness) measures false-alarm rates on authentic videos injected with fake watermarks. Across ten models spanning specialized detectors, transformer classifiers, and MLLMs, watermark manipulation induces accuracy changes of $-9.4$ to $+1.6$ pp (mean 6.6 pp; $p{<}0.01$ for 7/10 models on each task). A placebo control bounds inpainting-artefact confounds at $\le$2 pp, and a watermark-aware training augmentation recovers 3-4 pp on both tasks, together providing causal evidence that detectors actively rely on watermark cues. Per-generator breakdown shows that Sora 2 induces drops of $-11$ to $-14$ pp versus $-3$ to $-6$ pp for Pika and Open-Sora 2, indicating that watermark prominence, rather than detector architecture, is the principal driver of dependency. These results argue for watermark-aware evaluation and training in AIGC video detection. Dataset, evaluation code, and pretrained checkpoints will be released.

cs.CV↗

CoTEvol: Self-Evolving Chain-of-Thoughts for Data Synthesis in Mathematical Reasoning

Large Language Models (LLMs) exhibit strong mathematical reasoning when trained on high-quality Chain-of-Thought (CoT) that articulates intermediate steps, yet costly CoT curation hinders further progress. While existing remedies such as distillation from stronger LLMs and self-synthesis based on test-time search alleviate this issue, they often suffer from diminishing returns or high computing overhead.In this work, we propose CoTEvol, a genetic evolutionary framework that casts CoT generation as a population-based search over reasoning trajectories.Candidate trajectories are iteratively evolved through reflective global crossover at the trajectory level and local mutation guided by uncertainty at the step level, enabling holistic recombination and fine-grained refinement. Lightweight, task-aware fitness functions are designed to guide the evolutionary process toward accurate and diverse reasoning. Empirically, CoTEvol improves correct-CoT synthesis success by over 30% and enhances structural diversity, with markedly improved efficiency. LLMs trained on these evolutionary CoT data achieve an average gain of 6.6% across eight math benchmarks, outperforming previous distillation and self-synthesis approaches. These results underscore the promise of evolutionary CoT synthesis as a scalable and effective method for mathematical reasoning tasks.

cs.AI↗