Search arXiv⌕ Search

arXiv subjects

Shengjie Zhou

Publications and source records attributed to Shengjie Zhou.

7 recordsLinked to original sources

HySparse2: Hybrid Sparse Attention with Two-Level KV Sharing

Long-horizon and multi-turn agents typically generate short actions and process long observations from tools and environments. This growing context demands efficient prefill, compact KV-cache storage, and accurate long-context retrieval. To meet these demands, we introduce HySparse2, a hybrid sparse attention architecture with two-level KV sharing. At the outer level, KV Bridging adopts a YOCO-style self-decoder and cross-decoder structure, but bridges only full-attention layers. The self-decoder uses hybrid sliding-window attention (SWA), while the cross-decoder uses hybrid sparse attention. The KV caches for full-attention layers in the cross-decoder are generated from the hidden states of full-attention layers in the self-decoder. At the inner level, HySparse2 retains HySparse's core KV Reuse design with two refinements. First, it replaces block-level sparsity with token-level sparsity for finer long-context retrieval. Second, it removes the separate SWA branch from sparse layers and instead forces a sliding window of recent tokens into the sparse selection. This two-level KV sharing allows all cross-decoder KV caches to be constructed from self-decoder hidden states. Prefill can therefore exit after the self-decoder, skipping all cross-decoder layers. On an 80B-A3B MoE model, HySparse2 outperforms HySparse and Hybrid SWA on long-context retrieval and multi-turn agentic tasks, while substantially reducing prefill computation and KV-cache storage.

cs.CL↗

Improving Generalizability and Undetectability for Targeted Adversarial Attacks on Multimodal Pre-trained Models

Multimodal pre-trained models (e.g., ImageBind), which align distinct data modalities into a shared embedding space, have shown remarkable success across downstream tasks. However, their increasing adoption raises serious security concerns, especially regarding targeted adversarial attacks. In this paper, we show that existing targeted adversarial attacks on multimodal pre-trained models still have limitations in two aspects: generalizability and undetectability. Specifically, the crafted targeted adversarial examples (AEs) exhibit limited generalization to partially known or semantically similar targets in cross-modal alignment tasks (i.e., limited generalizability) and can be easily detected by simple anomaly detection methods (i.e., limited undetectability). To address these limitations, we propose a novel method called Proxy Targeted Attack (PTA), which leverages multiple source-modal and target-modal proxies to optimize targeted AEs, ensuring they remain evasive to defenses while aligning with multiple potential targets. We also provide theoretical analyses to highlight the relationship between generalizability and undetectability and to ensure optimal generalizability while meeting the specified requirements for undetectability. Furthermore, experimental results demonstrate that our PTA can achieve a high success rate across various related targets and remain undetectable against multiple anomaly detection methods.

cs.CV↗

KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents

As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts. Its 238 tasks are manually screened from more than 1,000 generated candidates and combine a user simulator, stateful tools, deterministic environment assertions, an open-source natural-language evaluator, and human trajectory verification. Evaluation of nine models, including DeepSeek-V4-Flash, GLM-5.2, and MiniMax-M3, shows substantial cross-domain variation: no model handles factual correction, identity consistency checking, and temporal conflict resolution reliably across all settings. In the simulated environments, missed conflicts can propagate to tool calls or synthetic protected-data flows. KC-Bench isolates this model-level behavior rather than ranking complete agent frameworks, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.

cs.AI↗

Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents

The paper introduces GUI-Owl-1.5, the latest native GUI agent model that features instruct/thinking variants in multiple sizes (2B/4B/8B/32B/235B) and supports a range of platforms (desktop, mobile, browser, and more) to enable cloud-edge collaboration and real-time interaction. GUI-Owl-1.5 achieves state-of-the-art results on more than 20+ GUI benchmarks on open-source models: (1) on GUI automation tasks, it obtains 56.5 on OSWorld, 71.6 on AndroidWorld, and 48.4 on WebArena; (2) on grounding tasks, it obtains 80.3 on ScreenSpotPro; (3) on tool-calling tasks, it obtains 47.6 on OSWorld-MCP, and 46.8 on MobileWorld; (4) on memory and knowledge tasks, it obtains 75.5 on GUI-Knowledge Bench. GUI-Owl-1.5 incorporates several key innovations: (1) Hybird Data Flywheel: we construct the data pipeline for UI understanding and trajectory generation based on a combination of simulated environments and cloud-based sandbox environments, in order to improve the efficiency and quality of data collection. (2) Unified Enhancement of Agent Capabilities: we use a unified thought-synthesis pipeline to enhance the model's reasoning capabilities, while placing particular emphasis on improving key agent abilities, including Tool/MCP use, memory and multi-agent adaptation; (3) Multi-platform Environment RL Scaling: We propose a new environment RL algorithm, MRPO, to address the challenges of multi-platform conflicts and the low training efficiency of long-horizon tasks. The GUI-Owl-1.5 models are open-sourced, and an online cloud-sandbox demo is available at https://github.com/X-PLUG/MobileAgent.

cs.AI↗

A2Eval: Agentic and Automated Evaluation for Embodied Brain

Current embodied VLM evaluation relies on static, expert-defined, manually annotated benchmarks that exhibit severe redundancy and coverage imbalance. This labor intensive paradigm drains computational and annotation resources, inflates costs, and distorts model rankings, ultimately stifling iterative development. To address this, we propose Agentic Automatic Evaluation (A2Eval), the first agentic framework that automates benchmark curation and evaluation through two collaborative agents. The Data Agent autonomously induces capability dimensions and assembles a balanced, compact evaluation suite, while the Eval Agent synthesizes and validates executable evaluation pipelines, enabling fully autonomous, high-fidelity assessment. Evaluated across 10 benchmarks and 13 models, A2Eval compresses evaluation suites by 85%, reduces overall computational costs by 77%, and delivers a 4.6x speedup while preserving evaluation quality. Crucially, A2Eval corrects systematic ranking biases, improves human alignment to Spearman's rho=0.85, and maintains high ranking fidelity (Kendall's tau=0.81), establishing a new standard for high-fidelity, low-cost embodied assessment. Our code and data will be public soon.

cs.CL↗

Three-Dimensional Fermi Surface, Van Hove Singularity and Enhancement of Superconductivity in Infinite-Layer Nickelates

Recent experiments reveal a three-dimensional (3D) Fermi surface with a clear $k_z$ dispersion in infinite-layer nickelates, distinguishing them from their cuprate superconductor counterparts. However, the impact of this difference on the superconducting properties of nickelates remains unclear. Here, we employ a combined random-phase-approximation and dynamical-mean-field-theory (RPA+DMFT) approach to solve the linearized gap equation for superconductivity. We find that, compared to the cuprate-like two-dimensional (2D) single-orbital Fermi surface, the van Hove singularities on the 3D Fermi surface of infinite-layer nickelates strengthen spin fluctuations by driving the system closer to antiferromagnetic instabilities, thereby significantly enhancing superconductivity. Our findings underscore the critical role of the van Hove singularities in shaping the superconducting properties of infinite-layer nickelates and, more broadly, highlight the importance of subtle Fermi surface features in modeling material-specific unconventional superconductors.

cond-mat.supr-con↗

Sensitive dependence of pairing symmetry on Ni-$e_g$ crystal field splitting in the nickelate superconductor La$_3$Ni$_2$O$_7$

The discovery of high-temperature superconductivity in La$_3$Ni$_2$O$_7$ under pressure has drawn great attention. However, consensus has not been reached on its pairing symmetry in theory. By combining density-functional-theory (DFT), maximally-localized-Wannier-function, and linearized gap equation with random-phase-approximation, we find that the pairing symmetry of La$_3$Ni$_2$O$_7$ is $d_{xy}$, if its DFT band structure is accurately reproduced by a downfolded bilayer two-orbital model. More importantly, we reveal that the pairing symmetry of La$_3$Ni$_2$O$_7$ sensitively depends on the crystal field splitting between two Ni-$e_g$ orbitals. A slight increase in Ni-$e_g$ crystal field splitting alters the pairing symmetry from $d_{xy}$ to $s_{\pm}$. Such a transition is associated with the change in inverse Fermi velocity and susceptibility, while the shape of Fermi surface remains almost unchanged. Our work highlights the sensitive dependence of pairing symmetry on low-energy electronic structure in multi-orbital superconductors, which calls for care in the downfolding procedure when one calculates their pairing symmetry.

cond-mat.supr-con↗