arXiv · 2507.00938
WebArxiv: A Reproducible Benchmark for Evaluating Multimodal Web Agents on arXiv Tasks
Abstract
Foundation models now enable autonomous agents to interact with real-world websites, but existing benchmarks emphasize general-purpose browsing, underrepresent research-oriented environments and scholarly discovery workflows, and often depend on live sites whose changing content and structure undermine reproducibility. arXiv provides a realistic, reproducible, hierarchically structured, information-centric testbed without privacy-sensitive interactions. We introduce WebArxiv, a static-snapshot benchmark comprising 510 time-invariant tasks, each with a unique deterministic ground truth. Its diverse, realistic scholarly tasks go beyond simple information lookup and rule following to emphasize multi-constraint paper retrieval, fine-grained content extraction, and cross-paper comparison. Evaluations of a range of foundation-model-based web agents show that WebArxiv remains challenging. Behavioral analysis reveals that agents over-rely on fixed interaction histories, causing incomplete or repetitive reasoning. We therefore equip agents with a lightweight dynamic-memory mechanism for adaptive retrieval and reasoning over relevant context. The benchmark and code are available at https://anonymous.4open.science/r/74E4423BVNW/README.md.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zihao Sun, Zijing Shi, Ling Chen. 2026-09-22. WebArxiv: A Reproducible Benchmark for Evaluating Multimodal Web Agents on arXiv Tasks. https://arxiv.org/abs/2507.00938
Cite the original work for its findings. Save a collection to share your selection of sources.