VeriWeb: Verifiable Long-Chain Web Benchmark for Agentic Information-Seeking
Shunyu Liu , Minghao Liu , Huichi Zhou , Zhenyu Cui , Yang Zhou , Yuhao Zhou , Jialiang Gao , Heng Zhou , Yunhao Yang , Wendong Fan , puzhen zhang , Ge Zhang , Jiajun Shi , Weihao Xuan , Jiaxing Huang , Shuang Luo , Fang Wu , Heli Qi , Qingcheng Zeng , Junjie Wang , Aosong Feng , Jindi Lv , Sicong Jiang , Ziqi Ren , Wangchunshu Zhou , Zhenfei Yin , Wenlong Zhang , Guohao Li , Wenhao Yu , Lei Ma , Lei Bai , Qunshu Lin , Mingli Song , Dacheng Tao
- 🏛 Institutions
- NTU , ZJU , University of Tokyo , Shanghai AI Laboratory , Google DeepMind , University of Alberta
- 📅 Date
- August 6, 2025
- 📑 Publisher
- arXiv
- 💻 Env
- Web
- 🔑 Keywords
VeriWeb is a web benchmark for long-chain information-seeking tasks that decomposes each problem into interdependent, verifiable subtasks instead of relying only on final-answer checks. It contains 302 human-annotated tasks across five domains and is designed to stress both coverage-oriented search and multi-hop context tracking in realistic web environments.
- AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human DemonstrationsNovember 24, 2024 · ACL 2025
- OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human DemonstrationsSeptember 2, 2026 · arXiv
- AppAgent: Multimodal Agents as Smartphone UsersDecember 21, 2023 · CHI 2025
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgetsSeptember 11, 2026 · arXiv
- Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step SupervisionSeptember 2, 2026 · arXiv
- Discriminative World Models for Web AgentsSeptember 2, 2026 · arXiv
- Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection GuardrailsSeptember 2, 2026 · arXiv
- When and What to Teach: Budget-Aware Online Adaptation for Web AgentsAugust 31, 2026 · arXiv
- SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill AbstractionAugust 31, 2026 · arXiv
- Learning Simple Test-Time Environments for LLM Web AgentsAugust 29, 2026 · arXiv
- WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM AgentsAugust 25, 2026 · arXiv
- BrowserForge: Scaling Web Episode via Parallel Browser SandboxesAugust 25, 2026 · arXiv
- Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent LearningAugust 22, 2026 · arXiv
- Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and WebAugust 22, 2026 · arXiv
- ComponentBench: Diagnosing Component-Level Failures in Computer-Use AgentsAugust 18, 2026 · COLM 2026
- StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt InjectionAugust 6, 2026 · arXiv
- Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web AgentsAugust 6, 2026 · arXiv
- LoginTrap: Uncovering Task-Agnostic Phishing-Style Indirect Prompt Injection Attacks against LLM-based Web AgentsAugust 5, 2026 · arXiv
- Qwen-CUA: Native Computer Use for (almost) EverythingAugust 3, 2026 · arXiv
- MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action DistillationJuly 31, 2026 · arXiv
- Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI AgentsJuly 30, 2026 · arXiv
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward ModelsJuly 30, 2026 · arXiv
- EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition UnderstandingJuly 19, 2026 · arXiv
- StructAgent: Harness Long-horizon Digital Agents with Unified Causal StructureJuly 13, 2026 · arXiv