ClawBench: Can AI Agents Complete Everyday Online Tasks?
Yuxuan Zhang , Yubo Wang , Yipeng Zhu , Penghui Du , Junwen Miao , Xuan Lu , Wendong Xu , Yunzhuo Hao , Songcheng Cai , Xiaochen Wang , Huaisong Zhang , Xian Wu , Yi Lu , Minyi Lei , Kai Zou , Huifeng Yin , Ping Nie , Liang Chen , Dongfu Jiang , Wenhu Chen , Kelsey R. Allen
- 🏛 Institutions
- UBC , Vector Institute , CMU , UWaterloo , SJTU , ZJU , HKUST , Tsinghua
- 📅 Date
- April 9, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- Web
- 🔑 Keywords
TLDR
ClawBench evaluates browser agents on 283 everyday online tasks (V1 153 + V2 130) across 163 live production websites spanning purchases, bookings, and job applications. A lightweight interception layer blocks final submissions for safe evaluation while preserving end-to-end interaction. Its two-stage interception and judge protocol exposes a persistent gap in real-world web automation.
Related papers (24)
- AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?October 21, 2024 · EMNLP 2024 (Poster)
- An Illusion of Progress? Assessing the Current State of Web AgentsApril 2, 2025 · COLM 2025
- Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional FieldsJune 9, 2026 · arXiv
- ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental SynthesisMay 24, 2026 · arXiv
- CocoaBench: Evaluating Unified Digital Agents in the WildApril 13, 2026 · arXiv
- HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration TasksApril 10, 2026 · arXiv
- Gym-Anything: Turn any Software into an Agent EnvironmentApril 7, 2026 · arXiv
- AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI AgentsMarch 19, 2026 · arXiv
- OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive TasksJanuary 28, 2026 · arXiv
- MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World EnvironmentJanuary 28, 2026 · Findings of ACL 2026
- LongHorizonUI: A Unified Framework for Robust long-horizon Task Automation of GUI AgentJanuary 26, 2026 · ICLR 2026 (Poster)
- MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented EnvironmentsDecember 22, 2025 · arXiv
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgetsSeptember 11, 2026 · arXiv
- ComponentBench: Diagnosing Component-Level Failures in Computer-Use AgentsAugust 18, 2026 · COLM 2026
- StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt InjectionAugust 6, 2026 · arXiv
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward ModelsJuly 30, 2026 · arXiv
- EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition UnderstandingJuly 19, 2026 · arXiv
- StructAgent: Harness Long-horizon Digital Agents with Unified Causal StructureJuly 13, 2026 · arXiv
- WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent EvaluationJuly 7, 2026 · arXiv
- Odysseys: Benchmarking Web Agents on Realistic Long Horizon TasksApril 27, 2026 · arXiv
- WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent BenchmarkApril 13, 2026 · arXiv
- The Amazing Agent Race: Strong Tool Users, Weak NavigatorsApril 11, 2026 · arXiv
- GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game AgentsApril 8, 2026 · arXiv
- WebSP-Eval: Evaluating Web Agents on Website Security and Privacy TasksApril 7, 2026 · arXiv