GUI Agents Papers
Star · 902

TurkingBench: A Challenge Benchmark for Web Agents

Kevin Xu , Yeganeh Kordi , Tanay Nayak , Adi Asija , Yizhong Wang , Kate Sanders , Adam Byerly , Jingyu Zhang , Benjamin Van Durme , Daniel Khashabi

🏛 Institutions
JHU , Brown University , University of Washington
📅 Date
March 18, 2024
📑 Publisher
NAACL 2025 (Oral)
💻 Env
Web
🔑 Keywords
TLDR

TurkingBench is a web-agent benchmark built from real crowdsourcing task pages instead of synthetic websites, with 158 tasks and 32.2K instantiated examples. It evaluates both language-only and multimodal models through an action-execution layer that maps model outputs to webpage actions, and shows large remaining performance gaps on these realistic web tasks.

Open paper Publisher Report issue
Related papers (24)