GUI Agents Papers
Star · 902

ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents

Tianchen Guan , Xinlei Lin , Royce Cheng-Yue , Xiangjun Wang , Shuyan Zhou

🏛 Institutions
Duke University , Amazon AGI SF Lab
📅 Date
August 18, 2026
📑 Publisher
COLM 2026
💻 Env
Web
🔑 Keywords
TLDR

ComponentBench targets the under-instrumented middle layer between long-horizon workflow benchmarks and atomic GUI-grounding tests, with a library-agnostic ontology of 97 canonical UI components instantiated as 2,910 programmatically verified, component-centered tasks on modern web UIs.

Open paper Report issue
Related papers (24)