GUI Agents Papers
Star · 902

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Jingbo Zhou , Yusai Zhao , Qi Bao , Jingjia Cao , Zhenghai Chen , Chang Gao , Kaiqi Guo , Muxin Guo , Mingxuan Li , Xinjiang Lu , Yanru Ma , Yixiong Xiao , Zenghui Zhang , Le Zhang , Hua Wu

🏛 Institutions
Unknown
📅 Date
July 29, 2026
📑 Publisher
arXiv
💻 Env
Desktop
🔑 Keywords
TLDR

OmegaUse-OfficeVal benchmarks agents on 100 long-horizon office-suite tasks with code-based verifiers and task-level human labor and price signals. The economic annotations enable comparisons of deliverable quality, inference cost, and execution time against human work.

Open paper arXiv Report issue
Related papers (24)