EntWorld: A Holistic Environment and Benchmark for Verifiable Enterprise GUI Agents
Ying Mo , Yu Bai , Dapeng Sun , Yuqian Shi , Yukai Miao , Li Chen , Dan Li
- 🏛 Institutions
- Zhongguancun Laboratory , Tsinghua
- 📅 Date
- January 25, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- Desktop
- 🔑 Keywords
TLDR
EntWorld introduces a verifiable enterprise-agent environment and a 1,756-task benchmark spanning six business domains such as CRM, ITIL, and ERP. It synthesizes workflows from database schemas and uses SQL-based deterministic verification instead of visual matching, and current top models still trail human performance by a large margin.
Related papers (24)
- ScienceBoard: Evaluating Multimodal Autonomous Agents in Realistic Scientific WorkflowsMay 26, 2025 · ICLR 2026 (Poster)
- WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?March 11, 2024 · ICML 2024
- WebArena: A Realistic Web Environment for Building Autonomous AgentsJuly 25, 2023 · ICLR 2024 (Poster)
- WebShop: Towards Scalable Real-World Web Interaction with Grounded Language AgentsJuly 31, 2022 · NeurIPS 2022
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionSeptember 9, 2026 · arXiv
- FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?September 7, 2026 · arXiv
- CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI AgentsSeptember 4, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- CUADebug: Diagnosing and Repairing Computer-Use Agent FailuresJuly 31, 2026 · arXiv
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward ModelsJuly 30, 2026 · arXiv
- OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic GroundingJuly 29, 2026 · arXiv
- Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional FieldsJune 9, 2026 · arXiv
- WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application EnvironmentsApril 30, 2026 · arXiv
- The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use AgentsApril 12, 2026 · arXiv
- HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration TasksApril 10, 2026 · arXiv
- Gym-Anything: Turn any Software into an Agent EnvironmentApril 7, 2026 · arXiv
- HippoCamp: Benchmarking Contextual Agents on Personal ComputersApril 1, 2026 · arXiv
- PIRA-Bench: A Transition from Reactive GUI Agents to GUI-based Proactive Intent Recommendation AgentsMarch 9, 2026 · arXiv
- OSExpert: Computer-Use Agents Learning Professional Skills via ExplorationMarch 9, 2026 · arXiv
- When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use AgentsFebruary 9, 2026 · arXiv
- When Actions Go Off-Task: Detecting and Correcting Misaligned Actions in Computer-Use AgentsFebruary 9, 2026 · arXiv
- OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive TasksJanuary 28, 2026 · arXiv
- MirrorGuard: Toward Secure Computer-Use Agents via Simulation-to-Real Reasoning CorrectionJanuary 19, 2026 · arXiv
- ShowUI-π: Flow-based Generative Models as GUI Dexterous HandsDecember 31, 2025 · arXiv