CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use Agents
Marta Sumyk , Oleksandr Kosovan
- 🏛 Institutions
- Ukrainian Catholic University
- 📅 Date
- March 11, 2026
- 📑 Publisher
- HEAL @ CHI 2026 Workshop
- 💻 Env
- Desktop
- 🔑 Keywords
TLDR
CUAAudit studies vision-language models as autonomous judges of desktop-agent task success from observable interactions alone. Across multiple operating-system benchmarks, it finds that even strong VLM auditors degrade on harder environments and disagree substantially with one another, highlighting limits of model-based auditing.
Related papers (24)
- Same Outcomes, Different Journeys: A Trace-Level Framework for Comparing Human and GUI-Agent Behavior in Production Search SystemsApril 9, 2026 · arXiv
- GUIDE: Interpretable GUI Agent Evaluation via Hierarchical DiagnosisApril 6, 2026 · arXiv
- MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic EnvironmentsFebruary 3, 2026 · arXiv
- WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph RepresentationOctober 22, 2025 · NeurIPS 2025 Workshop on Multi-Turn Interactions in Large Language Models
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-JudgeSeptember 18, 2025 · NeurIPS 2025 Datasets & Benchmarks Track (Poster)
- You Don’t Know Until You Click: Automated GUI Testing for Production-Ready Software EvaluationAugust 17, 2025 · SEA @ NeurIPS 2025 (Poster)
- An Illusion of Progress? Assessing the Current State of Web AgentsApril 2, 2025 · COLM 2025
- GUI Agents: A SurveyDecember 18, 2024 · Findings of ACL 2025
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionSeptember 9, 2026 · arXiv
- FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?September 7, 2026 · arXiv
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use AgentsSeptember 6, 2026 · arXiv
- From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use AgentsSeptember 4, 2026 · arXiv
- CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI AgentsSeptember 4, 2026 · arXiv
- OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human DemonstrationsSeptember 2, 2026 · arXiv
- Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop ApplicationsSeptember 1, 2026 · arXiv
- SIR: Self-improving Red-teaming for Compute Use AgentsAugust 31, 2026 · arXiv
- CURA: Certified Runtime Alarms for Computer-Use AgentsAugust 28, 2026 · arXiv
- ASIL: Replacing Screenshot-and-Click with Structured State and Semantic ActionsAugust 27, 2026 · arXiv
- LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI AgentsAugust 26, 2026 · arXiv
- Reflection with Action-Induced Visual Differences for Desktop GUI AgentsAugust 25, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model TrajectoriesAugust 23, 2026 · arXiv
- Spine-Branch Coordination for Multi-agent Computer UseAugust 22, 2026 · arXiv
- Inducing Task Models from Computer-Use TracesAugust 20, 2026 · arXiv