You Don’t Know Until You Click: Automated GUI Testing for Production-Ready Software Evaluation
Yutong Bian , Xianhao Lin , Yupeng Xie , Tianyang Liu , Mingchen Zhuge , Siyuan Lu , Haoming Tang , Jinlin Wang , Jiayi Zhang , Jiaqi Chen , Xiangru Tang , Yongxin Ni , Sirui Hong , Chenglin Wu
- 🏛 Institutions
- DeepWisdom , Fudan , HKUST(GZ) , UC San Diego , KAUST , Westlake University , Stanford , Yale University , NUS
- 📅 Date
- August 17, 2025
- 📑 Publisher
- SEA @ NeurIPS 2025 (Poster)
- 💻 Env
- General GUI
- 🔑 Keywords
TLDR
RealDevWorld is an evaluation framework for repository-scale software generation that judges whether produced applications actually work when interacted with through their GUIs. It pairs a 194-task benchmark, RealDevBench, with AppEvalPilot, an agent-as-a-judge system for functional, visual, and runtime evaluation, and reports strong alignment with expert human assessments.
Related papers (24)
- CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use AgentsMarch 11, 2026 · HEAL @ CHI 2026 Workshop
- Mind2Web 2: Evaluating Agentic Search with Agent-as-a-JudgeSeptember 18, 2025 · NeurIPS 2025 Datasets & Benchmarks Track (Poster)
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgetsSeptember 11, 2026 · arXiv
- TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI AgentsSeptember 9, 2026 · arXiv
- FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?September 7, 2026 · arXiv
- Selective Knowledge Control for Continual GUI Agent Learning over Application StreamsSeptember 6, 2026 · arXiv
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use AgentsSeptember 6, 2026 · arXiv
- From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use AgentsSeptember 4, 2026 · arXiv
- Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI AgentsSeptember 3, 2026 · arXiv
- Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime OptimizationSeptember 2, 2026 · arXiv
- SIR: Self-improving Red-teaming for Compute Use AgentsAugust 31, 2026 · arXiv
- Iron: Intent-Aligned and Retrospective Dual Learning Framework for Enhancing Generalist Virtual AgentsAugust 28, 2026 · arXiv
- UI-Venus-2 Technical ReportAugust 27, 2026 · arXiv
- Task-Adaptive Rubrics for GUI Reward ModelingAugust 25, 2026 · arXiv
- CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI AgentsAugust 23, 2026 · arXiv
- GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI GroundingAugust 22, 2026 · arXiv
- Beyond Success and Failure: Length-Aware Contrastive Learning for GUI AgentsAugust 22, 2026 · arXiv
- Software Engineering for and with GUI AgentAugust 10, 2026 · arXiv
- Hallucination-Free GUI Grounding via Regression-Free Layout-Aware MatchingAugust 10, 2026 · arXiv
- GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMsAugust 4, 2026 · arXiv
- Naive Visual Memory is Not Enough: A Failure-Mode Study of GUI AgentsJune 12, 2026 · arXiv
- Demo2Tutorial: From Human Experience to Multimodal Software TutorialsJune 2, 2026 · CVPR 2026
- STaR-KV: Spatio-Temporal Adaptive Re-weighting for KV Cache Compression in GUI Vision-Language ModelsJune 1, 2026 · arXiv
- GUI-C²: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement LearningMay 29, 2026 · arXiv