GUI Agents: A Survey
Dang Nguyen , Jian Chen , Yu Wang , Gang Wu , Namyong Park , Zhengmian Hu , Hanjia Lyu , Junda Wu , Ryan Aponte , Yu Xia , Xintong Li , Jing Shi , Hongjie Chen , Viet Dac Lai , Zhouhang Xie , Sungchul Kim , Ruiyi Zhang , Tong Yu , Mehrab Tanjim , Nesreen K. Ahmed , Puneet Mathur , Seunghyun Yoon , Lina Yao , Branislav Kveton , Jihyung Kil , Thien Huu Nguyen , Trung Bui , Tianyi Zhou , Ryan A. Rossi , Franck Dernoncourt
- 🏛 Institutions
- UMD , State University of New York at Buffalo , University of Oregon , Adobe Research , University of Rochester , UC San Diego , CMU , Dolby Labs , Cisco Research , University of New South Wales
- 📅 Date
- December 18, 2024
- 📑 Publisher
- Findings of ACL 2025
- 💻 Env
- General GUI
- 🔑 Keywords
TLDR
This survey organizes GUI-agent research around benchmarks, evaluation metrics, architectures, and training methods for agents powered by large foundation models. It proposes a unified perception-reasoning-planning-acting framework and highlights the open problems that remain across the stack.
Related papers (24)
- OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser UseDecember 20, 2024 · ACL 2025
- A Survey of WebAgents: Towards Next-Generation AI Agents for Web Automation with Large Foundation ModelsMarch 30, 2025 · KDD 2025
- Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime OptimizationSeptember 2, 2026 · arXiv
- Software Engineering for and with GUI AgentAugust 10, 2026 · arXiv
- GUIDE: Interpretable GUI Agent Evaluation via Hierarchical DiagnosisApril 6, 2026 · arXiv
- Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-MakingJanuary 30, 2026 · arXiv
- A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?May 16, 2025 · arXiv
- A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement LearningApril 29, 2025 · arXiv
- Towards Trustworthy GUI Agents: A SurveyMarch 30, 2025 · arXiv
- GUI Agents with Foundation Models: A Comprehensive SurveyNovember 7, 2024 · arXiv
- Same Outcomes, Different Journeys: A Trace-Level Framework for Comparing Human and GUI-Agent Behavior in Production Search SystemsApril 9, 2026 · arXiv
- CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use AgentsMarch 11, 2026 · HEAL @ CHI 2026 Workshop
- MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic EnvironmentsFebruary 3, 2026 · arXiv
- WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph RepresentationOctober 22, 2025 · NeurIPS 2025 Workshop on Multi-Turn Interactions in Large Language Models
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and ProspectsApril 28, 2025 · TMLR 2025
- An Illusion of Progress? Assessing the Current State of Web AgentsApril 2, 2025 · COLM 2025
- Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital PlatformsNovember 17, 2024 · arXiv
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgetsSeptember 11, 2026 · arXiv
- TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI AgentsSeptember 9, 2026 · arXiv
- FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?September 7, 2026 · arXiv
- Selective Knowledge Control for Continual GUI Agent Learning over Application StreamsSeptember 6, 2026 · arXiv
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use AgentsSeptember 6, 2026 · arXiv
- From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use AgentsSeptember 4, 2026 · arXiv
- Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI AgentsSeptember 3, 2026 · arXiv