GUI Agents with Foundation Models: A Comprehensive Survey
Shuai Wang , Weiwen Liu , Jingxuan Chen , Yuqi Zhou , Weinan Gan , Xingshan Zeng , Yuhan Che , Shuai Yu , Xinlong Hao , Kun Shao , Bin Wang , Chuhan Wu , Yasheng Wang , Ruiming Tang , Jianye Hao
- 🏛 Institutions
- Huawei Noah's Ark Lab
- 📅 Date
- November 7, 2024
- 📑 Publisher
- arXiv
- 💻 Env
- General GUI
- 🔑 Keywords
TLDR
This survey organizes foundation-model GUI agents around data resources, agent construction, taxonomy, and industrial applications. It also summarizes open challenges around the benchmark-reality gap, agent self-evolution, and inference efficiency.
Related papers (24)
- Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital PlatformsNovember 17, 2024 · arXiv
- Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime OptimizationSeptember 2, 2026 · arXiv
- Software Engineering for and with GUI AgentAugust 10, 2026 · arXiv
- How Smart Is Your GUI Agent? A Framework for the Future of Software InteractionFebruary 12, 2026 · arXiv
- A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?May 16, 2025 · arXiv
- A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement LearningApril 29, 2025 · arXiv
- Towards Trustworthy GUI Agents: A SurveyMarch 30, 2025 · arXiv
- OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser UseDecember 20, 2024 · ACL 2025
- GUI Agents: A SurveyDecember 18, 2024 · Findings of ACL 2025
- Mapping the Design Space of User Experience for Computer Use AgentsFebruary 7, 2026 · IUI 2026
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and ProspectsApril 28, 2025 · TMLR 2025
- A Survey of WebAgents: Towards Next-Generation AI Agents for Web Automation with Large Foundation ModelsMarch 30, 2025 · KDD 2025
- WebSuite: Systematically Evaluating Why Web Agents FailJune 1, 2024 · arXiv
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgetsSeptember 11, 2026 · arXiv
- TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI AgentsSeptember 9, 2026 · arXiv
- FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?September 7, 2026 · arXiv
- Selective Knowledge Control for Continual GUI Agent Learning over Application StreamsSeptember 6, 2026 · arXiv
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use AgentsSeptember 6, 2026 · arXiv
- From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use AgentsSeptember 4, 2026 · arXiv
- Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI AgentsSeptember 3, 2026 · arXiv
- SIR: Self-improving Red-teaming for Compute Use AgentsAugust 31, 2026 · arXiv
- Iron: Intent-Aligned and Retrospective Dual Learning Framework for Enhancing Generalist Virtual AgentsAugust 28, 2026 · arXiv
- UI-Venus-2 Technical ReportAugust 27, 2026 · arXiv
- Task-Adaptive Rubrics for GUI Reward ModelingAugust 25, 2026 · arXiv