OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use
Xueyu Hu , Tao Xiong , Biao Yi , Zishu Wei , Ruixuan Xiao , Yurun Chen , Jiasheng Ye , Meiling Tao , Xiangxin Zhou , Ziyu Zhao , Yuhuai Li , Shengze Xu , Shenzhi Wang , Xinchen Xu , Shuofei Qiao , Zhaokai Wang , Kun Kuang , Tieyong Zeng , Liang Wang , Jiwei Li , Yuchen Eleanor Jiang , Wangchunshu Zhou , Guoyin Wang , Keting Yin , Zhou Zhao , Hongxia Yang , Fan Wu , Shengyu Zhang , Fei Wu
- 🏛 Institutions
- ZJU , Fudan , OPPO AI Center , University of Chinese Academy of Sciences , Institute of Automation , CAS , CUHK , Tsinghua , SJTU , 01.AI , PolyU
- 📅 Date
- December 20, 2024
- 📑 Publisher
- ACL 2025
- 💻 Env
- General GUI
- 🔑 Keywords
TLDR
This survey reviews MLLM-based OS agents across computers, phones, and browsers, covering their environments, observation and action spaces, capabilities, and system designs. It also organizes the benchmark landscape and highlights open problems such as safety, privacy, personalization, and self-evolution.
Related papers (24)
- GUI Agents: A SurveyDecember 18, 2024 · Findings of ACL 2025
- A Survey of WebAgents: Towards Next-Generation AI Agents for Web Automation with Large Foundation ModelsMarch 30, 2025 · KDD 2025
- A Survey on the Safety and Security Threats of Computer-Using Agents: JARVIS or Ultron?May 16, 2025 · arXiv
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use AgentsSeptember 6, 2026 · arXiv
- Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime OptimizationSeptember 2, 2026 · arXiv
- SIR: Self-improving Red-teaming for Compute Use AgentsAugust 31, 2026 · arXiv
- Software Engineering for and with GUI AgentAugust 10, 2026 · arXiv
- Human-Guided Harm Recovery for Computer Use AgentsApril 20, 2026 · arXiv
- Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element InjectionApril 9, 2026 · arXiv
- LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial ScenariosFebruary 3, 2026 · arXiv
- SafePred: A Predictive Guardrail for Computer-Using Agents via World ModelsFebruary 2, 2026 · arXiv
- Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-MakingJanuary 30, 2026 · arXiv
- GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI AgentsMay 19, 2025 · arXiv
- A Survey on GUI Agents with Foundation Models Enhanced by Reinforcement LearningApril 29, 2025 · arXiv
- Towards Trustworthy GUI Agents: A SurveyMarch 30, 2025 · arXiv
- GUI Agents with Foundation Models: A Comprehensive SurveyNovember 7, 2024 · arXiv
- WiP: Characterizing and Defending Against Mobile-Agent-Driven MFA AutomationSeptember 2, 2026 · arXiv
- Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection GuardrailsSeptember 2, 2026 · arXiv
- WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM AgentsAugust 25, 2026 · arXiv
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial EnvironmentsAugust 25, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android AppsAugust 18, 2026 · arXiv
- StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt InjectionAugust 6, 2026 · arXiv
- SeerGuard: A Safety Framework for Mobile GUI Agents via World Model PredictionJuly 17, 2026 · arXiv