Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial Environments
Guo Gan , Yilun Zhao , Cong Chen , Jinbiao Wei , Tingyu Song , Zheyuan Yang , Lin Fu , Hong Zhou
- 🏛 Institutions
- Zhejiang University , Yale University , University of Chinese Academy of Sciences , Tongji University
- 📅 Date
- August 25, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- Mobile
- 🔑 Keywords
TLDR
AnTrap is an AndroidWorld-based benchmark that injects controllable runtime anomalies, organized into four categories with ten subcategories, into Android GUI-agent trajectories. Testing 16 leading GUI agents reveals substantial robustness degradation under these anomalies, and adversarial RL training can address many single-step traps but struggles with deep contextual failures such as state deadlock.
Related papers (24)
- Don't Act Blindly: Robust GUI Automation via Action-Effect Verification and Self-CorrectionApril 7, 2026 · ACL 2026
- Generalization in Online Reinforcement Learning for Mobile AgentsMarch 8, 2026 · arXiv
- AgentCPM‑GUI: Building Mobile‑Use Agents with Reinforcement Fine‑TuningJune 2, 2025 · EMNLP 2025 System Demonstrations
- The Blind Spot of Agent Safety: How Benign User Instructions Expose Critical Vulnerabilities in Computer-Use AgentsApril 12, 2026 · arXiv
- Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal AttacksMarch 4, 2026 · arXiv
- WiP: Characterizing and Defending Against Mobile-Agent-Driven MFA AutomationSeptember 2, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android AppsAugust 18, 2026 · arXiv
- MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent ResearchMay 25, 2026 · arXiv
- CORA: Conformal Risk-Controlled Agents for Safeguarded Mobile GUI AutomationApril 10, 2026 · arXiv
- UI-Voyager: A Self-Evolving GUI Agent Learning via Failed ExperienceMarch 25, 2026 · arXiv
- Adaptive Milestone Reward for GUI AgentsFebruary 12, 2026 · arXiv
- STEP: Success-Rate-Aware Trajectory-Efficient Policy OptimizationNovember 17, 2025 · Findings of ACL 2026
- Mobile GUI Agents under Real-world Threats: Are We There Yet?July 6, 2025 · MobiSys 2026
- GUI-R1: A Generalist R1-Style Vision-Language Action Model for GUI AgentsApril 14, 2025 · arXiv
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningMarch 27, 2025 · arXiv
- MobileSafetyBench: Evaluating Safety of Autonomous Agents in Mobile Device ControlOctober 23, 2024 · arXiv
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous AgentsMay 23, 2024 · ICLR 2025 (Poster)
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use AgentsSeptember 6, 2026 · arXiv
- Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection GuardrailsSeptember 2, 2026 · arXiv
- SIR: Self-improving Red-teaming for Compute Use AgentsAugust 31, 2026 · arXiv
- WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM AgentsAugust 25, 2026 · arXiv
- Beyond Success and Failure: Length-Aware Contrastive Learning for GUI AgentsAugust 22, 2026 · arXiv
- StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt InjectionAugust 6, 2026 · arXiv