CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications
Linqiang Guo , Li Gu , Zihuan Jiang , Zhixiang Chi , Siobhan Reid , Ziqiang Wang , Yuanhao Yu , Wei Liu , Yang Wang , Tse-Hsun (Peter) Chen
- 🏛 Institutions
- Concordia University , Mila , University of Toronto , McMaster University
- 📅 Date
- August 12, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- Mobile
- 🔑 Keywords
Mobile GUI agents stay brittle on applications absent from source training, so CoAdapt-GUI performs test-time adaptation that jointly updates a structured workflow context and the policy from the agent's own target-app rollouts and rewards. The workflow context retains transferable procedures, failure modes, and verification rules while excluding app-bound source details, and policy adaptation uses a task-context-matched group-relative objective over a LoRA adapter on a frozen VLM. It reaches 45.0% on AndroidWorld-Generalization against 37.5% for the reported Policy-Only TTA baseline, and lifts AndroidWorld Plus from 38.6% to 52.9%.
- Generalization in Online Reinforcement Learning for Mobile AgentsMarch 8, 2026 · arXiv
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial EnvironmentsAugust 25, 2026 · arXiv
- Don't Act Blindly: Robust GUI Automation via Action-Effect Verification and Self-CorrectionApril 7, 2026 · ACL 2026
- OmegaUse: Building a General-Purpose GUI Agent for Autonomous Task ExecutionJanuary 28, 2026 · arXiv
- AgentCPM‑GUI: Building Mobile‑Use Agents with Reinforcement Fine‑TuningJune 2, 2025 · EMNLP 2025 System Demonstrations
- ZeroGUI: Automating Online GUI Learning at Zero Human CostMay 29, 2025 · arXiv
- GUI-R1: A Generalist R1-Style Vision-Language Action Model for GUI AgentsApril 14, 2025 · arXiv
- UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement LearningMarch 27, 2025 · arXiv
- Learning Simple Test-Time Environments for LLM Web AgentsAugust 29, 2026 · arXiv
- Beyond Success and Failure: Length-Aware Contrastive Learning for GUI AgentsAugust 22, 2026 · arXiv
- AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction RefinementMarch 18, 2026 · arXiv
- Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal AttacksMarch 4, 2026 · arXiv
- CGL: Advancing Continual GUI Learning via Reinforcement Fine-TuningMarch 3, 2026 · arXiv
- ARPO:End-to-End Policy Optimization for GUI Agents with Experience ReplayMay 22, 2025 · arXiv
- BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI AgentsSeptember 11, 2026 · arXiv
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionSeptember 9, 2026 · arXiv
- APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI AgentsSeptember 7, 2026 · arXiv
- Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool ActionsSeptember 6, 2026 · arXiv
- ElderBench: Benchmarking Autonomous Mobile Agents for Older AdultsSeptember 4, 2026 · arXiv
- WiP: Characterizing and Defending Against Mobile-Agent-Driven MFA AutomationSeptember 2, 2026 · arXiv
- GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent EnvironmentsAugust 30, 2026 · arXiv
- ActReal: System-Level Mobile Agents Challenge Mobile Automation DetectionAugust 30, 2026 · arXiv
- WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement LearningAugust 27, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv