CoCo-Agent: A Comprehensive Cognitive MLLM Agent for Smartphone GUI Automation
Xinbei Ma , Zhuosheng Zhang , Hai Zhao
- 🏛 Institutions
- SJTU
- 📅 Date
- February 19, 2024
- 📑 Publisher
- Findings of ACL 2024
- 💻 Env
- Mobile
- 🔑 Keywords
TLDR
CoCo-Agent is a smartphone GUI agent built around comprehensive environment perception (CEP) and conditional action prediction (CAP). The paper reports state-of-the-art performance on AITW and META-GUI, arguing that richer multimodal environment modeling improves mobile action selection.
Related papers (24)
- ClickAgent: Enhancing UI Location Capabilities of Autonomous AgentsOctober 9, 2024 · SIGDIAL 2025
- DigiRL: Training In-The-Wild Device-Control Agents with Autonomous Reinforcement LearningJune 14, 2024 · NeurIPS 2024 Main Conference Track
- CogAgent: A Visual Language Model for GUI AgentsDecember 14, 2023 · CVPR 2024 (Highlight)
- You Only Look at Screens: Multimodal Chain-of-Action AgentsSeptember 20, 2023 · Findings of ACL 2024
- Android in the Wild: A Large-Scale Dataset for Android Device ControlJuly 19, 2023 · NeurIPS 2023 Datasets and Benchmarks Track
- META-GUI: Towards Multi-modal Conversational Agents on Mobile GUIMay 23, 2022 · EMNLP 2022
- BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI AgentsSeptember 11, 2026 · arXiv
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionSeptember 9, 2026 · arXiv
- APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI AgentsSeptember 7, 2026 · arXiv
- Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool ActionsSeptember 6, 2026 · arXiv
- ElderBench: Benchmarking Autonomous Mobile Agents for Older AdultsSeptember 4, 2026 · arXiv
- WiP: Characterizing and Defending Against Mobile-Agent-Driven MFA AutomationSeptember 2, 2026 · arXiv
- GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent EnvironmentsAugust 30, 2026 · arXiv
- ActReal: System-Level Mobile Agents Challenge Mobile Automation DetectionAugust 30, 2026 · arXiv
- WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement LearningAugust 27, 2026 · arXiv
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial EnvironmentsAugust 25, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data SynthesisAugust 24, 2026 · arXiv
- Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and WebAugust 22, 2026 · arXiv
- Benchmarking General Mobile Assistants in Challenging Real-World ScenariosAugust 21, 2026 · arXiv
- Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and AggregationAugust 21, 2026 · arXiv
- MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android AppsAugust 18, 2026 · arXiv
- CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI ApplicationsAugust 12, 2026 · arXiv
- The Next Screenshot Knows: Gated Hindsight Distillation for Mobile GUI AgentsAugust 6, 2026 · arXiv