You Only Look at Screens: Multimodal Chain-of-Action Agents
- 🏛 Institutions
- SJTU , Meta
- 📅 Date
- September 20, 2023
- 📑 Publisher
- Findings of ACL 2024
- 💻 Env
- Mobile
- 🔑 Keywords
TLDR
Auto-GUI is a screenshot-only mobile GUI agent that avoids environment parsing and application-specific APIs. The paper introduces a chain-of-action prompting technique and evaluates the method on AITW, a device-control benchmark with 30K unique instructions.
Related papers (24)
- MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent ResearchMay 25, 2026 · arXiv
- ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental SynthesisMay 24, 2026 · arXiv
- GUITester: Enabling GUI Agents for Exploratory Defect DiscoveryJanuary 8, 2026 · arXiv
- GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI AgentMay 22, 2025 · ACL 2025
- LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration BenchmarkApril 18, 2025 · arXiv
- ClickAgent: Enhancing UI Location Capabilities of Autonomous AgentsOctober 9, 2024 · SIGDIAL 2025
- AutoDroid: LLM-powered Task Automation in AndroidAugust 29, 2023 · MobiCom 2024
- Android in the Wild: A Large-Scale Dataset for Android Device ControlJuly 19, 2023 · NeurIPS 2023 Datasets and Benchmarks Track
- LongHorizonUI: A Unified Framework for Robust long-horizon Task Automation of GUI AgentJanuary 26, 2026 · ICLR 2026 (Poster)
- WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting PointFebruary 12, 2025 · arXiv
- WebWalker: Benchmarking LLMs in Web TraversalJanuary 13, 2025 · arXiv
- The BrowserGym Ecosystem for Web Agent ResearchDecember 6, 2024 · TMLR
- SheetCopilot: Bringing Software Productivity to the Next Level through Large Language ModelsMay 30, 2023 · NeurIPS 2023
- Grounding Open-Domain Instructions to Automate Web Support TasksMarch 30, 2021 · NAACL 2021
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionSeptember 9, 2026 · arXiv
- APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI AgentsSeptember 7, 2026 · arXiv
- Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool ActionsSeptember 6, 2026 · arXiv
- ElderBench: Benchmarking Autonomous Mobile Agents for Older AdultsSeptember 4, 2026 · arXiv
- GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent EnvironmentsAugust 30, 2026 · arXiv
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial EnvironmentsAugust 25, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- Benchmarking General Mobile Assistants in Challenging Real-World ScenariosAugust 21, 2026 · arXiv
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward ModelsJuly 30, 2026 · arXiv
- AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure InteractionJune 22, 2026 · arXiv