Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents
Siqi Fan , Minghao Li , Xiaoqian Ma , Wenhui Tan , Xiusheng Huang , Juntong Wu , Liujie Zhang , Shuo Shang , Weihang Chen
- 🏛 Institutions
- Unknown
- 📅 Date
- August 4, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- Desktop
- 🔑 Keywords
TLDR
This work examines hybrid computer-use agents that can choose screenshots or text tools on OSWorld-MCP. It identifies an adoption gap in tool use and shows that training with a compatible post-tool observation policy can reduce image-context cost while improving the reported operating point.
Related papers (24)
- EE-MCP: Self-Evolving MCP-GUI Agents via Automated Environment Generation and Experience LearningApril 10, 2026 · arXiv
- Mobile-Agent-v3.5: Multi-platform Fundamental GUI AgentsFebruary 15, 2026 · arXiv
- MAI-UI Technical Report: Real-World Centric Foundation GUI AgentsDecember 26, 2025 · arXiv
- The Amazing Agent Race: Strong Tool Users, Weak NavigatorsApril 11, 2026 · arXiv
- The Tool Illusion: Rethinking Tool Use in Web AgentsApril 3, 2026 · arXiv
- LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial ScenariosFebruary 3, 2026 · arXiv
- ToolTok: Tool Tokenization for Efficient and Generalizable GUI AgentsJanuary 30, 2026 · arXiv
- MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented EnvironmentsDecember 22, 2025 · arXiv
- A Multimodal GUI Architecture for Interfacing with LLM-Based Conversational AssistantsAugust 31, 2025 · arXiv
- TurkingBench: A Challenge Benchmark for Web AgentsMarch 18, 2024 · NAACL 2025 (Oral)
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionSeptember 9, 2026 · arXiv
- FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?September 7, 2026 · arXiv
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use AgentsSeptember 6, 2026 · arXiv
- From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use AgentsSeptember 4, 2026 · arXiv
- CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI AgentsSeptember 4, 2026 · arXiv
- OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human DemonstrationsSeptember 2, 2026 · arXiv
- Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop ApplicationsSeptember 1, 2026 · arXiv
- SIR: Self-improving Red-teaming for Compute Use AgentsAugust 31, 2026 · arXiv
- CURA: Certified Runtime Alarms for Computer-Use AgentsAugust 28, 2026 · arXiv
- ASIL: Replacing Screenshot-and-Click with Structured State and Semantic ActionsAugust 27, 2026 · arXiv
- LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI AgentsAugust 26, 2026 · arXiv
- Reflection with Action-Induced Visual Differences for Desktop GUI AgentsAugust 25, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model TrajectoriesAugust 23, 2026 · arXiv