CUA-Skill: Develop Skills for Computer Using Agent
Tianyi Chen , Yinheng Li , Michael Solodko , Sen Wang , Nan Jiang , Tingyuan Cui , Junheng Hao , Jongwoo Ko , Sara Abdali , Leon Xu , Suzhen Zheng , Hao Fan , Pashmina Cameron , Justin Wagle , Kazuhito Koishida
- 🏛 Institutions
- Microsoft
- 📅 Date
- January 28, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- Desktop
- 🔑 Keywords
TLDR
CUA-Skill builds a reusable skill base for computer-use agents by encoding human computer-use knowledge as parameterized skills plus execution and composition graphs. The resulting CUA-Skill Agent improves robustness and reaches strong performance on WindowsAgentArena through dynamic skill retrieval and memory-aware recovery.
Related papers (24)
- From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use AgentsSeptember 4, 2026 · arXiv
- WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application EnvironmentsApril 30, 2026 · arXiv
- IntentCUA: Learning Intent-level Representations for Skill Abstraction and Multi-Agent Planning in Computer-Use AgentsFebruary 19, 2026 · AAMAS 2026
- ANCHOR: Branch-Point Data Generation for GUI AgentsFebruary 6, 2026 · arXiv
- GUI-360: A Comprehensive Dataset and Benchmark for Computer-Using AgentsNovember 6, 2025 · arXiv
- The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer UseNovember 15, 2024 · arXiv
- ASSISTGUI: Task-Oriented Desktop Graphical User Interface AutomationDecember 20, 2023 · CVPR 2024 (Poster)
- SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill AbstractionAugust 31, 2026 · arXiv
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionSeptember 9, 2026 · arXiv
- FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?September 7, 2026 · arXiv
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use AgentsSeptember 6, 2026 · arXiv
- CUA-Universe: A Scalable and Dynamic Environment for Hybrid GUI+CLI AgentsSeptember 4, 2026 · arXiv
- OmegaUse-SOP: SOP Engineering for Professional Computer Use from Human DemonstrationsSeptember 2, 2026 · arXiv
- Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop ApplicationsSeptember 1, 2026 · arXiv
- SIR: Self-improving Red-teaming for Compute Use AgentsAugust 31, 2026 · arXiv
- CURA: Certified Runtime Alarms for Computer-Use AgentsAugust 28, 2026 · arXiv
- ASIL: Replacing Screenshot-and-Click with Structured State and Semantic ActionsAugust 27, 2026 · arXiv
- LocalLSTC: A Long Short-Term Control Architecture for Locally Deployed GUI AgentsAugust 26, 2026 · arXiv
- Reflection with Action-Induced Visual Differences for Desktop GUI AgentsAugust 25, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model TrajectoriesAugust 23, 2026 · arXiv
- Spine-Branch Coordination for Multi-agent Computer UseAugust 22, 2026 · arXiv
- Inducing Task Models from Computer-Use TracesAugust 20, 2026 · arXiv
- Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use AgentsAugust 4, 2026 · arXiv