TreeCUA: Efficiently Scaling GUI Automation with Tree-Structured Verifiable Evolution
Deyang Jiang , Jing Huang , Xuanle Zhao , Lei Chen , Liming Zheng , Fanfan Liu , Haibo Qiu , Peng Shi , Zhixiong Zeng
- 🏛 Institutions
- Meituan
- 📅 Date
- February 10, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- General GUI
- 🔑 Keywords
TLDR
TreeCUA tackles the scaling bottleneck in GUI planning by organizing exploration trajectories as reusable tree structures with verification, summarization, and evaluation. The resulting data supports TreeCUA-DPO, which improves planning quality and out-of-domain generalization.
Related papers (24)
- Demo2Tutorial: From Human Experience to Multimodal Software TutorialsJune 2, 2026 · CVPR 2026
- Executable Agentic Memory for GUI AgentMay 12, 2026 · arXiv
- SSL: Sweet Spot Learning for Differentiated Guidance in Agentic OptimizationJanuary 30, 2026 · arXiv
- TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI AgentsApril 17, 2025 · AAAI 2026
- Spine-Branch Coordination for Multi-agent Computer UseAugust 22, 2026 · arXiv
- MobileWorldBench: Towards Semantic World Modeling For Mobile AgentsDecember 16, 2025 · arXiv
- WebATLAS: An LLM Agent with Experience-Driven Memory and Action SimulationOctober 26, 2025 · NeurIPS 2025 Workshop on Language Agents and World Models
- Building a Stable Planner: An Extended Finite State Machine Based Planning Module for Mobile GUI AgentMay 20, 2025 · arXiv
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and ProspectsApril 28, 2025 · TMLR 2025
- UFO2: The Desktop AgentOSApril 20, 2025 · arXiv
- WebRollback: Enhancing Web Agents with Explicit Rollback MechanismsApril 16, 2025 · EACL 2026 (Oral)
- LiteWebAgent: The Open-Source Suite for VLM-Based Web-Agent ApplicationsMarch 4, 2025 · NAACL 2025 System Demonstrations
- WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work TasksJuly 7, 2024 · NeurIPS 2024 Datasets and Benchmarks Track (Poster)
- A Real-World WebAgent with Planning, Long Context Understanding, and Program SynthesisJuly 24, 2023 · ICLR 2024 (Oral)
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgetsSeptember 11, 2026 · arXiv
- TRACE: Trajectory-robust Admission with Evidence Ordering for Efficient GUI AgentsSeptember 9, 2026 · arXiv
- FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?September 7, 2026 · arXiv
- Selective Knowledge Control for Continual GUI Agent Learning over Application StreamsSeptember 6, 2026 · arXiv
- AgentHijack: Visual Patch Attacks on Multimodal Computer-Use AgentsSeptember 6, 2026 · arXiv
- From Interaction Traces to Persistent Skills: Online Evolution for Computer-Use AgentsSeptember 4, 2026 · arXiv
- Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI AgentsSeptember 3, 2026 · arXiv
- Efficient GUI Agents: A Systems Survey of Observation, Memory, Action, and Runtime OptimizationSeptember 2, 2026 · arXiv
- SIR: Self-improving Red-teaming for Compute Use AgentsAugust 31, 2026 · arXiv
- Iron: Intent-Aligned and Retrospective Dual Learning Framework for Enhancing Generalist Virtual AgentsAugust 28, 2026 · arXiv