StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure
Wenyi Wu , Sibo Zhu , Kun Zhou , Aayush Salvi , Zixuan Song , Biwei Huang
- 🏛 Institutions
- UC San Diego , Aether AI Lab
- 📅 Date
- July 13, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- Desktop Web
- 🔑 Keywords
TLDR
StructAgent organizes long-horizon computer use around a unified, evidence-backed representation of task progress shared by its planner, actor, and verifier. Verifier-gated state transitions support checkpointing, evidence-based completion, and targeted recovery across OSWorld-Verified and other digital environments.
Related papers (24)
- Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI AutomationNovember 27, 2025 · CVPR 2026
- StateAct: Program State, before Pixels, for Long-Horizon Computer-Use AgentsJuly 24, 2026 · arXiv
- Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional FieldsJune 9, 2026 · arXiv
- HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration TasksApril 10, 2026 · arXiv
- ClawBench: Can AI Agents Complete Everyday Online Tasks?April 9, 2026 · arXiv
- Gym-Anything: Turn any Software into an Agent EnvironmentApril 7, 2026 · arXiv
- The Art of Building Verifiers for Computer Use AgentsApril 5, 2026 · arXiv
- GPA: Learning GUI Process Automation from DemonstrationsApril 2, 2026 · arXiv
- ContractSkill: Repairable Contract-Based Skills for Multimodal Web AgentsMarch 20, 2026 · arXiv
- M^2: Dual-Memory Augmentation for Long-Horizon Web Agents via Trajectory Summarization and Insight RetrievalFebruary 28, 2026 · arXiv
- ANCHOR: Branch-Point Data Generation for GUI AgentsFebruary 6, 2026 · arXiv
- OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive TasksJanuary 28, 2026 · arXiv
- AssistantBench: Can Web Agents Solve Realistic and Time-Consuming Tasks?October 21, 2024 · EMNLP 2024 (Poster)
- CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI AgentsAugust 23, 2026 · arXiv
- Qwen-CUA: Native Computer Use for (almost) EverythingAugust 3, 2026 · arXiv
- MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action DistillationJuly 31, 2026 · arXiv
- Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI AgentsJuly 30, 2026 · arXiv
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward ModelsJuly 30, 2026 · arXiv
- SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RLJuly 13, 2026 · arXiv
- ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental SynthesisMay 24, 2026 · arXiv
- MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI AgentsMay 18, 2026 · arXiv
- Planner Matters! An Efficient and Unbalanced Multi-agent Collaboration Framework for Long-horizon PlanningMay 4, 2026 · arXiv
- CocoaBench: Evaluating Unified Digital Agents in the WildApril 13, 2026 · arXiv
- AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI AgentsMarch 19, 2026 · arXiv