OpenWebVoyager: Building Multimodal Web Agents via Iterative Real-World Exploration, Feedback and Optimization
Hongliang He , Wenlin Yao , Kaixin Ma , Wenhao Yu , Hongming Zhang , Tianqing Fang , Zhenzhong Lan , Dong Yu
- 🏛 Institutions
- ZJU , Tencent AI Lab (Seattle) , Westlake University
- 📅 Date
- October 25, 2024
- 📑 Publisher
- ACL 2025
- 💻 Env
- Web
- 🔑 Keywords
TLDR
OpenWebVoyager is a multimodal web agent that improves itself through repeated cycles of real-world exploration, feedback collection, and policy optimization. It starts from imitation learning, mines open-web trajectories, and shows stronger performance after each optimization round.
Related papers (24)
- SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill AbstractionAugust 31, 2026 · arXiv
- Thinking vs. Doing: Agents that Reason by Scaling Test-Time InteractionJune 9, 2025 · SEA @ NeurIPS 2025 (Oral)
- Agent-E: From Autonomous Web Navigation to Foundational Design Principles in Agentic SystemsJuly 17, 2024 · arXiv
- Large Language Models Can Self-Improve At Web Agent TasksMay 30, 2024 · arXiv
- Autonomous Evaluation and Refinement of Digital AgentsApril 9, 2024 · COLM 2024
- BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI AgentsSeptember 11, 2026 · arXiv
- OSExpert: Computer-Use Agents Learning Professional Skills via ExplorationMarch 9, 2026 · arXiv
- Compress to Focus: Efficient Coordinate Compression for Policy Optimization in Multi-Turn GUI AgentsJanuary 14, 2026 · arXiv
- UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI AgentsMay 27, 2025 · NeurIPS 2025 (Poster)
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgetsSeptember 11, 2026 · arXiv
- Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step SupervisionSeptember 2, 2026 · arXiv
- Discriminative World Models for Web AgentsSeptember 2, 2026 · arXiv
- Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection GuardrailsSeptember 2, 2026 · arXiv
- When and What to Teach: Budget-Aware Online Adaptation for Web AgentsAugust 31, 2026 · arXiv
- Learning Simple Test-Time Environments for LLM Web AgentsAugust 29, 2026 · arXiv
- WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM AgentsAugust 25, 2026 · arXiv
- BrowserForge: Scaling Web Episode via Parallel Browser SandboxesAugust 25, 2026 · arXiv
- Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent LearningAugust 22, 2026 · arXiv
- Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and WebAugust 22, 2026 · arXiv
- ComponentBench: Diagnosing Component-Level Failures in Computer-Use AgentsAugust 18, 2026 · COLM 2026
- StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt InjectionAugust 6, 2026 · arXiv
- Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web AgentsAugust 6, 2026 · arXiv
- LoginTrap: Uncovering Task-Agnostic Phishing-Style Indirect Prompt Injection Attacks against LLM-based Web AgentsAugust 5, 2026 · arXiv
- Qwen-CUA: Native Computer Use for (almost) EverythingAugust 3, 2026 · arXiv