Tree Search for Language Model Agents
Jing Yu Koh , Stephen McAleer , Daniel Fried , Ruslan Salakhutdinov
- 🏛 Institutions
- CMU
- 📅 Date
- July 1, 2024
- 📑 Publisher
- TMLR 2025
- 💻 Env
- Web
- 🔑 Keywords
TLDR
This paper adds inference-time best-first tree search to language-model web agents by searching directly in the environment and guiding expansion with a model-based value function. On top of a GPT-4o baseline it reports a 39.7% relative gain on VisualWebArena and a 28.0% relative gain on WebArena, showing that web-agent performance scales with additional test-time search.
Related papers (24)
- Agent Alpha: Tree Search Unifying Generation, Exploration and Evaluation for Computer-Use AgentsFebruary 3, 2026 · arXiv
- Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web AgentsAugust 6, 2026 · arXiv
- WebOperator: Action-Aware Tree Search for Autonomous Agents in Web EnvironmentDecember 14, 2025 · arXiv
- WALT: Web Agents that Learn ToolsOctober 1, 2025 · ICLR 2026 (Poster)
- LiteWebAgent: The Open-Source Suite for VLM-Based Web-Agent ApplicationsMarch 4, 2025 · NAACL 2025 System Demonstrations
- Attacking Vision-Language Computer Agents via Pop-upsNovember 4, 2024 · ACL 2025
- ExACT: Teaching AI Agents to Explore with Reflective-MCTS and Exploratory LearningOctober 2, 2024 · ICLR 2025 (Poster)
- Dissecting Adversarial Robustness of Multimodal LM AgentsJune 18, 2024 · ICLR 2025 (Poster)
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web TasksJanuary 24, 2024 · ACL 2024
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgetsSeptember 11, 2026 · arXiv
- Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step SupervisionSeptember 2, 2026 · arXiv
- Discriminative World Models for Web AgentsSeptember 2, 2026 · arXiv
- Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection GuardrailsSeptember 2, 2026 · arXiv
- When and What to Teach: Budget-Aware Online Adaptation for Web AgentsAugust 31, 2026 · arXiv
- SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill AbstractionAugust 31, 2026 · arXiv
- Learning Simple Test-Time Environments for LLM Web AgentsAugust 29, 2026 · arXiv
- WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM AgentsAugust 25, 2026 · arXiv
- BrowserForge: Scaling Web Episode via Parallel Browser SandboxesAugust 25, 2026 · arXiv
- Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent LearningAugust 22, 2026 · arXiv
- Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and WebAugust 22, 2026 · arXiv
- ComponentBench: Diagnosing Component-Level Failures in Computer-Use AgentsAugust 18, 2026 · COLM 2026
- StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt InjectionAugust 6, 2026 · arXiv
- LoginTrap: Uncovering Task-Agnostic Phishing-Style Indirect Prompt Injection Attacks against LLM-based Web AgentsAugust 5, 2026 · arXiv
- Qwen-CUA: Native Computer Use for (almost) EverythingAugust 3, 2026 · arXiv