Why Do LLM-based Web Agents Fail? A Hierarchical Planning Perspective
Mohamed Aghzal , Gregory J. Stein , Ziyu Yao
- 🏛 Institutions
- George Mason University
- 📅 Date
- March 15, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- Web
- 🔑 Keywords
TLDR
This paper analyzes web-agent failures through a three-layer hierarchy of high-level planning, low-level execution, and replanning rather than relying only on end-to-end success. It finds that structured PDDL plans improve strategic planning over natural-language plans, but that execution and grounding remain the dominant reliability bottlenecks.
Related papers (24)
- SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web AgentsOctober 11, 2025 · arXiv
- WebSuite: Systematically Evaluating Why Web Agents FailJune 1, 2024 · arXiv
- VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?April 9, 2024 · COLM 2024
- GPT-4V(ision) is a Generalist Web Agent, if GroundedJanuary 3, 2024 · ICML 2024
- Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy OptimizationFebruary 14, 2026 · arXiv
- Agent S: An Open Agentic Framework that Uses Computers Like a HumanOctober 10, 2024 · ICLR 2025 (Poster)
- Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMsApril 8, 2024 · ECCV 2024 (Poster)
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgetsSeptember 11, 2026 · arXiv
- Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step SupervisionSeptember 2, 2026 · arXiv
- Discriminative World Models for Web AgentsSeptember 2, 2026 · arXiv
- Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection GuardrailsSeptember 2, 2026 · arXiv
- When and What to Teach: Budget-Aware Online Adaptation for Web AgentsAugust 31, 2026 · arXiv
- SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill AbstractionAugust 31, 2026 · arXiv
- Learning Simple Test-Time Environments for LLM Web AgentsAugust 29, 2026 · arXiv
- WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM AgentsAugust 25, 2026 · arXiv
- BrowserForge: Scaling Web Episode via Parallel Browser SandboxesAugust 25, 2026 · arXiv
- Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent LearningAugust 22, 2026 · arXiv
- Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and WebAugust 22, 2026 · arXiv
- ComponentBench: Diagnosing Component-Level Failures in Computer-Use AgentsAugust 18, 2026 · COLM 2026
- StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt InjectionAugust 6, 2026 · arXiv
- Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web AgentsAugust 6, 2026 · arXiv
- LoginTrap: Uncovering Task-Agnostic Phishing-Style Indirect Prompt Injection Attacks against LLM-based Web AgentsAugust 5, 2026 · arXiv
- Qwen-CUA: Native Computer Use for (almost) EverythingAugust 3, 2026 · arXiv
- MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action DistillationJuly 31, 2026 · arXiv