WebGraphEval: Multi-Turn Trajectory Evaluation for Web Agents using Graph Representation
Yaoyao Qian , Yuanli Wang , Jinda Zhang , Yun Zong , Meixu Chen , Hanhan Zhou , Jindan Huang , Yifan Zeng , Xinyu Hu , Chan Hee Song , Danqing Zhang
- 🏛 Institutions
- Northeastern University , Boston University , University of Victoria , University of Minnesota , George Washington University , Tufts University , Oregon State University , University of Texas at San Antonio , OSU , PathOnAI.org
- 📅 Date
- October 22, 2025
- 📑 Publisher
- NeurIPS 2025 Workshop on Multi-Turn Interactions in Large Language Models
- 💻 Env
- Web
- 🔑 Keywords
TLDR
WebGraphEval evaluates web agents by converting many interaction trajectories into a unified weighted action graph instead of scoring only final success or conformity to one reference path. This graph view highlights redundancy, inefficiency, and critical decision points across agents and benchmark runs.
Related papers (24)
- Same Outcomes, Different Journeys: A Trace-Level Framework for Comparing Human and GUI-Agent Behavior in Production Search SystemsApril 9, 2026 · arXiv
- AI Planning Framework for LLM-Based Web AgentsMarch 13, 2026 · arXiv
- An Illusion of Progress? Assessing the Current State of Web AgentsApril 2, 2025 · COLM 2025
- GUIDE: Interpretable GUI Agent Evaluation via Hierarchical DiagnosisApril 6, 2026 · arXiv
- CUAAudit: Meta-Evaluation of Vision-Language Models as Auditors of Autonomous Computer-Use AgentsMarch 11, 2026 · HEAL @ CHI 2026 Workshop
- MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic EnvironmentsFebruary 3, 2026 · arXiv
- Modular and Multi-Path-Aware Offline Benchmarking for Mobile GUI AgentsDecember 14, 2025 · arXiv
- SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy TasksMay 17, 2025 · ICLR 2026 (Poster)
- GUI Agents: A SurveyDecember 18, 2024 · Findings of ACL 2025
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgetsSeptember 11, 2026 · arXiv
- Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step SupervisionSeptember 2, 2026 · arXiv
- Discriminative World Models for Web AgentsSeptember 2, 2026 · arXiv
- Beyond the Verdict: Evidence-Aligned Evaluation of Visual Prompt-Injection GuardrailsSeptember 2, 2026 · arXiv
- When and What to Teach: Budget-Aware Online Adaptation for Web AgentsAugust 31, 2026 · arXiv
- SCAFFOLD: Self-Improving Web Agents via Recursive Parametric Skill AbstractionAugust 31, 2026 · arXiv
- Learning Simple Test-Time Environments for LLM Web AgentsAugust 29, 2026 · arXiv
- WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM AgentsAugust 25, 2026 · arXiv
- BrowserForge: Scaling Web Episode via Parallel Browser SandboxesAugust 25, 2026 · arXiv
- Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent LearningAugust 22, 2026 · arXiv
- Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and WebAugust 22, 2026 · arXiv
- ComponentBench: Diagnosing Component-Level Failures in Computer-Use AgentsAugust 18, 2026 · COLM 2026
- StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt InjectionAugust 6, 2026 · arXiv
- Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web AgentsAugust 6, 2026 · arXiv
- LoginTrap: Uncovering Task-Agnostic Phishing-Style Indirect Prompt Injection Attacks against LLM-based Web AgentsAugust 5, 2026 · arXiv