GUI-PRA: Process Reward Agent for GUI Tasks
Tao Xiong , Xavier Hu , Yurun Chen , Yuhang Liu , Changqiao Wu , Pengzhi Gao , Wei Liu , Jian Luan , Shengyu Zhang
- 🏛 Institutions
- Zhejiang University , MiLM Plus , Xiaomi
- 📅 Date
- September 27, 2025
- 📑 Publisher
- arXiv
- 💻 Env
- Mobile
- 🔑 Keywords
TLDR
GUI-PRA turns GUI process-reward evaluation from passive scoring into active investigation. It synthesizes state-specific verification criteria from experience and uses them to navigate visual tools and gather grounded evidence, improving success over standard process reward models on AndroidWorld and Mobile-MiniWoB++.
Related papers (24)
- GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data SynthesisAugust 24, 2026 · arXiv
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward ModelsJuly 30, 2026 · arXiv
- ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental SynthesisMay 24, 2026 · arXiv
- ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI AgentsApril 13, 2026 · arXiv
- Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple ActionsApril 8, 2026 · arXiv
- AndroTMem: From Interaction Trajectories to Anchored Memory in Long-Horizon GUI AgentsMarch 19, 2026 · arXiv
- Video-Based Reward Modeling for Computer-Use AgentsMarch 10, 2026 · arXiv
- MobileBench-OL: A Comprehensive Chinese Benchmark for Evaluating Mobile GUI Agents in Real-World EnvironmentJanuary 28, 2026 · Findings of ACL 2026
- MobileWorld: Benchmarking Autonomous Mobile Agents in Agent-User Interactive and MCP-Augmented EnvironmentsDecember 22, 2025 · arXiv
- UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI AgentsMay 27, 2025 · NeurIPS 2025 (Poster)
- A3: Android Agent Arena for Mobile GUI Agents with Essential-State Procedural EvaluationJanuary 2, 2025 · arXiv
- Task-Adaptive Rubrics for GUI Reward ModelingAugust 25, 2026 · arXiv
- CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI AgentsAugust 23, 2026 · arXiv
- GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMsAugust 4, 2026 · arXiv
- SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use AgentsJuly 25, 2026 · arXiv
- StructAgent: Harness Long-horizon Digital Agents with Unified Causal StructureJuly 13, 2026 · arXiv
- Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional FieldsJune 9, 2026 · arXiv
- MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI AgentsMay 18, 2026 · arXiv
- CocoaBench: Evaluating Unified Digital Agents in the WildApril 13, 2026 · arXiv
- HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration TasksApril 10, 2026 · arXiv
- ClawBench: Can AI Agents Complete Everyday Online Tasks?April 9, 2026 · arXiv
- Gym-Anything: Turn any Software into an Agent EnvironmentApril 7, 2026 · arXiv
- IntentScore: Intent-Conditioned Action Evaluation for Computer-Use AgentsApril 6, 2026 · arXiv
- The Art of Building Verifiers for Computer Use AgentsApril 5, 2026 · arXiv