Task-Adaptive Rubrics for GUI Reward Modeling
Tao Xiong , Xavier Hu , Wenkai Wang , Qinzhuo Wu , Changqiao Wu , Pengzhi Gao , Wei Liu , Jian Luan , Shengyu Zhang
- 🏛 Institutions
- Zhejiang University , MiLM Plus , Xiaomi
- 📅 Date
- August 25, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- General GUI
- 🔑 Keywords
TLDR
AdaptRubric makes GUI reward verification task-adaptive by routing each instruction to a task family for coarse rubric retrieval and then generating instance-level checks for concrete values, scopes, and constraints. It improves offline reward F1 by 3.6 points over the matched baseline average and raises downstream task success by 4.23 points in online reinforcement learning.
Related papers (24)
- OS-Themis: A Scalable Critic Framework for Generalist GUI RewardsMarch 19, 2026 · arXiv
- ProgRM: Build Better GUI Agents with Progress RewardsMay 23, 2025 · arXiv
- GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data SynthesisAugust 24, 2026 · arXiv
- SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use AgentsJuly 25, 2026 · arXiv
- ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI AgentsApril 13, 2026 · arXiv
- Android Coach: Improve Online Agentic Training Efficiency with Single State Multiple ActionsApril 8, 2026 · arXiv
- UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI AgentsMay 27, 2025 · NeurIPS 2025 (Poster)
- Beyond Success and Failure: Length-Aware Contrastive Learning for GUI AgentsAugust 22, 2026 · arXiv
- GUI-C²: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement LearningMay 29, 2026 · arXiv
- LiteGUI: Distilling Compact GUI Agents with Reinforcement LearningMay 8, 2026 · arXiv
- CGL: Advancing Continual GUI Learning via Reinforcement Fine-TuningMarch 3, 2026 · arXiv
- GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RLFebruary 25, 2026 · arXiv
- Building Autonomous GUI Navigation via Agentic-Q Estimation and Step-Wise Policy OptimizationFebruary 14, 2026 · arXiv
- Autonomous Continual Learning of Computer-Use Agents for Environment AdaptationFebruary 10, 2026 · arXiv
- SSL: Sweet Spot Learning for Differentiated Guidance in Agentic OptimizationJanuary 30, 2026 · arXiv
- MagicGUI-RMS: A Multi-Agent Reward Model System for Self-Evolving GUI Agents via Automated Feedback RefluxJanuary 19, 2026 · arXiv
- GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI AgentsJanuary 14, 2026 · arXiv
- From Off-Policy to On-Policy: Enhancing GUI Agents via Bi-level Expert-to-Policy AssimilationJanuary 9, 2026 · arXiv
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement LearningDecember 2, 2025 · arXiv
- HiconAgent: History Context-aware Policy Optimization for GUI AgentsDecember 1, 2025 · arXiv
- Training High-Level Schedulers with Execution-Feedback Reinforcement Learning for Long-Horizon GUI AutomationNovember 27, 2025 · CVPR 2026
- UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time GroundingJuly 29, 2025 · CVPR 2026 Findings
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI AgentsMay 21, 2025 · NeurIPS 2025 (Poster)
- Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement LearningMay 18, 2025 · NeurIPS 2025 (Poster)