Succeed or Learn Slowly: Sample Efficient Off-Policy Reinforcement Learning for Mobile App Control
Georgios Papoudakis , Thomas Coste , Jianye Hao , Jun Wang , Kun Shao
- 🏛 Institutions
- Huawei Noah’s Ark Lab , UCL
- 📅 Date
- September 1, 2025
- 📑 Publisher
- NeurIPS 2025 (Poster)
- 💻 Env
- Mobile
- 🔑 Keywords
TLDR
SoLS is an off-policy RL algorithm for mobile app control that updates directly on successful samples but applies conservative regularized updates on negative ones to avoid policy degradation in sparse-reward settings. With Successful Transition Replay, it improves AndroidWorld performance substantially while using far less compute than GPT-4o-based baselines.
Related papers (24)
- BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI AgentsSeptember 11, 2026 · arXiv
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial EnvironmentsAugust 25, 2026 · arXiv
- SE-GA: Memory-Augmented Self-Evolution for GUI AgentsMay 16, 2026 · ICML 2026
- UI-Voyager: A Self-Evolving GUI Agent Learning via Failed ExperienceMarch 25, 2026 · arXiv
- HATS: Hardness-Aware Trajectory Synthesis for GUI AgentsMarch 12, 2026 · CVPR 2026
- Adaptive Milestone Reward for GUI AgentsFebruary 12, 2026 · arXiv
- STEP: Success-Rate-Aware Trajectory-Efficient Policy OptimizationNovember 17, 2025 · Findings of ACL 2026
- MobileRL: Online Agentic Reinforcement Learning for Mobile GUI AgentsSeptember 10, 2025 · ICLR 2026 (Poster)
- AppVLM: A Lightweight Vision Language Model for Online App ControlFebruary 10, 2025 · arXiv
- AndroidWorld: A Dynamic Benchmarking Environment for Autonomous AgentsMay 23, 2024 · ICLR 2025 (Poster)
- Agentic Reward Modeling: Verifying GUI Agent via Online Proactive InteractionJanuary 31, 2026 · arXiv
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionSeptember 9, 2026 · arXiv
- APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI AgentsSeptember 7, 2026 · arXiv
- Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool ActionsSeptember 6, 2026 · arXiv
- ElderBench: Benchmarking Autonomous Mobile Agents for Older AdultsSeptember 4, 2026 · arXiv
- WiP: Characterizing and Defending Against Mobile-Agent-Driven MFA AutomationSeptember 2, 2026 · arXiv
- GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent EnvironmentsAugust 30, 2026 · arXiv
- ActReal: System-Level Mobile Agents Challenge Mobile Automation DetectionAugust 30, 2026 · arXiv
- WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement LearningAugust 27, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data SynthesisAugust 24, 2026 · arXiv
- Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and WebAugust 22, 2026 · arXiv
- Benchmarking General Mobile Assistants in Challenging Real-World ScenariosAugust 21, 2026 · arXiv
- Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and AggregationAugust 21, 2026 · arXiv