GUI Agents Papers
Star · 902

Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI Navigation

Heng Qu , Yike Liu , Renren Jin , Wenzong Zhang , Pengzhi Gao , Wei Liu , Jian Luan

🏛 Institutions
MiLM Plus , Xiaomi
📅 Date
May 26, 2026
📑 Publisher
ICML 2026 (Poster)
💻 Env
Mobile
🔑 Keywords
TLDR

This study examines data scaling, benchmarking, and reasoning for VLM-based mobile GUI agents. It introduces HyperTrack, a dataset of more than 16,000 tasks across over 650 Chinese mobile apps, and GUIEvalKit for unified offline evaluation, finding that reinforcement-based fine-tuning is especially effective for out-of-domain generalization.

Open paper OpenReview Code Report issue
Related papers (24)