FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
Qinglong Yang , Haoming Li , Haotian Zhao , Xiaokai Yan , Jingtao Ding , Fengli Xu , Yong Li
- 🏛 Institutions
- Tsinghua
- 📅 Date
- June 9, 2025
- 📑 Publisher
- ICLR 2026 (Poster)
- 💻 Env
- Mobile
- 🔑 Keywords
TLDR
FingerTip 20K is a mobile benchmark built from 20K real-life Android demonstrations collected over long-term usage rather than isolated tasks. It focuses on proactive task suggestion and personalized execution, and shows that current mobile agents make poor use of user context and preference information compared with humans.
Related papers (24)
- PSPA-Bench: A Personalized Benchmark for Smartphone GUI AgentMarch 31, 2026 · arXiv
- SecAgent: Efficient Mobile GUI Agent with Semantic ContextMarch 9, 2026 · arXiv
- Turing Test on Screen: A Benchmark for Mobile GUI Agent HumanizationFebruary 24, 2026 · arXiv
- AmbiBench: Benchmarking Mobile GUI Agents Beyond One-Shot Instructions in the WildFebruary 12, 2026 · arXiv
- MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic EnvironmentsFebruary 3, 2026 · arXiv
- SwipeGen: Bridging the Execution Gap in GUI Agents via Human-like Swipe SynthesisJanuary 26, 2026 · arXiv
- SMAN-Bench: A Cross-System Benchmark for Mobile Agents under Single- and Multi-path, Ambiguous, and Noisy TasksJanuary 26, 2026 · ICLR 2026 (Poster)
- PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric RecordsJanuary 14, 2026 · arXiv
- MobileWorldBench: Towards Semantic World Modeling For Mobile AgentsDecember 16, 2025 · arXiv
- NaturalGAIA: Pushing the Frontiers of GUI Agents with a Challenging Benchmark and High-Quality Trajectory DatasetAugust 2, 2025 · arXiv
- LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration BenchmarkApril 18, 2025 · arXiv
- AndroidLab: Training and Systematic Benchmarking of Android Autonomous AgentsOctober 31, 2024 · ACL 2025
- GUI-World: A Video Benchmark and Dataset for Multimodal GUI-oriented UnderstandingJune 16, 2024 · ICLR 2025 (Poster)
- LlamaTouch: A Faithful and Scalable Testbed for Mobile UI Task AutomationApril 12, 2024 · UIST 2024
- Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMsApril 8, 2024 · ECCV 2024 (Poster)
- SeeClick: Harnessing GUI Grounding for Advanced Visual GUI AgentsJanuary 17, 2024 · ACL 2024
- GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI NavigationNovember 13, 2023 · arXiv
- Android in the Wild: A Large-Scale Dataset for Android Device ControlJuly 19, 2023 · NeurIPS 2023 Datasets and Benchmarks Track
- A Dataset for Interactive Vision-Language Navigation with Unknown Command FeasibilityFebruary 4, 2022 · ECCV 2022
- Screen2Words: Automatic Mobile UI Summarization with Multimodal LearningAugust 6, 2021 · UIST 2021
- Widget Captioning: Generating Natural Language Description for Mobile User Interface ElementsNovember 30, 2020 · EMNLP 2020
- Mapping Natural Language Instructions to Mobile UI Action SequencesJuly 31, 2020 · ACL 2020
- WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent BenchmarkApril 13, 2026 · arXiv
- Gym-Anything: Turn any Software into an Agent EnvironmentApril 7, 2026 · arXiv