GUIPilot: A Consistency-Based Mobile GUI Testing Approach for Detecting Application-Specific Bugs
Ruofan Liu , Xiwen Teoh , Yun Lin , Guanjie Chen , Ruofei Ren , Denys Poshyvanyk , Jin Song Dong
- 🏛 Institutions
- SJTU , NUS , William & Mary
- 📅 Date
- June 9, 2025
- 📑 Publisher
- ISSTA 2025
- 💻 Env
- Mobile
- 🔑 Keywords
TLDR
GUIPilot detects application-specific mobile GUI bugs by checking both screen-level and process-level consistency between design mock-ups and the implemented app. It solves a widget alignment problem to find layout inconsistencies and uses a vision-language model with visual prompts to execute and validate specified GUI transitions, reaching 94.5% precision and 99.6% recall on 80 apps and 160 mock-ups and surfacing nine confirmed bugs in a trading app with 32 million users.
Related papers (24)
- Towards Automated Crowdsourced Testing via Personified-LLMMarch 25, 2026 · FSE 2026
- GUITester: Enabling GUI Agents for Exploratory Defect DiscoveryJanuary 8, 2026 · arXiv
- AUITestAgent: Automatic Requirements Oriented GUI Function TestingJuly 12, 2024 · arXiv
- Seeing is Believing: Vision-driven Non-crash Functional Bug Detection for Mobile AppsJuly 3, 2024 · arXiv
- CocoaBench: Evaluating Unified Digital Agents in the WildApril 13, 2026 · arXiv
- Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element InjectionApril 9, 2026 · arXiv
- ToolTok: Tool Tokenization for Efficient and Generalizable GUI AgentsJanuary 30, 2026 · arXiv
- How do Visual Attributes Influence Web Agents? A Comprehensive Evaluation of User Interface Design FactorsJanuary 29, 2026 · arXiv
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive PerceptionDecember 26, 2025 · arXiv
- Iris: Breaking GUI Complexity with Adaptive Focus and Self-RefiningDecember 13, 2024 · arXiv
- Visual Grounding for User InterfacesJune 16, 2024 · NAACL 2024 Industry Track
- Dual-View Visual Contextualization for Web NavigationFebruary 6, 2024 · CVPR 2024 (Poster)
- BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI AgentsSeptember 11, 2026 · arXiv
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionSeptember 9, 2026 · arXiv
- APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI AgentsSeptember 7, 2026 · arXiv
- Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool ActionsSeptember 6, 2026 · arXiv
- ElderBench: Benchmarking Autonomous Mobile Agents for Older AdultsSeptember 4, 2026 · arXiv
- WiP: Characterizing and Defending Against Mobile-Agent-Driven MFA AutomationSeptember 2, 2026 · arXiv
- GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent EnvironmentsAugust 30, 2026 · arXiv
- ActReal: System-Level Mobile Agents Challenge Mobile Automation DetectionAugust 30, 2026 · arXiv
- WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement LearningAugust 27, 2026 · arXiv
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial EnvironmentsAugust 25, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data SynthesisAugust 24, 2026 · arXiv