Seeing is Believing: Vision-driven Non-crash Functional Bug Detection for Mobile Apps
Zhe Liu , Cheng Li , Chunyang Chen , Junjie Wang , Mengzhuo Chen , Boyu Wu , Yawen Wang , Jun Hu , Qing Wang
- 🏛 Institutions
- Institute of Software , CAS , University of Chinese Academy of Sciences , TUM
- 📅 Date
- July 3, 2024
- 📑 Publisher
- arXiv
- 💻 Env
- Mobile
- 🔑 Keywords
TLDR
This paper introduces Trident, a vision-driven mobile GUI testing system with Explorer, Monitor, and Detector agents for finding non-crash functional bugs from screenshot sequences and transition logic. It evaluates on 590 non-crash bugs, reports large recall and precision gains over 12 baselines, and finds 43 new Google Play bugs, 31 of which were fixed.
Related papers (24)
- Towards Automated Crowdsourced Testing via Personified-LLMMarch 25, 2026 · FSE 2026
- GUITester: Enabling GUI Agents for Exploratory Defect DiscoveryJanuary 8, 2026 · arXiv
- GUIPilot: A Consistency-Based Mobile GUI Testing Approach for Detecting Application-Specific BugsJune 9, 2025 · ISSTA 2025
- AUITestAgent: Automatic Requirements Oriented GUI Function TestingJuly 12, 2024 · arXiv
- MobileExperts: A Dynamic Tool-Enabled Agent Team in Mobile DevicesJuly 4, 2024 · arXiv
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent CollaborationJune 3, 2024 · NeurIPS 2024
- World-Model-Augmented Web Agents with Action CorrectionFebruary 17, 2026 · arXiv
- BlueLM-GUI Technical Report: A Real-Device-Centric Flywheel for Self-Improving Mobile GUI AgentsSeptember 11, 2026 · arXiv
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionSeptember 9, 2026 · arXiv
- APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI AgentsSeptember 7, 2026 · arXiv
- Improving Proficiency and Efficiency of Android GUI Agents via Self-Generating Tool ActionsSeptember 6, 2026 · arXiv
- ElderBench: Benchmarking Autonomous Mobile Agents for Older AdultsSeptember 4, 2026 · arXiv
- WiP: Characterizing and Defending Against Mobile-Agent-Driven MFA AutomationSeptember 2, 2026 · arXiv
- GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent EnvironmentsAugust 30, 2026 · arXiv
- ActReal: System-Level Mobile Agents Challenge Mobile Automation DetectionAugust 30, 2026 · arXiv
- WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement LearningAugust 27, 2026 · arXiv
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial EnvironmentsAugust 25, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data SynthesisAugust 24, 2026 · arXiv
- Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and WebAugust 22, 2026 · arXiv
- Benchmarking General Mobile Assistants in Challenging Real-World ScenariosAugust 21, 2026 · arXiv
- Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and AggregationAugust 21, 2026 · arXiv
- MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android AppsAugust 18, 2026 · arXiv
- CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI ApplicationsAugust 12, 2026 · arXiv