GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
Zichuan Fu , Shirong Wang , Wenlin Zhang , Guojing Li , Yimin Deng , Jingtong Gao , Junjia Qi , Hanyu Yan , Yefeng Zheng , Xiaopeng Li , Wanyu Wang , Xian Wu , Xiangyu Zhao
- 🏛 Institutions
- CityU , Tencent Jarvis Lab , Westlake University
- 📅 Date
- August 4, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- General GUI
- 🔑 Keywords
TLDR
GUI-Lens turns GUI grounding into an iterative visual-search process: OCR and detected components provide coordinate references, a general-purpose VLM selects progressively enlarged crops, and separate verification gates reject bad crop or click proposals. Across four grounding benchmarks and three VLM backends, the paper reports gains of up to 24.9 percentage points, with GPT-5.5 reaching 87.9% on ScreenSpot-Pro.
Related papers (24)
- GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI GroundingAugust 22, 2026 · arXiv
- Hallucination-Free GUI Grounding via Regression-Free Layout-Aware MatchingAugust 10, 2026 · arXiv
- GUI-C²: Coarse-to-Fine GUI Grounding via Difficulty-Aware Reinforcement LearningMay 29, 2026 · arXiv
- UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI GroundingApril 15, 2026 · arXiv
- GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding ModelsApril 15, 2026 · arXiv
- See, Point, Refine: Multi-Turn Approach to GUI Grounding with Visual FeedbackApril 14, 2026 · arXiv
- What's Missing in Screen-to-Action? Towards a UI-in-the-Loop Paradigm for Multimodal GUI ReasoningApril 8, 2026 · Findings of ACL 2026
- Towards GUI Agents: Vision-Language Diffusion Models for GUI GroundingMarch 27, 2026 · CVPR 2026
- AdaZoom-GUI: Adaptive Zoom-based GUI Grounding with Instruction RefinementMarch 18, 2026 · arXiv
- Zoom to Essence: Trainless GUI Grounding by Inferring upon Interface ElementsMarch 15, 2026 · arXiv
- Moving Beyond Sparse Grounding with Complete Screen Parsing SupervisionFebruary 15, 2026 · arXiv
- Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal FusionFebruary 6, 2026 · arXiv
- POINTS-GUI-G: GUI-Grounding JourneyFebruary 6, 2026 · arXiv
- SSL: Sweet Spot Learning for Differentiated Guidance in Agentic OptimizationJanuary 30, 2026 · arXiv
- V2P: Visual Attention Calibration for GUI Grounding via Background Suppression and Center PeakingJanuary 11, 2026 · arXiv
- MVP: Multiple View Prediction Improves GUI GroundingDecember 9, 2025 · arXiv
- Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI GroundingDecember 5, 2025 · arXiv
- Beyond Clicking: A Step Towards Generalist GUI Grounding via Text DraggingNovember 7, 2025 · arXiv
- GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual GroundingOctober 5, 2025 · arXiv
- UI-AGILE: Advancing GUI Agents with Effective Reinforcement Learning and Precise Inference-Time GroundingJuly 29, 2025 · CVPR 2026 Findings
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI AgentsMay 21, 2025 · NeurIPS 2025 (Poster)
- Scaling Computer‑Use Grounding via User Interface Decomposition and SynthesisMay 19, 2025 · NeurIPS 2025 Datasets and Benchmarks Track (Spotlight)
- Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement LearningMay 18, 2025 · NeurIPS 2025 (Poster)
- Visual Test-time Scaling for GUI Agent GroundingMay 1, 2025 · ICCV 2025