MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning
Haozhan Shen , Shilin Yan , Hongwei Xue , Shuaiqi Lu , Xiaojun Tang , Guannan Zhang , Tiancheng Zhao , Jianwei Yin
- 🏛 Institutions
- Accio Team , Alibaba Group , Zhejiang University , ZJU-BJ
- 📅 Date
- March 12, 2026
- 📑 Publisher
- arXiv
- 💻 Env
- 🔑 Keywords
TLDR
MM-CondChain is a benchmark for visually grounded deep compositional reasoning built from multi-layer conditional chains whose steps are programmatically verified through VPIR. It spans natural images, charts, and GUI trajectories, and shows that even the strongest MLLMs remain weak on deep chained reasoning.
Related papers (24)
- CocoaBench: Evaluating Unified Digital Agents in the WildApril 13, 2026 · arXiv
- VRL-Bench: Benchmarking agents on computer control tasks under finite trial budgetsSeptember 11, 2026 · arXiv
- JarvisGUI: Towards Cross-Device GUI Agents with Dynamic Task CompositionSeptember 9, 2026 · arXiv
- FinCUABuild: Can Agents Build Reliable Benchmarks for Dynamic Financial Computer Use?September 7, 2026 · arXiv
- APPSim-Bench: Bridging Real-world Apps and Reproducible Evaluation for Mobile GUI AgentsSeptember 7, 2026 · arXiv
- ElderBench: Benchmarking Autonomous Mobile Agents for Older AdultsSeptember 4, 2026 · arXiv
- GUI-CC: Benchmarking Contextual Consistency of GUI World Models as Agent EnvironmentsAugust 30, 2026 · arXiv
- Are Android GUI Agents Robust Against Runtime Anomalies? AnTrap: Evaluating Agents in Dynamic Adversarial EnvironmentsAugust 25, 2026 · arXiv
- ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across DevicesAugust 25, 2026 · arXiv
- Benchmarking General Mobile Assistants in Challenging Real-World ScenariosAugust 21, 2026 · arXiv
- ComponentBench: Diagnosing Component-Level Failures in Computer-Use AgentsAugust 18, 2026 · COLM 2026
- StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt InjectionAugust 6, 2026 · arXiv
- CUADebug: Diagnosing and Repairing Computer-Use Agent FailuresJuly 31, 2026 · arXiv
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward ModelsJuly 30, 2026 · arXiv
- OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic GroundingJuly 29, 2026 · arXiv
- EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition UnderstandingJuly 19, 2026 · arXiv
- WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent EvaluationJuly 7, 2026 · arXiv
- Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional FieldsJune 9, 2026 · arXiv
- Benchmarking Living-Screen-Native GUI Agents on Short-Video PlatformsJune 3, 2026 · arXiv
- Scaling, Benchmarking, and Reasoning of Vision-Language Agents for Mobile GUI NavigationMay 26, 2026 · ICML 2026 (Poster)
- AndroidDaily: A Verifiable Benchmark for Mobile GUI Agents on Real-World Closed-Source ApplicationsMay 26, 2026 · arXiv
- MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent ResearchMay 25, 2026 · arXiv
- ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental SynthesisMay 24, 2026 · arXiv
- WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application EnvironmentsApril 30, 2026 · arXiv