Explain Before You Answer: A Survey on Compositional Visual Reasoning

  • 2025-08-27 08:55:54
  • Fucai Ke, Joy Hsu, Zhixi Cai, Zixian Ma, Xin Zheng, Xindi Wu, Sukai Huang, Weiqing Wang, Pari Delir Haghighi, Gholamreza Haffari, Ranjay Krishna, Jiajun Wu, Hamid Rezatofighi
  • 0

Abstract

Compositional visual reasoning has emerged as a key research frontier inmultimodal AI, aiming to endow machines with the human-like ability todecompose visual scenes, ground intermediate concepts, and perform multi-steplogical inference. While early surveys focus on monolithic vision-languagemodels or general multimodal reasoning, a dedicated synthesis of the rapidlyexpanding compositional visual reasoning literature is still missing. We fillthis gap with a comprehensive survey spanning 2023 to 2025 that systematicallyreviews 260+ papers from top venues (CVPR, ICCV, NeurIPS, ICML, ACL, etc.). Wefirst formalize core definitions and describe why compositional approachesoffer advantages in cognitive alignment, semantic fidelity, robustness,interpretability, and data efficiency. Next, we trace a five-stage paradigmshift: from prompt-enhanced language-centric pipelines, through tool-enhancedLLMs and tool-enhanced VLMs, to recently minted chain-of-thought reasoning andunified agentic VLMs, highlighting their architectural designs, strengths, andlimitations. We then catalog 60+ benchmarks and corresponding metrics thatprobe compositional visual reasoning along dimensions such as groundingaccuracy, chain-of-thought faithfulness, and high-resolution perception.Drawing on these analyses, we distill key insights, identify open challenges(e.g., limitations of LLM-based reasoning, hallucination, a bias towarddeductive reasoning, scalable supervision, tool integration, and benchmarklimitations), and outline future directions, including world-model integration,human-AI collaborative reasoning, and richer evaluation protocols. By offeringa unified taxonomy, historical roadmap, and critical outlook, this survey aimsto serve as a foundational reference and inspire the next generation ofcompositional visual reasoning research.

 

Quick Read (beta)

loading the full paper ...