Multimodal Reasoning
3개 벤치마크 · 논문 1,039편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation
Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation
WebQA: Multihop and Multimodal QA
Papers
Reason Through the Latent! Making Latent Visual Reasoning Necessary
Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply tha…
Multimodal ReasoningVisual ReasoningMobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control
Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-language-action (VLA) systems on mobile robots, due to the persistent gap between high-level semanti…
Reinforcement LearningInstruction FollowingMultimodal ReasoningDecision MakingLearning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step deduction. Existing free-form traces obscure the decisions that determine the answer, and trajectory-le…
Reinforcement LearningMultimodal ReasoningReactivating Test-Time Scaling for Plane Geometry Problem Solving
Plane geometry problem (PGP) solving has become a critical benchmark for multimodal reasoning because it requires accurate visual perception and precise multi-step symbolic deduction. Although test-time scaling (TTS) has…
Mathematical ReasoningMultimodal ReasoningVisual GroundingTAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film
Films communicate through deliberate creative choices, including lighting, color, composition, editing, dialogue, music, and sound. Humans naturally interpret these signals as directorial intent, yet current multimodal l…
Multimodal ReasoningCoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions
Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT fo…
Multimodal Reasoning