paper-with-me

Multimodal Reasoning

3개 벤치마크 · 논문 1,039편 · 이 태스크의 논문 보기 →

Benchmarks

REBUS

결과 8개

MATH-V

결과 4개

AlgoPuzzleVQA

결과 1개

Most implemented

WebQA: Multihop and Multimodal QA

2021-09-01 · 구현 3개

Papers

Reason Through the Latent! Making Latent Visual Reasoning Necessary

2026-09-06 · Suhyeong Park, Junha Jung, Jaewoo Kang hf

Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply tha…

Multimodal ReasoningVisual Reasoning

MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control

2026-09-05 · Ting Huang, Yue Huang, Zeyu Zhang, Shuicheng Yan 외 hf

Grounding natural-language instructions into reliable and executable actions remains a fundamental challenge for vision-language-action (VLA) systems on mobile robots, due to the persistent gap between high-level semanti…

Reinforcement LearningInstruction FollowingMultimodal ReasoningDecision Making

Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry

2026-08-31 · Jiani Guo, Junjie Wang, Jie Wu, Pengxiang Zhao 외 arxiv

Multimodal geometry reasoning requires VLMs to extract precise visual relations and preserve them through multi-step deduction. Existing free-form traces obscure the decisions that determine the answer, and trajectory-le…

Reinforcement LearningMultimodal Reasoning

Reactivating Test-Time Scaling for Plane Geometry Problem Solving

2026-08-31 · Xiaoqiang Kang, Shengen Wu, Maizhen Ning, Xiaobo Jin 외 arxiv

Plane geometry problem (PGP) solving has become a critical benchmark for multimodal reasoning because it requires accurate visual perception and precise multi-step symbolic deduction. Although test-time scaling (TTS) has…

Mathematical ReasoningMultimodal ReasoningVisual Grounding

TAKE 85: Testing Audiovisual filmmaKer's intEnt across 85 Hours of Film

2026-08-30 · Kaishuu Shinozaki-Conefrey, Olivier Pascaud, Robin Courant, Xi Wang 외 arxiv

Films communicate through deliberate creative choices, including lighting, color, composition, editing, dialogue, music, and sound. Humans naturally interpret these signals as directorial intent, yet current multimodal l…

Multimodal Reasoning

CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions

2026-08-29 · Tsung-Han Wu, Heekyung Lee, Anya Ji, Haoming Chen 외 hf

Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT fo…

Multimodal Reasoning

전체 1,039편 보기 →