paper-with-me

홈 › Papers

Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation

2025-04-04 · Jaewoo Park, Jungyang Park, Dongju Jang, Jiwan Chung, Byungwoo Yoo, Jaewoo Shin, Seonjoon Park, Taehyeong Kim, Youngjae Yu

With the rapid advancement of mathematical reasoning capabilities in large language models (LLMs), AI systems are increasingly being adopted in educational settings to support students' comprehension of problem-solving processes. However, a critical component remains underexplored in current LLM-generated explanations: visual explanation. In real-world instructional contexts, human tutors routinely employ visual aids-such as diagrams, markings, and highlights-to enhance conceptual clarity. To bridge this gap, we introduce a novel task of visual solution explanation, which requires not only solving problems but also generating explanations that incorporate newly introduced visual elements essential for understanding (e.g., auxiliary lines, annotations, or geometric constructions). To evaluate model performance on this task, we propose MathExplain, a multimodal benchmark consisting of 997 math problems annotated with visual keypoints and corresponding explanatory text that references those elements. Our empirical results show that while some closed-source models demonstrate promising capabilities on visual solution-explaining, current open-source general-purpose models perform inconsistently, particularly in identifying relevant visual components and producing coherent keypoint-based explanations. We expect that visual solution-explaining and the MathExplain dataset will catalyze further research on multimodal LLMs in education and advance their deployment as effective, explanation-oriented AI tutors. Code and data will be released publicly.

📄 PDF Abstract BibTeX arXiv:2504.03197

Code (0)

등록된 구현이 없습니다.

Tasks

MathMathematical Reasoning

Similar Papers 제목 키워드 기반

Classroom-Inspired Multi-Mentor Distillation with Adaptive Learning Strategies

2024-09-30 · Shalini Sarode, Muhammad Saif Ullah Khan, Tahira Shehzadi, Didier Stricker 외

We propose ClassroomKD, a novel multi-mentor knowledge distillation framework inspired by classroom environments to enhance knowledge transfer between student and multiple mentors. Unlike traditional methods that rely on…

2D Human Pose Estimationimage-classificationImage ClassificationKnowledge Distillation+2

MENTOR: Mixture-of-Experts Network with Task-Oriented Perturbation for Visual Reinforcement Learning

2024-10-19 · Suning Huang, Zheyu Zhang, Tianhai Liang, Yihan Xu 외

Visual deep reinforcement learning (RL) enables robots to acquire skills from visual input for unstructured tasks. However, current algorithms suffer from low sample efficiency, limiting their practical applicability. In…

Deep Reinforcement LearningMixture-of-ExpertsReinforcement Learning (RL)

Robotic Surgery Remote Mentoring via AR with 3D Scene Streaming and Hand Interaction

2022-04-09 · Yonghao Long, Chengkun Li, Qi Dou

With the growing popularity of robotic surgery, education becomes increasingly important and urgently needed for the sake of patient safety. However, experienced surgeons have limited accessibility due to their busy clin…

Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation

2024-11-26 · Xu Zheng, Haiwei Xue, Jialei Chen, Yibo Yan 외

Simultaneously using multimodal inputs from multiple sensors to train segmentors is intuitively advantageous but practically challenging. A key challenge is unimodal bias, where multimodal segmentors over rely on certain…

Supporting Student Decisions on Learning Recommendations: An LLM-Based Chatbot with Knowledge Graph Contextualization for Conversational Explainability and Mentoring

2024-01-16 · Hasan Abu-Rasheed, Mohamad Hussam Abdulsalam, Christian Weber, Madjid Fathi

Student commitment towards a learning recommendation is not separable from their understanding of the reasons it was recommended to them; and their ability to modify it based on that understanding. Among explainability a…

Chatbot