paper-with-me

홈 › Papers

Cascaded Self-Evaluation Augmented Training for Efficient Multimodal Large Language Models

2025-01-10 · Zheqi Lv, Wenkai Wang, Jiawei Wang, Shengyu Zhang, Fei Wu

Efficient Multimodal Large Language Models (EMLLMs) have rapidly advanced recently. Incorporating Chain-of-Thought (CoT) reasoning and step-by-step self-evaluation has improved their performance. However, limited parameters often hinder EMLLMs from effectively using self-evaluation during inference. Key challenges include synthesizing evaluation data, determining its quantity, optimizing training and inference strategies, and selecting appropriate prompts. To address these issues, we introduce Self-Evaluation Augmented Training (SEAT). SEAT uses more powerful EMLLMs for CoT reasoning, data selection, and evaluation generation, then trains EMLLMs with the synthesized data. However, handling long prompts and maintaining CoT reasoning quality are problematic. Therefore, we propose Cascaded Self-Evaluation Augmented Training (Cas-SEAT), which breaks down lengthy prompts into shorter, task-specific cascaded prompts and reduces costs for resource-limited settings. During data synthesis, we employ open-source 7B-parameter EMLLMs and annotate a small dataset with short prompts. Experiments demonstrate that Cas-SEAT significantly boosts EMLLMs' self-evaluation abilities, improving performance by 19.68%, 55.57%, and 46.79% on the MathVista, Math-V, and We-Math datasets, respectively. Additionally, our Cas-SEAT Dataset serves as a valuable resource for future research in enhancing EMLLM self-evaluation.

📄 PDF Abstract BibTeX arXiv:2501.05662

Code (0)

등록된 구현이 없습니다.

Tasks

Math

Similar Papers 제목 키워드 기반

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

2026-07-27 · Hangjie Yuan, Yichen Qian, Zhiwei Tang, Xianzhe Xu 외 hf

Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from …

Visual Question AnsweringInstruction Following

Trust Before Fusion: QIMG-7 and Source-Aware Resolution for Polluted Multimodal RAG

2026-07-12 · Saadeldine Eletter, Owais Aijaz, Preslav Nakov arxiv

Multimodal retrieval-augmented generation (RAG) is often evaluated with clean evidence, yet real retrieval can return topically relevant but unreliable content: false text and misleading images from corrupted metadata, e…

Style Transfer

Multimodal Priors-Augmented Text-Driven 3D Human-Object Interaction Generation

2026-02-11 · Yin Wang, Ziyao Zhang, Zhiying Leng, Haitian Liu 외 arxiv

We address the challenging task of text-driven 3D human-object interaction (HOI) motion generation. Existing methods primarily rely on a direct text-to-HOI mapping, which suffers from three key limitations due to the sig…

Decoupling Semantics and Logic: A Training-Free Coarse-to-Fine Pipeline for Video Retrieval-Augmented Generation

2026-06-06 · Jiaxin Dai, Zehang Wei, Jiamin Yan, Xiang Xiang arxiv

This paper presents our system description for the 2nd Workshop on Multimodal Augmented Generation via MultimodAl Retrieval (MAGMaR). Addressing the critical challenges of cross-lingual long-video comprehension, strict p…

Information RetrievalSemantic RetrievalLogical ReasoningVideo Retrieval

SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data

2024-02-10 · Hsuan-Fu Wang, Yi-Jen Shih, Heng-Jui Chang, Layne Berry 외

The recently proposed visually grounded speech model SpeechCLIP is an innovative framework that bridges speech and text through images via CLIP without relying on text transcription. On this basis, this paper introduces …

Keyword ExtractionMulti-Task LearningRepresentation LearningRetrieval