paper-with-me

Papers

KBE-DME: Dynamic Multimodal Evaluation via Knowledge Enhanced Benchmark Evolution

2025-10-24 · Junzhe Zhang, Huixuan Zhang, Xiaojun Wan arxiv

The rapid progress of multimodal large language models (MLLMs) calls for more reliable evaluation protocols. Existing static benchmarks suffer from the potential risk of data contamination and saturation, leading to inflated or misleading performance evaluations. To address these issues, we first apply Graph formulation to represent a static or dynamic VQA sample. With the formulation, we propose Knowledge-enhanced Benchmark Evolution(KBE), a dynamic multimodal evaluation framework. KBE first analyzes the original static benchmark, then expands it by integrating multimodal knowledge, transforming the static benchmark into a controllable, dynamic evolving version. Crucially, KBE can both reconstruct questions by Re-selecting visual information in the original image and expand existing questions with external textual knowledge. It enables difficulty-controllable evaluation by adjusting the degree of question exploration. Extensive experiments demonstrate that KBE alleviates the risk of data contamination, data saturation, and provides a more comprehensive assessment of MLLM capabilities.

📄 PDF Abstract BibTeX arXiv:2510.21182

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Dynamic Knowledge Integration for Enhanced Vision-Language Reasoning

2025-01-15 · Julian Perry, Surasakdi Siripong, Thanakorn Phonchai

Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multimodal tasks, but their performance is often constrained by the lack of external knowledge integration, limiting their ability to hand…

Question AnsweringVisual Question AnsweringWorld Knowledge

CUE-M: Contextual Understanding and Enhanced Search with Multimodal Large Language Model

2024-11-19 · Dongyoung Go, Taesun Whang, Chanhee Lee, Hwa-Yeon Kim 외

The integration of Retrieval-Augmented Generation (RAG) with Multimodal Large Language Models (MLLMs) has revolutionized information retrieval and expanded the practical applications of AI. However, current systems strug…

Information RetrievalLanguage ModelingLanguage ModellingLarge Language Model+4

Knowledge-enhanced Multimodal ECG Representation Learning with Arbitrary-Lead Inputs

2025-02-25 · Che Liu, Cheng Ouyang, Zhongwei Wan, Haozhe Wang 외

Recent advances in multimodal ECG representation learning center on aligning ECG signals with paired free-text reports. However, suboptimal alignment persists due to the complexity of medical language and the reliance on…

Representation Learningzero-shot-classificationZero-Shot Learning

MUOT_3M: A 3 Million Frame Multimodal Underwater Benchmark and the MUTrack Tracking Method

2026-02-20 · Ahsan Baidar Bakht, Mohamad Alansari, Muhayy Ud Din, Muzammal Naseer 외 arxiv

Underwater Object Tracking (UOT) is crucial for efficient marine robotics, large scale ecological monitoring, and ocean exploration; however, progress has been hindered by the scarcity of large, multimodal, and diverse d…

Knowledge DistillationObject Tracking

KGCE: Knowledge-Augmented Dual-Graph Evaluator for Cross-Platform Educational Agent Benchmarking with Multimodal Language Models

2026-01-04 · Zixian Liu, Sihao Liu, Yuqi Zhao arxiv

With the rapid adoption of multimodal large language models (MLMs) in autonomous agents, cross-platform task execution capabilities in educational settings have garnered significant attention. However, existing benchmark…