paper-with-me

Papers

MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence

2026-05-12 · Yifan Chen, Fei Yin, Qingyan Bai, Zicheng Lin, Yujiu Yang arxiv

We introduce MMCL-Bench, a benchmark for multimodal context learning: learning task-local rules, procedures, and empirical patterns from visual or mixed-modality teaching context and applying them to new visual instances. Unlike text-only context learning or standard multimodal question answering, this setting requires models to recover and localize relevant evidence from images, screenshots, manuals, videos, and frame sequences before they can reason over the learned context. MMCL-Bench contains 102 tasks spanning three categories: rule system application, procedural task execution, and empirical discovery and induction. We evaluate frontier multimodal models with strict rubric-based scoring and find that current systems remain far from robust multimodal context learning, with even the strongest model solving fewer than one-third of tasks under strict evaluation. Diagnostic ablations and error analysis show that failures arise throughout the context-to-answer pipeline, including context anchoring, visual evidence extraction, context reasoning, and response construction. MMCL-Bench thus highlights multimodal context learning as an important unsolved capability bottleneck for current multimodal models.

📄 PDF Abstract BibTeX arXiv:2605.12703

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

2026-06-08 · Muhammad Umer Sheikh, Hassan Abid, Khawar Shehzad, Ufaq Khan 외 arxiv

Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate QA benchmarks are small, mostly textual, and cover a narrow range of …

Question Answering

Recent Advances of Multimodal Continual Learning: A Comprehensive Survey

2024-10-07 · Dianzhi Yu, Xinni Zhang, Yankai Chen, Aiwei Liu 외

Continual learning (CL) aims to empower machine learning models to learn continually from new data, while building upon previously acquired knowledge without forgetting. As machine learning models have evolved from small…

Continual LearningSurvey

Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition

2024-07-22 · Jinfu Liu, Chen Chen, Mengyuan Liu

Skeleton-based action recognition has garnered significant attention due to the utilization of concise and resilient skeletons. Nevertheless, the absence of detailed body information in skeletons restricts performance, w…

Action RecognitionContrastive LearningSkeleton Based Action Recognition

On the Generalization of Multi-modal Contrastive Learning

2023-06-07 · Qi Zhang, Yifei Wang, Yisen Wang

Multi-modal contrastive learning (MMCL) has recently garnered considerable interest due to its superior performance in visual tasks, achieved by embedding multi-modal data, such as visual-language pairs. However, there s…

Contrastive Learning

MMCLIP: Cross-modal Attention Masked Modelling for Medical Language-Image Pre-Training

2024-07-28 · Biao Wu, Yutong Xie, Zeyu Zhang, Minh Hieu Phan 외

Vision-and-language pretraining (VLP) in the medical field utilizes contrastive learning on image-text pairs to achieve effective transfer across tasks. Yet, current VLP approaches with the masked modeling strategy face …

Contrastive LearningLanguage ModelingLanguage ModellingMasked Language Modeling