paper-with-me

Papers

Mixture of Probes: Learning from Privileged Modalities in Multimodal LLMs Through Probing

2026-07-09 · Dominick Reilly, Qiyu Wu, Hiromi Wakaki, Srijan Das, Yuki Mistufuji arxiv

Multimodal Large Language Models (MLLMs) are typically designed under the assumption that all modalities available during training will also be accessible at inference. However, many real-world settings violate this assumption, requiring models to operate under a privileged modality setting, where auxiliary modalities are available only during training. While these modalities contain valuable information, existing MLLMs largely fail to leverage them effectively, as they treat modalities as interchangeable inputs rather than sources of complementary supervision. We propose Mixture of Probes (MoP), a novel framework that disentangles modality-specific and modality-general signals within the MLLM, allowing the model to preserve modality-dependent structure while learning transferable representations across modalities. At its core, MoP achieves this through a structured probing mechanism that extracts and organizes information from intermediate representations of a shared modality encoder, rather than relying only on final-layer alignment as done in existing MLLMs. To support this disentanglement, we further introduce MoP Cross-modal Training (MoP-X), a training strategy for MoP centered around a probe disentanglement loss that prevents probe collapse and encourages cross-modal learning. We evaluate MoP across two domains spanning eight tasks and four modalities under a comprehensive evaluation protocol tailored to the privileged modality setting, where each modality is independently treated as the sole input at inference time. MoP consistently outperforms strong MLLM baselines, achieving up to 65% relative improvement, demonstrating that auxiliary modalities, even when unavailable at inference, can provide substantial gains when effectively leveraged during training. Code, model checkpoints, and evaluation protocols will be made available at https://github.com/Sony/MoP.

📄 PDF Abstract BibTeX arXiv:2607.08839

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Distilling Privileged Multimodal Information for Expression Recognition using Optimal Transport

2024-01-27 · Muhammad Haseeb Aslam, Muhammad Osama Zeeshan, Soufiane Belharbi, Marco Pedersoli 외

Deep learning models for multimodal expression recognition have reached remarkable performance in controlled laboratory environments because of their ability to learn complementary and redundant semantic information. How…

DiversityKnowledge DistillationOrdinal Classification

Graph Distillation for Action Detection with Privileged Modalities

2017-11-30 · ECCV 2018 9 · Zelun Luo, Jun-Ting Hsieh, Lu Jiang, Juan Carlos Niebles 외

We propose a technique that tackles action detection in multimodal videos under a realistic and challenging condition in which only limited training data and partially observed modalities are available. Common methods in…

Action ClassificationAction DetectionTransfer Learning

On the effectiveness of multimodal privileged knowledge distillation in two vision transformer based diagnostic applications

2025-08-06 · Simon Baur, Alexandra Benova, Emilio Dolgener Cantú, Jackie Ma arxiv

Deploying deep learning models in clinical practice often requires leveraging multiple data modalities, such as images, text, and structured data, to achieve robust and trustworthy decisions. However, not all modalities …

Knowledge Distillation

Multi Teacher Privileged Knowledge Distillation for Multimodal Expression Recognition

2024-08-16 · Muhammad Haseeb Aslam, Marco Pedersoli, Alessandro Lameiras Koerich, Eric Granger

Human emotion is a complex phenomenon conveyed and perceived through facial expressions, vocal tones, body language, and physiological signals. Multimodal emotion recognition systems can perform well because they can lea…

Emotion RecognitionKnowledge DistillationMultimodal Emotion Recognition

Inertial Hallucinations -- When Wearable Inertial Devices Start Seeing Things

2022-07-14 · Alessandro Masullo, Toby Perrett, Tilo Burghardt, Ian Craddock 외

We propose a novel approach to multimodal sensor fusion for Ambient Assisted Living (AAL) which takes advantage of learning using privileged information (LUPI). We address two major shortcomings of standard multimodal ap…

HallucinationSensor FusionTriplet