paper-with-me

홈 › Papers

Training Multimodal Systems for Classification with Multiple Objectives

2020-08-26 · Jason Armitage, Shramana Thakur, Rishi Tripathi, Jens Lehmann, Maria Maleshkova

We learn about the world from a diverse range of sensory information. Automated systems lack this ability as investigation has centred on processing information presented in a single form. Adapting architectures to learn from multiple modalities creates the potential to learn rich representations of the world - but current multimodal systems only deliver marginal improvements on unimodal approaches. Neural networks learn sampling noise during training with the result that performance on unseen data is degraded. This research introduces a second objective over the multimodal fusion process learned with variational inference. Regularisation methods are implemented in the inner training loop to control variance and the modular structure stabilises performance as additional neurons are added to layers. This framework is evaluated on a multilabel classification task with textual and visual inputs to demonstrate the potential for multiple objectives and probabilistic methods to lower variance and improve generalisation.

📄 PDF Abstract BibTeX arXiv:2008.11450

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationVariational Inference

Similar Papers 제목 키워드 기반

G$^{2}$D: Boosting Multimodal Learning with Gradient-Guided Distillation

2025-06-26 · Mohammed Rakib, Arunkumar Bagavathi

Multimodal learning aims to leverage information from diverse data modalities to achieve more comprehensive performance. However, conventional multimodal models often suffer from modality imbalance, where one or a few mo…

Knowledge DistillationModel Optimization

Adversarial Attacks on Multimodal Large Language Models: A Comprehensive Survey

2026-03-30 · Bhavuk Jain, Sercan Ö. Arık, Hardeo K. Thakur arxiv

Multimodal large language models (MLLMs) integrate information from multiple modalities such as text, images, audio, and video, enabling complex capabilities such as visual question answering and audio translation. While…

Visual Question Answering

Balancing Multimodal Domain Generalization via Gradient Modulation and Projection

2026-03-15 · Hongzhao Li, Guohao Shen, Shupan Li, Mingliang Xu 외 arxiv

Multimodal Domain Generalization (MMDG) leverages the complementary strengths of multiple modalities to enhance model generalization on unseen domains. A central challenge in multimodal learning is optimization imbalance…

Domain Generalization

Effective Backdoor Mitigation in Vision-Language Models Depends on the Pre-training Objective

2023-11-25 · Sahil Verma, Gantavya Bhatt, Avi Schwarzschild, Soumye Singhal 외

Despite the advanced capabilities of contemporary machine learning (ML) models, they remain vulnerable to adversarial and backdoor attacks. This vulnerability is particularly concerning in real-world deployments, where c…

zero-shot-classificationZero-Shot Learning

Toward Universal Text-to-Music Retrieval

2022-11-26 · Seungheon Doh, Minz Won, Keunwoo Choi, Juhan Nam

This paper introduces effective design choices for text-to-music retrieval systems. An ideal text-based retrieval system would support various input queries such as pre-defined tags, unseen tags, and sentence-level descr…

Music ClassificationRetrievalSentenceTAG