paper-with-me

홈 › Papers

What is Multimodality?

2021-03-10 · ACL (mmsr, IWCS) 2021 6 · Letitia Parcalabescu, Nils Trost, Anette Frank

The last years have shown rapid developments in the field of multimodal machine learning, combining e.g., vision, text or speech. In this position paper we explain how the field uses outdated definitions of multimodality that prove unfit for the machine learning era. We propose a new task-relative definition of (multi)modality in the context of multimodal machine learning that focuses on representations and information that are relevant for a given machine learning task. With our new definition of multimodality we aim to provide a missing foundation for multimodal research, an important component of language grounding and a crucial milestone towards NLU.

📄 PDF Abstract BibTeX arXiv:2103.06304

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningPosition

Similar Papers 제목 키워드 기반

Exact Probability Landscapes of Stochastic Phenotype Switching in Feed-Forward Loops: Phase Diagrams of Multimodality

2021-04-06 · Anna Terebus, Farid Manuchehrfar, Youfang Cao, Jie Liang

Feed-forward loops (FFLs) are among the most ubiquitously found motifs of reaction networks in nature. However, little is known about their stochastic behavior and the variety of network phenotypes they can exhibit. In t…

What Is Missing in Multilingual Visual Reasoning and How to Fix It

2024-03-03 · Yueqi Song, Simran Khanuja, Graham Neubig

NLP models today strive for supporting multiple languages and modalities, improving accessibility for diverse users. In this paper, we evaluate their multilingual, multimodal capabilities by testing on a visual reasoning…

Image CaptioningVisual Reasoning

Reliable Multimodality Eye Disease Screening via Mixture of Student's t Distributions

2023-03-17 · Ke Zou, Tian Lin, Xuedong Yuan, Haoyu Chen 외

Multimodality eye disease screening is crucial in ophthalmology as it integrates information from diverse sources to complement their respective performances. However, the existing methods are weak in assessing the relia…

Decision Making

Unispeaker: A Unified Approach for Multimodality-driven Speaker Generation

2025-01-11 · Zhengyan Sheng, Zhihao Du, Heng Lu, Shiliang Zhang 외

Recent advancements in personalized speech generation have brought synthetic speech increasingly close to the realism of target speakers' recordings, yet multimodal speaker generation remains on the rise. This paper intr…

Diversity

Multimodality in Meta-Learning: A Comprehensive Survey

2021-09-28 · Yao Ma, Shilin Zhao, Weixiao Wang, Yaoman Li 외

Meta-learning has gained wide popularity as a training framework that is more data-efficient than traditional machine learning methods. However, its generalization ability in complex task distributions, such as multimoda…

Few-Shot LearningMeta-LearningSurveyZero-Shot Learning