paper-with-me

Papers

Read, Look or Listen? What's Needed for Solving a Multimodal Dataset

2023-07-06 · Netta Madvil, Yonatan Bitton, Roy Schwartz

The prevalence of large-scale multimodal datasets presents unique challenges in assessing dataset quality. We propose a two-step method to analyze multimodal datasets, which leverages a small seed of human annotation to map each multimodal instance to the modalities required to process it. Our method sheds light on the importance of different modalities in datasets, as well as the relationship between them. We apply our approach to TVQA, a video question-answering dataset, and discover that most questions can be answered using a single modality, without a substantial bias towards any specific modality. Moreover, we find that more than 70% of the questions are solvable using several different single-modality strategies, e.g., by either looking at the video or listening to the audio, highlighting the limited integration of multiple modalities in TVQA. We leverage our annotation and analyze the MERLOT Reserve, finding that it struggles with image-based questions compared to text and audio, but also with auditory speaker identification. Based on our observations, we introduce a new test set that necessitates multiple modalities, observing a dramatic drop in model performance. Our methodology provides valuable insights into multimodal datasets and highlights the need for the development of more robust models.

📄 PDF Abstract BibTeX arXiv:2307.04532

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringSpeaker IdentificationVideo Question Answering

Similar Papers 제목 키워드 기반

Speech language models lack important brain-relevant semantics

2023-11-08 · Subba Reddy Oota, Emin Çelik, Fatma Deniz, Mariya Toneva

Despite known differences between reading and listening in the brain, recent work has shown that text-based language models predict both text-evoked and speech-evoked brain activity to an impressive degree. This poses th…

Language ModelingLanguage Modelling

Discovering the Italian literature: interactive access to audio indexed text resources

2014-05-01 · LREC 2014 5 · Vincenzo Galat{\`a}, Alberto Benin, Piero Cosi, Giuseppe Riccardo Leone 외

In this paper we present a web interface to study Italian through the access to read Italian literature. The system allows to browse the content, search for specific words and listen to the correct pronunciation produced…

Cultural Vocal Bursts Intensity PredictionSentencetext-to-speechText to Speech

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

2026-08-20 · Alin-Ionut Popa arxiv

Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still trip them up under direct visual inference, even when the page is already sitting…

Visual Question Answering

Look, Listen and Learn

2017-05-23 · ICCV 2017 10 · Relja Arandjelović, Andrew Zisserman

We consider the question: what can be learnt by looking at and listening to a large number of unlabelled videos? There is a valuable, but so far untapped, source of information contained in the video itself -- the corres…

Audio ClassificationGeneral ClassificationSound Classification

Predicting Musical Sophistication from Music Listening Behaviors: A Preliminary Study

2018-08-22 · Bruce Ferwerda, Mark Graus

Psychological models are increasingly being used to explain online behavioral traces. Aside from the commonly used personality traits as a general user model, more domain dependent models are gaining attention. The use o…