paper-with-me

Papers

Shifting the Baseline: Single Modality Performance on Visual Navigation & QA

2018-11-01 · Jesse Thomason, Daniel Gordon, Yonatan Bisk

We demonstrate the surprising strength of unimodal baselines in multimodal domains, and make concrete recommendations for best practices in future research. Where existing work often compares against random or majority class baselines, we argue that unimodal approaches better capture and reflect dataset biases and therefore provide an important comparison when assessing the performance of multimodal techniques. We present unimodal ablations on three recent datasets in visual navigation and QA, seeing an up to 29% absolute gain in performance over published baselines.

📄 PDF Abstract BibTeX arXiv:1811.00613

Code (0)

등록된 구현이 없습니다.

Tasks

Question AnsweringVisual Navigation

Similar Papers 제목 키워드 기반

Shifting the Baseline: Single Modality Performance on Visual Navigation \& QA

2019-06-01 · NAACL 2019 6 · Jesse Thomason, Daniel Gordon, Yonatan Bisk

We demonstrate the surprising strength of unimodal baselines in multimodal domains, and make concrete recommendations for best practices in future research. Where existing work often compares against random or majority c…

Visual Navigation

On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI

2025-07-31 · David Restrepo, Ira Ktena, Maria Vakalopoulou, Stergios Christodoulidis 외 arxiv

Clinical decision-making relies on the integrated analysis of medical images and the associated clinical reports. While Vision-Language Models (VLMs) can offer a unified framework for such tasks, they can exhibit strong …

Binary Classification

OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder

2026-05-02 · Detao Bai, Shimin Yao, Weixuan Chen, Chengen Lai 외 arxiv

Recent advances in omni-modal large language models have enabled remarkable progress in joint vision-audio understanding. However, prevailing architectures rely on modality-specific encoders with a \emph{video-coarse, au…

Sign Language RecognitionComputational EfficiencySpeaker Identification

CMSBERT-CLR: Context-driven Modality Shifting BERT with Contrastive Learning for linguistic, visual, acoustic Representations

2022-08-21 · Junghun Kim, Jihie Kim

Multimodal sentiment analysis has become an increasingly popular research area as the demand for multimodal online content is growing. For multimodal sentiment analysis, words can have different meanings depending on the…

Contrastive LearningMultimodal Sentiment AnalysisSentenceSentiment Analysis

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test

2025-05-28 · Guangfu Hao, Frederic Alexandre, Shan Yu

Cognitive flexibility has been extensively studied in human cognition but remains relatively unexplored in the context of Visual Large Language Models (VLLMs). This study assesses the cognitive flexibility of state-of-th…