paper-with-me

홈 › Papers

On the Contributions of Visual and Textual Supervision in Low-Resource Semantic Speech Retrieval

2019-04-24 · Ankita Pasad, Bowen Shi, Herman Kamper, Karen Livescu

Recent work has shown that speech paired with images can be used to learn semantically meaningful speech representations even without any textual supervision. In real-world low-resource settings, however, we often have access to some transcribed speech. We study whether and how visual grounding is useful in the presence of varying amounts of textual supervision. In particular, we consider the task of semantic speech retrieval in a low-resource setting. We use a previously studied data set and task, where models are trained on images with spoken captions and evaluated on human judgments of semantic relevance. We propose a multitask learning approach to leverage both visual and textual modalities, with visual supervision in the form of keyword probabilities from an external tagger. We find that visual grounding is helpful even in the presence of textual supervision, and we analyze this effect over a range of sizes of transcribed data sets. With ~5 hours of transcribed speech, we obtain 23% higher average precision when also using visual supervision.

📄 PDF Abstract BibTeX arXiv:1904.10947

Code (0)

등록된 구현이 없습니다.

Tasks

RetrievalVisual Grounding

Similar Papers 제목 키워드 기반

Segment-Phrase Table for Semantic Segmentation, Visual Entailment and Paraphrasing

2015-09-27 · ICCV 2015 12 · Hamid Izadinia, Fereshteh Sadeghi, Santosh Kumar Divvala, Yejin Choi 외

We introduce Segment-Phrase Table (SPT), a large collection of bijective associations between textual phrases and their corresponding segmentations. Leveraging recent progress in object recognition and natural language s…

Natural Language UnderstandingObject RecognitionSemantic SegmentationVisual Entailment

Multi-View Synergistic Learning with Vision-Language Adaption for Low-Resource Biomedical Image Classification

2026-04-27 · Xiaoliu Luo, Minxue Xiao, Ting Xie, Mengzhu Wang 외 arxiv

Accurate biomedical image classification under low-resource conditions remains challenging due to limited annotations, subtle inter-class visual differences, and complex disease semantics. While vision--language models o…

parameter-efficient fine-tuningRepresentation LearningContrastive LearningImage Classification

BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models

2025-08-09 · Jianting Tang, Yubo Wang, Haoyu Cao, Linli Xu arxiv

Mainstream Multimodal Large Language Models (MLLMs) achieve visual understanding by using a vision projector to bridge well-pretrained vision encoders and large language models (LLMs). The inherent gap between visual and…

Textual Supervision for Visually Grounded Spoken Language Understanding

2020-10-06 · Findings of the Association for Computational Linguistics 2020 · Bertrand Higy, Desmond Elliott, Grzegorz Chrupała

Visually-grounded models of spoken language understanding extract semantic information directly from speech, without relying on transcriptions. This is useful for low-resource languages, where transcriptions can be expen…

Spoken Language Understanding

Hierarchical Semantic Correspondence Networks for Video Paragraph Grounding

2023-01-01 · CVPR 2023 1 · Chaolei Tan, Zihang Lin, Jian-Fang Hu, Wei-Shi Zheng 외

Video Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of…

DecoderSentenceVideo Grounding