On the Contributions of Visual and Textual Supervision in Low-Resource Semantic Speech Retrieval
Recent work has shown that speech paired with images can be used to learn semantically meaningful speech representations even without any textual supervision. In real-world low-resource settings, however, we often have access to some transcribed speech. We study whether and how visual grounding is useful in the presence of varying amounts of textual supervision. In particular, we consider the task of semantic speech retrieval in a low-resource setting. We use a previously studied data set and task, where models are trained on images with spoken captions and evaluated on human judgments of semantic relevance. We propose a multitask learning approach to leverage both visual and textual modalities, with visual supervision in the form of keyword probabilities from an external tagger. We find that visual grounding is helpful even in the presence of textual supervision, and we analyze this effect over a range of sizes of transcribed data sets. With ~5 hours of transcribed speech, we obtain 23% higher average precision when also using visual supervision.
Code (0)
등록된 구현이 없습니다.
Tasks
RetrievalVisual GroundingSimilar Papers 제목 키워드 기반
Segment-Phrase Table for Semantic Segmentation, Visual Entailment and Paraphrasing
We introduce Segment-Phrase Table (SPT), a large collection of bijective associations between textual phrases and their corresponding segmentations. Leveraging recent progress in object recognition and natural language s…
Natural Language UnderstandingObject RecognitionSemantic SegmentationVisual EntailmentMulti-View Synergistic Learning with Vision-Language Adaption for Low-Resource Biomedical Image Classification
Accurate biomedical image classification under low-resource conditions remains challenging due to limited annotations, subtle inter-class visual differences, and complex disease semantics. While vision--language models o…
parameter-efficient fine-tuningRepresentation LearningContrastive LearningImage ClassificationBASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
Mainstream Multimodal Large Language Models (MLLMs) achieve visual understanding by using a vision projector to bridge well-pretrained vision encoders and large language models (LLMs). The inherent gap between visual and…
Textual Supervision for Visually Grounded Spoken Language Understanding
Visually-grounded models of spoken language understanding extract semantic information directly from speech, without relying on transcriptions. This is useful for low-resource languages, where transcriptions can be expen…
Spoken Language UnderstandingHierarchical Semantic Correspondence Networks for Video Paragraph Grounding
Video Paragraph Grounding (VPG) is an essential yet challenging task in vision-language understanding, which aims to jointly localize multiple events from an untrimmed video with a paragraph query description. One of…
DecoderSentenceVideo Grounding