paper-with-me

홈 › Papers

Spoken Moments: Learning Joint Audio-Visual Representations from Video Descriptions

2021-05-10 · CVPR 2021 1 · Mathew Monfort, SouYoung Jin, Alexander Liu, David Harwath, Rogerio Feris, James Glass, Aude Oliva

When people observe events, they are able to abstract key information and build concise summaries of what is happening. These summaries include contextual and semantic information describing the important high-level details (what, where, who and how) of the observed event and exclude background information that is deemed unimportant to the observer. With this in mind, the descriptions people generate for videos of different dynamic events can greatly improve our understanding of the key information of interest in each video. These descriptions can be captured in captions that provide expanded attributes for video labeling (e.g. actions/objects/scenes/sentiment/etc.) while allowing us to gain new insight into what people find important or necessary to summarize specific events. Existing caption datasets for video understanding are either small in scale or restricted to a specific domain. To address this, we present the Spoken Moments (S-MiT) dataset of 500k spoken captions each attributed to a unique short video depicting a broad range of different events. We collect our descriptions using audio recordings to ensure that they remain as natural and concise as possible while allowing us to scale the size of a large classification dataset. In order to utilize our proposed dataset, we present a novel Adaptive Mean Margin (AMM) approach to contrastive learning and evaluate our models on video/caption retrieval on multiple datasets. We show that our AMM approach consistently improves our results and that models trained on our Spoken Moments dataset generalize better than those trained on other video-caption datasets.

📄 PDF Abstract BibTeX arXiv:2105.04489

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningRetrievalVideo Understanding

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Joint Visual and Audio Learning for Video Highlight Detection

2021-01-01 · ICCV 2021 10 · Taivanbat Badamdorj, Mrigank Rochan, Yang Wang, Li Cheng

In video highlight detection, the goal is to identify the interesting moments within an unedited video. Although the audio component of the video provides important cues for highlight detection, the majority of exist…

Highlight Detection

Jointly Discovering Visual Objects and Spoken Words from Raw Sensory Input

2018-04-04 · ECCV 2018 9 · David Harwath, Adrià Recasens, Dídac Surís, Galen Chuang 외

In this paper, we explore neural network models that learn to associate segments of spoken audio captions with the semantically relevant portions of natural images that they refer to. We demonstrate that these audio-visu…

RetrievalSound Prompted Semantic SegmentationSpeech Prompted Semantic Segmentation

Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

2024-06-12 · Se Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim 외

In this paper, we introduce a novel Face-to-Face spoken dialogue model. It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an ava…

ChatbotLanguage ModelingLanguage ModellingLarge Language Model

Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech

2025-10-29 · Pedro Corrêa, João Lima, Victor Moreno, Lucas Ueda 외 arxiv

Advancements in spoken language processing have driven the development of spoken language models (SLMs), designed to achieve universal audio understanding by jointly learning text and audio representations for a wide ran…

Speech Emotion Recognition

Joint Multimodal Contrastive Learning for Robust Spoken Term Detection and Keyword Spotting

2025-12-16 · Ramesh Gundluru, Shubham Gupta, Sri Rama Murty K arxiv

Acoustic Word Embeddings (AWEs) improve the efficiency of speech retrieval tasks such as Spoken Term Detection (STD) and Keyword Spotting (KWS). However, existing approaches suffer from limitations, including unimodal su…

Contrastive LearningKeyword Spotting