paper-with-me

Papers

Vision-Speech Models: Teaching Speech Models to Converse about Images

2025-03-19 · Amélie Royer, Moritz Böhle, Gabriel de Marmiesse, Laurent Mazaré, Neil Zeghidour, Alexandre Défossez, Patrick Pérez

The recent successes of Vision-Language models raise the question of how to equivalently imbue a pretrained speech model with vision understanding, an important milestone towards building a multimodal speech model able to freely converse about images. Building such a conversational Vision-Speech model brings its unique challenges: (i) paired image-speech datasets are much scarcer than their image-text counterparts, (ii) ensuring real-time latency at inference is crucial thus bringing compute and memory constraints, and (iii) the model should preserve prosodic features (e.g., speaker tone) which cannot be inferred from text alone. In this work, we introduce MoshiVis, augmenting a recent dialogue speech LLM, Moshi, with visual inputs through lightweight adaptation modules. An additional dynamic gating mechanism enables the model to more easily switch between the visual inputs and unrelated conversation topics. To reduce training costs, we design a simple one-stage, parameter-efficient fine-tuning pipeline in which we leverage a mixture of image-text (i.e., "speechless") and image-speech samples. We evaluate the model on downstream visual understanding tasks with both audio and text prompts, and report qualitative samples of interactions with MoshiVis. Our inference code will be made available, as well as the image-speech data used for audio evaluation.

📄 PDF Abstract BibTeX arXiv:2503.15633

Code (1)

kyutai-labs/moshivis 공식 구현 jax

Tasks

parameter-efficient fine-tuning

Similar Papers 제목 키워드 기반

Infant directed speech is consistent with teaching

2016-06-01 · Baxter S. Eaves Jr., Naomi H. Feldman, Thomas L. Griffiths, Patrick Shafto

Infant-directed speech (IDS) has distinctive properties that differ from adult-directed speech (ADS). Why it has these properties -- and whether they are intended to facilitate language learning -- is matter of contentio…

CLeLfPC: a Large Open Multi-Speaker Corpus of French Cued Speech

2022-06-01 · LREC 2022 6 · Brigitte Bigi, Maryvonne Zimmermann, Carine André

Cued Speech is a communication system developed for deaf people to complement speechreading at the phonetic level with hands. This visual communication mode uses handshapes in different placements near the face in combin…

Transliteration

Speaker Attentive Speech Emotion Recognition

2021-04-15 · Clément Le Moine, Nicolas Obin, Axel Roebel

Speech Emotion Recognition (SER) task has known significant improvements over the last years with the advent of Deep Neural Networks (DNNs). However, even the most successful methods are still rather failing when adaptat…

Emotion RecognitionSpeech Emotion Recognition

Kallaama: A Transcribed Speech Dataset about Agriculture in the Three Most Widely Spoken Languages in Senegal

2024-04-02 · Elodie Gauthier, Aminata Ndiaye, Abdoulaye Guissé

This work is part of the Kallaama project, whose objective is to produce and disseminate national languages corpora for speech technologies developments, in the field of agriculture. Except for Wolof, which benefits from…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

On the Role of Visual Cues in Audiovisual Speech Enhancement

2020-04-25 · Zakaria Aldeneh, Anushree Prasanna Kumar, Barry-John Theobald, Erik Marchi 외

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target spee…

Self-Supervised LearningSpeech Enhancement