paper-with-me

홈 › Papers

Speech-Image Semantic Alignment Does Not Depend on Any Prior Classification Tasks

2020-10-29 · Masood S. Mortazavi

Semantically-aligned $(speech, image)$ datasets can be used to explore "visually-grounded speech". In a majority of existing investigations, features of an image signal are extracted using neural networks "pre-trained" on other tasks (e.g., classification on ImageNet). In still others, pre-trained networks are used to extract audio features prior to semantic embedding. Without "transfer learning" through pre-trained initialization or pre-trained feature extraction, previous results have tended to show low rates of recall in $speech \rightarrow image$ and $image \rightarrow speech$ queries. Choosing appropriate neural architectures for encoders in the speech and image branches and using large datasets, one can obtain competitive recall rates without any reliance on any pre-trained initialization or feature extraction: $(speech,image)$ semantic alignment and $speech \rightarrow image$ and $image \rightarrow speech$ retrieval are canonical tasks worthy of independent investigation of their own and allow one to explore other questions---e.g., the size of the audio embedder can be reduced significantly with little loss of recall rates in $speech \rightarrow image$ and $image \rightarrow speech$ queries.

📄 PDF Abstract BibTeX arXiv:2010.15288

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationRetrievalTransfer Learning

Similar Papers 제목 키워드 기반

Speech Vecalign: an Embedding-based Method for Aligning Parallel Speech Documents

2025-09-22 · Chutong Meng, Philipp Koehn arxiv

We present Speech Vecalign, a parallel speech document alignment method that monotonically aligns speech segment embeddings and does not depend on text transcriptions. Compared to the baseline method Global Mining, a var…

Speech-to-Speech Translation

The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders

2025-09-19 · Adrian Sauter, Willem Zuidema, Marianne de Heer Kloots arxiv

How does visual information included in training affect language processing in audio- and text-based deep learning models? We explore how such visual grounding affects model-internal representations of words, and find su…

Visual Grounding

JETS: Jointly Training FastSpeech2 and HiFi-GAN for End to End Text to Speech

2022-03-31 · Dan Lim, Sunghee Jung, Eesung Kim

In neural text-to-speech (TTS), two-stage system or a cascade of separately learned models have shown synthesis quality close to human speech. For example, FastSpeech2 transforms an input text to a mel-spectrogram and th…

text-to-speechText to Speech

Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models

2021-07-05 · Khazar Khorrami, Okko Räsänen

Systems that can find correspondences between multiple modalities, such as between speech and images, have great potential to solve different recognition and data analysis tasks in an unsupervised manner. This work studi…

Cross-Modal RetrievalObject LocalizationRetrievalSemantic Retrieval

PortaSpeech: Portable and High-Quality Generative Text-to-Speech

2021-09-30 · NeurIPS 2021 12 · Yi Ren, Jinglin Liu, Zhou Zhao

Non-autoregressive text-to-speech (NAR-TTS) models such as FastSpeech 2 and Glow-TTS can synthesize high-quality speech from the given text in parallel. After analyzing two kinds of generative NAR-TTS models (VAE and nor…

text-to-speechText to SpeechText-To-Speech SynthesisVocal Bursts Intensity Prediction+1