paper-with-me

홈 › Papers

Unsupervised Mismatch Localization in Cross-Modal Sequential Data with Application to Mispronunciations Localization

2022-05-05 · Wei Wei, Huang Hengguan, Gu Xiangming, Wang Hao, Wang Ye

Content mismatch usually occurs when data from one modality is translated to another, e.g. language learners producing mispronunciations (errors in speech) when reading a sentence (target text) aloud. However, most existing alignment algorithms assume that the content involved in the two modalities is perfectly matched, thus leading to difficulty in locating such mismatch between speech and text. In this work, we develop an unsupervised learning algorithm that can infer the relationship between content-mismatched cross-modal sequential data, especially for speech-text sequences. More specifically, we propose a hierarchical Bayesian deep learning model, dubbed mismatch localization variational autoencoder (ML-VAE), which decomposes the generative process of the speech into hierarchically structured latent variables, indicating the relationship between the two modalities. Training such a model is very challenging due to the discrete latent variables with complex dependencies involved. To address this challenge, we propose a novel and effective training procedure that alternates between estimating the hard assignments of the discrete latent variables over a specifically designed mismatch localization finite-state acceptor (ML-FSA) and updating the parameters of neural networks. In this work, we focus on the mismatch localization problem for speech and text, and our experimental results show that ML-VAE successfully locates the mismatch between text and speech, without the need for human annotations for model training.

📄 PDF Abstract BibTeX arXiv:2205.02670

Code (0)

등록된 구현이 없습니다.

Tasks

Sentence

Similar Papers 제목 키워드 기반

Teaching deep neural networks to localize single molecules for super-resolution microscopy

2019-06-27 · Artur Speiser, Lucas-Raphael Müller, Ulf Matti, Christopher J. Obara 외

Single-molecule localization fluorescence microscopy constructs super-resolution images by sequential imaging and computational localization of sparsely activated fluorophores. Accurate and efficient fluorophore localiza…

Bayesian InferenceSuper-Resolution

Open-World Distributed Robot Self-Localization with Transferable Visual Vocabulary and Both Absolute and Relative Features

2021-09-09 · Mitsuki Yoshida, Ryogo Yamamoto, Daiki Iwata, Kanji Tanaka

Visual robot self-localization is a fundamental problem in visual robot navigation and has been studied across various problem settings, including monocular and sequential localization. However, many existing studies foc…

Graph Neural NetworkRobot Navigation

Sound Source Localization is All about Cross-Modal Alignment

2023-09-19 · ICCV 2023 1 · Arda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh 외

Humans can easily perceive the direction of sound sources in a visual scene, termed sound source localization. Recent studies on learning-based sound source localization have mainly explored the problem from a localizati…

Allcross-modal alignmentCross-Modal RetrievalRetrieval+1

Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment

2024-07-18 · Arda Senocak, Hyeonggon Ryu, Junsik Kim, Tae-Hyun Oh 외

Recent studies on learning-based sound source localization have mainly focused on the localization performance perspective. However, prior work and existing benchmarks overlook a crucial aspect: cross-modal interaction, …

cross-modal alignmentCross-Modal RetrievalSound Source Localization

Trajectory-aware Cross-view Geo-localization with Sequential Observations

2026-07-16 · Tianyi Gao, Jiayu Lin, Danielle Beaulieu, Nathan Jacobs arxiv

Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods show that sequential queries such as video clips yield richer spatiotemporal cues than single images, yet…