paper-with-me

Papers

Exploring Deep Learning for Joint Audio-Visual Lip Biometrics

2021-04-17 · Meng Liu, Longbiao Wang, Kong Aik Lee, Hanyi Zhang, Chang Zeng, Jianwu Dang

Audio-visual (AV) lip biometrics is a promising authentication technique that leverages the benefits of both the audio and visual modalities in speech communication. Previous works have demonstrated the usefulness of AV lip biometrics. However, the lack of a sizeable AV database hinders the exploration of deep-learning-based audio-visual lip biometrics. To address this problem, we compile a moderate-size database using existing public databases. Meanwhile, we establish the DeepLip AV lip biometrics system realized with a convolutional neural network (CNN) based video module, a time-delay neural network (TDNN) based audio module, and a multimodal fusion module. Our experiments show that DeepLip outperforms traditional speaker recognition models in context modeling and achieves over 50% relative improvements compared with our best single modality baseline, with an equal error rate of 0.75% and 1.11% on the test datasets, respectively.

📄 PDF Abstract BibTeX arXiv:2104.08510

Code (1)

DanielMengLiu/DeepLip 공식 구현 pytorch

Tasks

Deep LearningSpeaker Recognition

Similar Papers 제목 키워드 기반

Multilingual Audio-Visual Smartphone Dataset And Evaluation

2021-09-09 · Hareesh Mandalapu, Aravinda Reddy P N, Raghavendra Ramachandra, K Sreenivasa Rao 외

Smartphones have been employed with biometric-based verification systems to provide security in highly sensitive applications. Audio-visual biometrics are getting popular due to their usability, and also it will be chall…

Speaker Recognition

AudioVisual Video Summarization

2021-05-17 · Bin Zhao, Maoguo Gong, Xuelong Li

Audio and vision are two main modalities in video data. Multimodal learning, especially for audiovisual learning, has drawn considerable attention recently, which can boost the performance of various computer vision task…

Video Summarization

Audio-Visual Speaker Verification via Joint Cross-Attention

2023-09-28 · R. Gnana Praveen, Jahangir Alam

Speaker verification has been widely explored using speech signals, which has shown significant improvement using deep models. Recently, there has been a surge in exploring faces and voices as they can offer more complem…

Speaker Verification

Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic Scenarios

2024-06-21 · Ya Jiang, Qing Wang, Jun Du, Maocheng Hu 외

This study presents an audio-visual information fusion approach to sound event localization and detection (SELD) in low-resource scenarios. We aim at utilizing audio and video modality information through cross-modal lea…

Data AugmentationSound Event Localization and Detection

Audio Atlas: Visualizing and Exploring Audio Datasets

2024-11-30 · Luca A. Lanzendörfer, Florian Grötschla, Uzeyir Valizada, Roger Wattenhofer

We introduce Audio Atlas, an interactive web application for visualizing audio data using text-audio embeddings. Audio Atlas is designed to facilitate the exploration and analysis of audio datasets using a contrastive em…

Management