paper-with-me

Papers

Multilingual Audio-Visual Smartphone Dataset And Evaluation

2021-09-09 · Hareesh Mandalapu, Aravinda Reddy P N, Raghavendra Ramachandra, K Sreenivasa Rao, Pabitra Mitra, S R Mahadeva Prasanna, Christoph Busch

Smartphones have been employed with biometric-based verification systems to provide security in highly sensitive applications. Audio-visual biometrics are getting popular due to their usability, and also it will be challenging to spoof because of their multimodal nature. In this work, we present an audio-visual smartphone dataset captured in five different recent smartphones. This new dataset contains 103 subjects captured in three different sessions considering the different real-world scenarios. Three different languages are acquired in this dataset to include the problem of language dependency of the speaker recognition systems. These unique characteristics of this dataset will pave the way to implement novel state-of-the-art unimodal or audio-visual speaker recognition systems. We also report the performance of the bench-marked biometric verification systems on our dataset. The robustness of biometric algorithms is evaluated towards multiple dependencies like signal noise, device, language and presentation attacks like replay and synthesized signals with extensive experiments. The obtained results raised many concerns about the generalization properties of state-of-the-art biometrics methods in smartphones.

📄 PDF Abstract BibTeX arXiv:2109.04138

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Recognition

Similar Papers 제목 키워드 기반

EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$

2026-08-24 · Jiawen Wang, Xiaoxue Gao, Zi Haur Pang, Nancy F. Chen arxiv

Recent large audio language models (LALMs) have achieved impressive progress in audio understanding. However, existing evaluations remain largely constrained to English and narrow audio domains. Prior benchmarks typicall…

mWhisper-Flamingo for Multilingual Audio-Visual Noise-Robust Speech Recognition

2025-02-03 · Andrew Rouditchenko, Samuel Thomas, Hilde Kuehne, Rogerio Feris 외

Audio-Visual Speech Recognition (AVSR) combines lip-based video with audio and can improve performance in noise, but most methods are trained only on English data. One limitation is the lack of large-scale multilingual v…

Audio-Visual Speech RecognitionDecoderRobust Speech Recognitionspeech-recognition+2

Face-voice Association in Multilingual Environments (FAME) Challenge 2024 Evaluation Plan

2024-04-14 · Muhammad Saad Saeed, Shah Nawaz, Muhammad Salman Tahir, Rohan Kumar Das 외

The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, the audio-visual systems are one of the widely used multimodal systems. In the recent years, associ…

Face-voice Association in Multilingual Environments (FAME) 2026 Challenge Evaluation Plan

2025-08-06 · Marta Moscati, Ahmed Abdullah, Muhammad Saad Saeed, Shah Nawaz 외 arxiv

The advancements of technology have led to the use of multimodal systems in various real-world applications. Among them, audio-visual systems are among the most widely used multimodal systems. In the recent years, associ…

Knowledge Distillation for Efficient Audio-Visual Video Captioning

2023-06-16 · Özkan Çaylı, Xubo Liu, Volkan Kılıç, Wenwu Wang

Automatically describing audio-visual content with texts, namely video captioning, has received significant attention due to its potential applications across diverse fields. Deep neural networks are the dominant methods…

Audio-Visual Video CaptioningCaption GenerationKnowledge DistillationVideo Captioning