paper-with-me

Papers

ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition

2025-06-05 · Thai-Binh Nguyen, Thi Van Nguyen, Quoc Truong Do, Chi Mai Luong

Audio-Visual Speech Recognition (AVSR) has gained significant attention recently due to its robustness against noise, which often challenges conventional speech recognition systems that rely solely on audio features. Despite this advantage, AVSR models remain limited by the scarcity of extensive datasets, especially for most languages beyond English. Automated data collection offers a promising solution. This work presents a practical approach to generate AVSR datasets from raw video, refining existing techniques for improved efficiency and accessibility. We demonstrate its broad applicability by developing a baseline AVSR model for Vietnamese. Experiments show the automatically collected dataset enables a strong baseline, achieving competitive performance with robust ASR in clean conditions and significantly outperforming them in noisy environments like cocktail parties. This efficient method provides a pathway to expand AVSR to more languages, particularly under-resourced ones.

📄 PDF Abstract BibTeX arXiv:2506.04635

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual Speech Recognitionspeech-recognitionSpeech RecognitionVisual Speech Recognition

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

An archaeological Catalog Collection Method Based on Large Vision-Language Models

2024-12-28 · Honglin Pang, Yi Chang, Tianjing Duan, Xi Yang

Archaeological catalogs, containing key elements such as artifact images, morphological descriptions, and excavation information, are essential for studying artifact evolution and cultural inheritance. These data are wid…

Large Multi-modal Model Cartographic Map Comprehension for Textual Locality Georeferencing

2025-07-11 · Kalana Wijegunarathna, Kristin Stock, Christopher B. Jones arxiv

Millions of biological sample records collected in the last few centuries archived in natural history collections are un-georeferenced. Georeferencing complex locality descriptions associated with these collection sample…

Challenges and Applications of Automated Extraction of Socio-political Events from Text (CASE 2023): Workshop and Shared Task Report

2023-12-02 · Ali Hürriyetoğlu, Hristo Tanev, Osman Mutlu, Surendrabikram Thapa 외

We provide a summary of the sixth edition of the CASE workshop that is held in the scope of RANLP 2023. The workshop consists of regular papers, three keynotes, working papers of shared task participants, and shared task…

Event Extraction

Towers of Babel: Combining Images, Language, and 3D Geometry for Learning Multimodal Vision

2021-08-12 · ICCV 2021 10 · Xiaoshi Wu, Hadar Averbuch-Elor, Jin Sun, Noah Snavely

The abundance and richness of Internet photos of landmarks and cities has led to significant progress in 3D vision over the past two decades, including automated 3D reconstructions of the world's landmarks from tourist p…

3D geometryDescriptiveImage CaptioningMultimodal Reasoning

How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space

2026-08-27 · Corey D. C. Heath arxiv

Aesthetics are an important part of the symbolism of artistic works. Although subjective, humans categorize art based on the emotion evoked regardless of modality. What remains under-explored is how AI models form their …