paper-with-me

홈 › Papers

A large-scale multimodal dataset of human speech recognition

2023-03-15 · Yao Ge, Chong Tang, Haobo Li, Zikang Zhang, Wenda Li, Kevin Chetty, Daniele Faccio, Qammer H. Abbasi, Muhammad Imran

Nowadays, non-privacy small-scale motion detection has attracted an increasing amount of research in remote sensing in speech recognition. These new modalities are employed to enhance and restore speech information from speakers of multiple types of data. In this paper, we propose a dataset contains 7.5 GHz Channel Impulse Response (CIR) data from ultra-wideband (UWB) radars, 77-GHz frequency modulated continuous wave (FMCW) data from millimetre wave (mmWave) radar, and laser data. Meanwhile, a depth camera is adopted to record the landmarks of the subject's lip and voice. Approximately 400 minutes of annotated speech profiles are provided, which are collected from 20 participants speaking 5 vowels, 15 words and 16 sentences. The dataset has been validated and has potential for the research of lip reading and multimodal speech recognition.

📄 PDF Abstract BibTeX arXiv:2303.08295

Code (0)

등록된 구현이 없습니다.

Tasks

Lip ReadingMotion Detectionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

A Large-Scale Chinese Multimodal NER Dataset with Speech Clues

2021-08-01 · ACL 2021 5 · Dianbo Sui, Zhengkun Tian, Yubo Chen, Kang Liu 외

In this paper, we aim to explore an uncharted territory, which is Chinese multimodal named entity recognition (NER) with both textual and acoustic contents. To achieve this, we construct a large-scale human-annotated Chi…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER+1

OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA

2025-10-07 · Firoj Alam, Ali Ezzat Shahroor, Md. Arid Hasan, Zien Sheikh Ali 외 arxiv

Large-scale multimodal models achieve strong results on tasks like Visual Question Answering (VQA), but they are often limited when queries require cultural and visual information, everyday knowledge, particularly in low…

Visual Question AnsweringObject Recognition

MIKU-PAL: An Automated and Standardized Multi-Modal Method for Speech Paralinguistic and Affect Labeling

2025-05-21 · Cheng Yifan, Zhang Ruoyi, Shi Jiatong

Acquiring large-scale emotional speech data with strong consistency remains a challenge for speech synthesis. This paper presents MIKU-PAL, a fully automated multimodal pipeline for extracting high-consistency emotional …

Emotion RecognitionFace DetectionLanguage ModelingLanguage Modelling+6

TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation

2025-10-08 · Jiaben Chen, Zixin Wang, Ailing Zeng, Yang Fu 외 arxiv

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offer…

Video Generation

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness

2026-07-09 · Iulia-Maria Udrea, Alexandra Diaconu, Bogdan Alexe arxiv

We introduce VSRo-200, the first large-scale dataset for visual speech recognition (lip reading) in Romanian, comprising 200 hours of real-world podcast videos. All samples are annotated with pseudo-labels generated by a…

Audio-Visual Speech RecognitionDomain GeneralizationLip Reading