paper-with-me

LRS2

Lip Reading Sentences 2

홈페이지 · 논문 115편

The Oxford-BBC Lip Reading Sentences 2 (LRS2) dataset is one of the largest publicly available datasets for lip reading sentences in-the-wild. The database consists of mainly news and talk shows from BBC programs. Each sentence is up to 100 characters in length. The training, validation and test sets are divided according to broadcast date. It is a challenging set since it contains thousands of speakers without speaker labels and large variation in head pose. The pre-training set contains 96,318 utterances, the training set contains 45,839 utterances, the validation set contains 1,082 utterances and the test set contains 1,242 utterances. Source: Audio-visual Recognition of Overlapped speech for the LRS2 dataset Image Source: https://www.robots.ox.ac.uk/~vgg/data/lip_reading/lrs2.html

VideosTextsAudio

벤치마크

Lipreading on LRS2 결과 50개
Unconstrained Lip-synchronization on LRS2 결과 27개
Automatic Speech Recognition (ASR) on LRS2 결과 18개
Audio-Visual Speech Recognition on LRS2 결과 8개
Speech Separation on LRS2 결과 8개
Visual Speech Recognition on LRS2 결과 6개
Image Manipulation on LRS2 결과 4개
Landmark-based Lipreading on LRS2 결과 2개
Speech Recognition on LRS2 결과 1개
Visual Keyword Spotting on LRS2 결과 1개