paper-with-me

홈 › Papers

SUTAV: A Turkish Audio-Visual Database

2012-05-01 · LREC 2012 5 · Ibrahim Saygin Topkaya, Hakan Erdogan

This paper contains information about the ''''''`Sabanci University Turkish Audio-Visual (SUTAV)'''''''' database. The main aim of collecting SUTAV database was to obtain a large audio-visual collection of spoken words, numbers and sentences in Turkish language. The database was collected between 2006 and 2010 during ''''''`Novel approaches in audio-visual speech recognition'''''''' project which is funded by The Scientific and Technological Research Council of Turkey (TUBITAK). First part of the database contains a large corpus of Turkish language and contains standart quality videos. The second part is relatively small compared to the first one and contains recordings of spoken digits in high quality videos. Although the main aim to collect SUTAV database was to obtain a database for audio-visual speech recognition applications, it also contains useful data that can be used in other kinds of multimodal research like biometric security and person verification. The paper presents information about the data collection process and the the spoken content. It also contains a sample evaluation protocol and recognition results that are obtained with a small portion of the database.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual Speech RecognitionPerson Identificationspeech-recognitionSpeech RecognitionVisual Speech Recognition

Similar Papers 제목 키워드 기반

Morpholex Turkish: A Morphological Lexicon for Turkish

2022-06-01 · gwll (LREC) 2022 6 · Bilge Arican, Aslı Kuzgun, Büşra Marşan, Deniz Baran Aslan 외

MorphoLex is a study in which root, prefix and suffixes of words are analyzed. With MorphoLex, many words can be analyzed according to certain rules and a useful database can be created. Due to the fact that Turkish is a…

How Does Audio Influence Visual Attention in Omnidirectional Videos? Database and Model

2024-08-10 · Yuxin Zhu, Huiyu Duan, Kaiwei Zhang, Yucheng Zhu 외

Understanding and predicting viewer attention in omnidirectional videos (ODVs) is crucial for enhancing user engagement in virtual and augmented reality applications. Although both audio and visual modalities are essenti…

PredictionSaliency Prediction

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction

2026-08-22 · Ahmet Tuğrul Bayrak, Fatma Nur Korkmaz, Bekir Berker Türker, Mustafa Sertaç Türkel 외 hf

Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchronous dialog systems. While existing research has explored multimodal approaches and large language mod…

Binary Classification

Turkish Emotion Voice Database (TurEV-DB)

2020-05-01 · LREC 2020 5 · Salih Firat Canpolat, Zuhal Ormano{\u{g}}lu, Deniz Zeyrek

We introduce the Turkish Emotion-Voice Database (TurEV-DB) which involves a corpus of over 1700 tokens based on 82 words uttered by human subjects in four different emotions (\textit{angry, calm, happy, sad}). Three mach…

BIG-bench Machine Learning

The AV-LASYN Database : A synchronous corpus of audio and 3D facial marker data for audio-visual laughter synthesis

2014-05-01 · LREC 2014 5 · H{\"u}seyin {\c{C}}akmak, J{\'e}r{\^o}me Urbain, Thierry Dutoit, Jo{\"e}lle Tilmanne

A synchronous database of acoustic and 3D facial marker data was built for audio-visual laughter synthesis. Since the aim is to use this database for HMM-based modeling and synthesis, the amount of collected data from on…

Dimensionality ReductionSpeech Synthesis