SAVAS: Collecting, Annotating and Sharing Audiovisual Language Resources for Automatic Subtitling
This paper describes the data collection, annotation and sharing activities carried out within the FP7 EU-funded SAVAS project. The project aims to collect, share and reuse audiovisual language resources from broadcasters and subtitling companies to develop large vocabulary continuous speech recognisers in specific domains and new languages, with the purpose of solving the automated subtitling needs of the media industry.
Code (0)
등록된 구현이 없습니다.
Tasks
Language ModellingSpeech RecognitionSimilar Papers 제목 키워드 기반
ESSYS* Sharing #UC: An Emotion-driven Audiovisual Installation
We present ESSYS* Sharing #UC, an audiovisual installation artwork that reflects upon the emotional context related to the university and the city of Coimbra, based on the data shared about them on Twitter. The installat…
DiversityRLDS: an Ecosystem to Generate, Share and Use Datasets in Reinforcement Learning
We introduce RLDS (Reinforcement Learning Datasets), an ecosystem for recording, replaying, manipulating, annotating and sharing data in the context of Sequential Decision Making (SDM) including Reinforcement Learning (R…
Decision MakingImitation LearningOffline RLreinforcement-learning+3A horizon line annotation tool for streamlining autonomous sea navigation experiments
Horizon line (or sea line) detection (HLD) is a critical component in multiple marine autonomous navigation tasks, such as identifying the navigation area (i.e., the sea), obstacle detection and geo-localization, and dig…
Autonomous Navigationgeo-localizationLine DetectionPosition+1Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing
Explainable NLP (ExNLP) has increasingly focused on collecting human-annotated textual explanations. These explanations are used downstream in three ways: as data augmentation to improve performance on a predictive task,…
Data AugmentationLocation-based Twitter Filtering for the Creation of Low-Resource Language Datasets in Indonesian Local Languages
Twitter contains an abundance of linguistic data from the real world. We examine Twitter for user-generated content in low-resource languages such as local Indonesian. For NLP to work in Indonesian, it must consider loca…
Cultural Vocal Bursts Intensity Prediction