paper-with-me

Papers

AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies

2024-02-20 · José-M. Acosta-Triana, David Gimeno-Gómez, Carlos-D. Martínez-Hinarejos

More than 7,000 known languages are spoken around the world. However, due to the lack of annotated resources, only a small fraction of them are currently covered by speech technologies. Albeit self-supervised speech representations, recent massive speech corpora collections, as well as the organization of challenges, have alleviated this inequality, most studies are mainly benchmarked on English. This situation is aggravated when tasks involving both acoustic and visual speech modalities are addressed. In order to promote research on low-resource languages for audio-visual speech technologies, we present AnnoTheia, a semi-automatic annotation toolkit that detects when a person speaks on the scene and the corresponding transcription. In addition, to show the complete process of preparing AnnoTheia for a language of interest, we also describe the adaptation of a pre-trained model for active speaker detection to Spanish, using a database not initially conceived for this type of task. The AnnoTheia toolkit, tutorials, and pre-trained models are available on GitHub.

📄 PDF Abstract BibTeX arXiv:2402.13152

Code (1)

joactr/annotheia 공식 구현 pytorch

Tasks

Active Speaker Detection

Similar Papers 제목 키워드 기반

GazeBehavior Annotation Toolkit (GBAT): AI-powered toolkit for automatic annotation of egocentric eye-tracking and video data of child-caregiver interaction

2026-05-21 · Iba Baig, Kevin Li, Yanbin Xu, Seiji Cattelain 외 arxiv

Video recordings of child-caregiver interactions enable investigation of attentional dynamics during naturalistic behavior. Such multimodal recording also allows researchers to examine how attention interacts with action…

Semi-automatic 3D Object Keypoint Annotation and Detection for the Masses

2022-01-19 · Kenneth Blomqvist, Jen Jen Chung, Lionel Ott, Roland Siegwart

Creating computer vision datasets requires careful planning and lots of time and effort. In robotics research, we often have to use standardized objects, such as the YCB object set, for tasks such as object tracking, pos…

ObjectObject TrackingPose Estimation

SANTLR: Speech Annotation Toolkit for Low Resource Languages

2019-08-02 · Xinjian Li, Zhong Zhou, Siddharth Dalmia, Alan W. black 외

While low resource speech recognition has attracted a lot of attention from the speech community, there are a few tools available to facilitate low resource speech collection. In this work, we present SANTLR: Speech Anno…

speech-recognitionSpeech Recognition

SLMotion - An extensible sign language oriented video analysis tool

2014-05-01 · LREC 2014 5 · Matti Karppa, Ville Viitaniemi, Marcos Luzardo, Jorma Laaksonen 외

We present a software toolkit called SLMotion which provides a framework for automatic and semiautomatic analysis, feature extraction and annotation of individual sign language videos, and which can easily be adapted to …

Sign Language Recognition

PyMIC: A deep learning toolkit for annotation-efficient medical image segmentation

2022-08-19 · Guotai Wang, Xiangde Luo, Ran Gu, Shuojue Yang 외

Background and Objective: Open-source deep learning toolkits are one of the driving forces for developing medical image segmentation models. Existing toolkits mainly focus on fully supervised segmentation and require ful…

Deep LearningImage SegmentationMedical Image SegmentationSegmentation+2