paper-with-me

홈 › Papers

AVATAR submission to the Ego4D AV Transcription Challenge

2022-11-18 · Paul Hongsuck Seo, Arsha Nagrani, Cordelia Schmid

In this report, we describe our submission to the Ego4D AudioVisual (AV) Speech Transcription Challenge 2022. Our pipeline is based on AVATAR, a state of the art encoder-decoder model for AV-ASR that performs early fusion of spectrograms and RGB images. We describe the datasets, experimental settings and ablations. Our final method achieves a WER of 68.40 on the challenge test set, outperforming the baseline by 43.7%, and winning the challenge.

📄 PDF Abstract BibTeX arXiv:2211.09966

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Methods 이 논문이 사용한 방법론

Test 설명 없음

Similar Papers 제목 키워드 기반

OxfordVGG Submission to the EGO4D AV Transcription Challenge

2023-07-18 · Jaesung Huh, Max Bain, Andrew Zisserman

This report presents the technical details of our submission on the EGO4D Audio-Visual (AV) Automatic Speech Recognition Challenge 2023 from the OxfordVGG team. We present WhisperX, a system for efficient speech transcri…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

CUED_speech at TREC 2020 Podcast Summarisation Track

2020-12-04 · Potsawee Manakul, Mark Gales

In this paper, we describe our approach for the Podcast Summarisation challenge in TREC 2020. Given a podcast episode with its transcription, the goal is to generate a summary that captures the most important information…

9th Workshop on Sign Language Translation and Avatar Technologies (SLTAT 2025)

2025-08-11 · Fabrizio Nunnari, Cristina Luna Jiménez, Rosalee Wolfe, John C. McDonald 외 arxiv

The Sign Language Translation and Avatar Technology (SLTAT) workshops continue a series of gatherings to share recent advances in improving deaf / human communication through non-invasive means. This 2025 edition, the 9t…

Sign Language TranslationSign Language Recognition

Self-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling

2022-02-07 · Puyuan Peng, David Harwath

In this paper, we describe our submissions to the ZeroSpeech 2021 Challenge and SUPERB benchmark. Our submissions are based on the recently proposed FaST-VGS model, which is a Transformer-based model that learns to assoc…

Language ModelingLanguage ModellingMasked Language ModelingRepresentation Learning+1

Tools for the Use of SignWriting as a Language Resource

2020-05-01 · LREC 2020 5 · Antonio F. G. Sevilla, Alberto D{\'\i}az Esteban, Jos{\'e} Mar{\'\i}a Lahoz-Bengoechea

Representation of linguistic data is an issue of utmost importance when developing language resources, but the lack of a standard written form in sign languages presents a challenge. Different notation systems exist, but…