paper-with-me

홈 › Papers

OxfordVGG Submission to the EGO4D AV Transcription Challenge

2023-07-18 · Jaesung Huh, Max Bain, Andrew Zisserman

This report presents the technical details of our submission on the EGO4D Audio-Visual (AV) Automatic Speech Recognition Challenge 2023 from the OxfordVGG team. We present WhisperX, a system for efficient speech transcription of long-form audio with word-level time alignment, along with two text normalisers which are publicly available. Our final submission obtained 56.0% of the Word Error Rate (WER) on the challenge test set, ranked 1st on the leaderboard. All baseline codes and models are available on https://github.com/m-bain/whisperX.

📄 PDF Abstract BibTeX arXiv:2307.09006

Code (1)

m-bain/whisperx 공식 구현 pytorch

Tasks

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

AVATAR submission to the Ego4D AV Transcription Challenge

2022-11-18 · Paul Hongsuck Seo, Arsha Nagrani, Cordelia Schmid

In this report, we describe our submission to the Ego4D AudioVisual (AV) Speech Transcription Challenge 2022. Our pipeline is based on AVATAR, a state of the art encoder-decoder model for AV-ASR that performs early fusio…

Decoder

CUED_speech at TREC 2020 Podcast Summarisation Track

2020-12-04 · Potsawee Manakul, Mark Gales

In this paper, we describe our approach for the Podcast Summarisation challenge in TREC 2020. Given a podcast episode with its transcription, the goal is to generate a summary that captures the most important information…

Self-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling

2022-02-07 · Puyuan Peng, David Harwath

In this paper, we describe our submissions to the ZeroSpeech 2021 Challenge and SUPERB benchmark. Our submissions are based on the recently proposed FaST-VGS model, which is a Transformer-based model that learns to assoc…

Language ModelingLanguage ModellingMasked Language ModelingRepresentation Learning+1

QiNiAn at SemEval-2022 Task 5: Multi-Modal Misogyny Detection and Classification

2022-07-01 · SemEval (NAACL) 2022 7 · Qin Gu, Nino Meisinger, Anna-Katharina Dick

In this paper, we describe our submission to the misogyny classification challenge at SemEval-2022. We propose two models for the two subtasks of the challenge: The first uses joint image and text classification to class…

ClassificationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONtext-classification+1

ON-TRAC Consortium Systems for the IWSLT 2022 Dialect and Low-resource Speech Translation Tasks

2022-05-04 · IWSLT (ACL) 2022 5 · Marcely Zanon Boito, John Ortega, Hugo Riguidel, Antoine Laurent 외

This paper describes the ON-TRAC Consortium translation systems developed for two challenge tracks featured in the Evaluation Campaign of IWSLT 2022: low-resource and dialect speech translation. For the Tunisian Arabic-E…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Machine Translationspeech-recognition+3