paper-with-me

Papers

WeCanTalk: A New Multi-language, Multi-modal Resource for Speaker Recognition

2022-06-01 · LREC 2022 6 · Karen Jones, Kevin Walker, Christopher Caruso, Jonathan Wright, Stephanie Strassel

The WeCanTalk (WCT) Corpus is a new multi-language, multi-modal resource for speaker recognition. The corpus contains Cantonese, Mandarin and English telephony and video speech data from over 200 multilingual speakers located in Hong Kong. Each speaker contributed at least 10 telephone conversations of 8-10 minutes’ duration collected via a custom telephone platform based in Hong Kong. Speakers also uploaded at least 3 videos in which they were both speaking and visible, along with one selfie image. At least half of the calls and videos for each speaker were in Cantonese, while their remaining recordings featured one or more different languages. Both calls and videos were made in a variety of noise conditions. All speech and video recordings were audited by experienced multilingual annotators for quality including presence of the expected language and for speaker identity. The WeCanTalk Corpus has been used to support the NIST 2021 Speaker Recognition Evaluation and will be published in the LDC catalog.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Speaker Recognition

Similar Papers 제목 키워드 기반

The 2021 NIST Speaker Recognition Evaluation

2022-04-21 · Seyed Omid Sadjadi, Craig Greenberg, Elliot Singer, Lisa Mason 외

The 2021 Speaker Recognition Evaluation (SRE21) was the latest cycle of the ongoing evaluation series conducted by the U.S. National Institute of Standards and Technology (NIST) since 1996. It was the second large-scale …

Data AugmentationFace RecognitionPerson RecognitionSpeaker Recognition

Towards automation in using multi-modal language resources: compatibility and interoperability for multi-modal features in Kachako

2012-05-01 · LREC 2012 5 · Yoshinobu Kano

Use of language resources including annotated corpora and tools is not easy for users, as it requires expert knowledge to determine which resources are compatible and interoperable. Sometimes it requires programming skil…

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages

2026-05-16 · Firoj Alam, Shammur Absar Chowdhury, Enamul Hoque Prince arxiv

Multimodal LLMs are evolving from vision-language to tri-modality that see, hear, and read, yet pipelines and benchmarks remain English-centric and compute-heavy. The tutorial offers an overview of this emerging research…

DravidianMultiModality: A Dataset for Multi-modal Sentiment Analysis in Tamil and Malayalam

2021-06-09 · Bharathi Raja Chakravarthi, Jishnu Parameswaran P. K, Premjith B, K. P Soman 외

Human communication is inherently multimodal and asynchronous. Analyzing human emotions and sentiment is an emerging field of artificial intelligence. We are witnessing an increasing amount of multimodal content in local…

Multimodal Sentiment AnalysisSentiment Analysis

Constructing Multimodal Datasets from Scratch for Rapid Development of a Japanese Visual Language Model

2024-10-30 · Keito Sasagawa, Koki Maeda, Issa Sugiura, Shuhei Kurita 외

To develop high-performing Visual Language Models (VLMs), it is essential to prepare multimodal resources, such as image-text pairs, interleaved data, and instruction data. While multimodal resources for English are abun…

Language ModelingLanguage Modelling