paper-with-me

Papers

VoxVietnam: a Large-Scale Multi-Genre Dataset for Vietnamese Speaker Recognition

2024-12-31 · Hoang Long Vu, Phuong Tuan Dat, Pham Thao Nhi, Nguyen Song Hao, Nguyen Thi Thu Trang

Recent research in speaker recognition aims to address vulnerabilities due to variations between enrolment and test utterances, particularly in the multi-genre phenomenon where the utterances are in different speech genres. Previous resources for Vietnamese speaker recognition are either limited in size or do not focus on genre diversity, leaving studies in multi-genre effects unexplored. This paper introduces VoxVietnam, the first multi-genre dataset for Vietnamese speaker recognition with over 187,000 utterances from 1,406 speakers and an automated pipeline to construct a dataset on a large scale from public sources. Our experiments show the challenges posed by the multi-genre phenomenon to models trained on a single-genre dataset, and demonstrate a significant increase in performance upon incorporating the VoxVietnam into the training process. Our experiments are conducted to study the challenges of the multi-genre phenomenon in speaker recognition and the performance gain when the proposed dataset is used for multi-genre training.

📄 PDF Abstract BibTeX arXiv:2501.00328

Code (0)

등록된 구현이 없습니다.

Tasks

DiversitySpeaker Recognition

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

MATT: A Multiple-instance Attention Mechanism for Long-tail Music Genre Classification

2022-09-09 · Xiaokai Liu, Menghua Zhang

Imbalanced music genre classification is a crucial task in the Music Information Retrieval (MIR) field for identifying the long-tail, data-poor genre based on the related music audio segments, which is very prevalent in …

Genre classificationInformation RetrievalMusic Genre ClassificationMusic Information Retrieval+1

GCDT: A Chinese RST Treebank for Multigenre and Multilingual Discourse Parsing

2022-10-19 · Siyao Peng, Yang Janet Liu, Amir Zeldes

A lack of large-scale human-annotated data has hampered the hierarchical discourse parsing of Chinese. In this paper, we present GCDT, the largest hierarchical discourse treebank for Mandarin Chinese in the framework of …

Discourse Parsing

ALHD: A Large-Scale and Multigenre Benchmark Dataset for Arabic LLM-Generated Text Detection

2025-10-03 · Ali Khairallah, Arkaitz Zubiaga arxiv

We introduce ALHD, the first large-scale comprehensive Arabic dataset explicitly designed to distinguish between human- and LLM-generated texts. ALHD spans three genres (news, social media, reviews), covering both MSA an…

Text Detection

GenRec: Unifying Video Generation and Recognition with Diffusion Models

2024-08-27 · Zejia Weng, Xitong Yang, Zhen Xing, Zuxuan Wu 외

Video diffusion models are able to generate high-quality videos by learning strong spatial-temporal priors on large-scale datasets. In this paper, we aim to investigate whether such priors derived from a generative proce…

Image to Video GenerationVideo GenerationVideo Recognition

On large-scale genre classification in symbolically encoded music by automatic identification of repeating patterns

2019-10-21 · Andres Ferraro, Kjell Lemström

The importance of repetitions in music is well-known. In this paper, we study music repetitions in the context of effective and efficient automatic genre classification in large-scale music-databases. We aim at enhancing…

General ClassificationGenre classification