paper-with-me

Papers Dialect Identification

“Dialect Identification” 태그가 달린 논문 189편 · 필터 해제

Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification

2025-05-30 · Badr M. Abdullah, Matthew Baas, Bernd Möbius, Dietrich Klakow

Arabic dialect identification (ADI) systems are essential for large-scale data collection pipelines that enable the development of inclusive speech technologies for Arabic language varieties. However, the reliability of …

Dialect IdentificationVoice Conversion

Resource-Aware Arabic LLM Creation: Model Adaptation, Integration, and Multi-Domain Testing

2024-12-23 · Prakash Aryan

This paper presents a novel approach to fine-tuning the Qwen2-1.5B model for Arabic language processing using Quantized Low-Rank Adaptation (QLoRA) on a system with only 4GB VRAM. We detail the process of adapting this l…

ArabicMMLUDialect IdentificationGPULanguage Modeling+5

HiTZ at VarDial 2025 NorSID: Overcoming Data Scarcity with Language Transfer and Automatic Data Annotation

2024-12-13 · Jaione Bengoetxea, Mikel Zubillaga, Ekhi Azurmendi, Maite Heredia 외

In this paper we present our submission for the NorSID Shared Task as part of the 2025 VarDial Workshop (Scherrer et al., 2025), consisting of three tasks: Intent Detection, Slot Filling and Dialect Identification, evalu…

Dialect IdentificationIntent Detectionslot-fillingSlot Filling+1

Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and Challenges

2024-10-04 · Nguyen Van Dinh, Thanh Chi Dang, Luan Thanh Nguyen, Kiet Van Nguyen

Vietnamese, a low-resource language, is typically categorized into three primary dialect groups that belong to Northern, Central, and Southern Vietnam. However, each province within these regions exhibits its own distinc…

Dialect IdentificationDiversityMulti-Dialect Vietnamesespeech-recognition+3

AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs

2024-09-17 · Basel Mousi, Nadir Durrani, Fatema Ahmad, Md. Arid Hasan 외

Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects …

Dialect IdentificationDiversityMachine TranslationTranslation

Literary and Colloquial Dialect Identification for Tamil using Acoustic Features

2024-08-27 · M. Nanmalar, P. Vijayalakshmi, T. Nagarajan

The evolution and diversity of a language is evident from it's various dialects. If the various dialects are not addressed in technological advancements like automatic speech recognition and speech synthesis, there is a …

Automatic Speech RecognitionDialect IdentificationLanguage Identificationspeech-recognition+2

Literary and Colloquial Tamil Dialect Identification

2024-08-25 · M. Nanmalar, P. Vijayalakshmi, T. Nagarajan

Culture and language evolve together. The old literary form of Tamil is used commonly for writing and the contemporary colloquial Tamil is used for speaking. Human-computer interaction applications require Colloquial Tam…

Dialect Identificationspeech-recognitionSpeech Recognition

dzNLP at NADI 2024 Shared Task: Multi-Classifier Ensemble with Weighted Voting and TF-IDF Features

2024-07-18 · Mohamed Lichouri, Khaled Lounnas, Boualem Nadjib Zahaf, Mehdi Ayoub Rabiai

This paper presents the contribution of our dzNLP team to the NADI 2024 shared task, specifically in Subtask 1 - Multi-label Country-level Dialect Identification (MLDID) (Closed Track). We explored various configurations…

Dialect Identification

NADI 2024: The Fifth Nuanced Arabic Dialect Identification Shared Task

2024-07-06 · Muhammad Abdul-Mageed, Amr Keleg, AbdelRahim Elmadany, Chiyu Zhang 외

We describe the findings of the fifth Nuanced Arabic Dialect Identification Shared Task (NADI 2024). NADI's objective is to help advance SoTA Arabic NLP by providing guidance, datasets, modeling opportunities, and standa…

Dialect IdentificationMachine TranslationTranslationvalid

Exploiting Dialect Identification in Automatic Dialectal Text Normalization

2024-07-03 · Bashar Alhafni, Sarah Al-Towaity, Ziyad Fawzy, Fatema Nassar 외

Dialectal Arabic is the primary spoken language used by native Arabic speakers in daily communication. The rise of social media platforms has notably expanded its use as a written language. However, Arabic dialects do no…

Dialect IdentificationText Normalization

Exploring Energy-Based Models for Out-of-Distribution Detection in Dialect Identification

2024-06-26 · Yaqian Hao, Chenguang Hu, Yingying Gao, Shilei Zhang 외

The diverse nature of dialects presents challenges for models trained on specific linguistic patterns, rendering them susceptible to errors when confronted with unseen or out-of-distribution (OOD) data. This study introd…

Dialect IdentificationOut-of-Distribution Detection

Towards Zero-Shot Text-To-Speech for Arabic Dialects

2024-06-24 · Khai Duy Doan, Abdul Waheed, Muhammad Abdul-Mageed

Zero-shot multi-speaker text-to-speech (ZS-TTS) systems have advanced for English, however, it still lags behind due to insufficient resources. We address this gap for Arabic, a language of more than 450 million native s…

Dialect IdentificationSpeech Synthesistext-to-speechText to Speech

Low-resource speech recognition and dialect identification of Irish in a multi-task framework

2024-05-02 · Liam Lonergan, Mengjie Qian, Neasa Ní Chiaráin, Christer Gobl 외

This paper explores the use of Hybrid CTC/Attention encoder-decoder models trained with Intermediate CTC (InterCTC) for Irish (Gaelic) low-resource speech recognition (ASR) and dialect identification (DID). Results are c…

DecoderDialect IdentificationLanguage ModelingLanguage Modelling+2

Sebastian, Basti, Wastl?! Recognizing Named Entities in Bavarian Dialectal Data

2024-03-19 · Siyao Peng, Zihang Sun, Huangyan Shan, Marie Kolm 외

Named Entity Recognition (NER) is a fundamental task to extract key information from texts, but annotated resources are scarce for dialects. This paper introduces the first dialectal NER dataset for German, BarNER, with …

ArticlesDialect IdentificationDiversityMulti-Task Learning+4

USTHB at NADI 2023 shared task: Exploring Preprocessing and Feature Engineering Strategies for Arabic Dialect Identification

2023-12-16 · Mohamed Lichouri, Khaled Lounnas, Aicha Zitouni, Houda Latrache 외

In this paper, we conduct an in-depth analysis of several key factors influencing the performance of Arabic Dialect Identification NADI'2023, with a specific focus on the first subtask involving country-level dialect ide…

Dialect IdentificationFeature Engineering

Self-supervised Adaptive Pre-training of Multilingual Speech Models for Language and Dialect Identification

2023-12-12 · Mohammed Maqsood Shaik, Dietrich Klakow, Badr M. Abdullah

Pre-trained Transformer-based speech models have shown striking performance when fine-tuned on various downstream tasks such as automatic speech recognition and spoken language identification (SLID). However, the problem…

Automatic Speech RecognitionDialect IdentificationFew-Shot LearningLanguage Identification+3

Mavericks at NADI 2023 Shared Task: Unravelling Regional Nuances through Dialect Identification using Transformer-based Approach

2023-11-30 · Vedant Deshpande, Yash Patwardhan, Kshitij Deshpande, Sudeep Mangalvedhekar 외

In this paper, we present our approach for the "Nuanced Arabic Dialect Identification (NADI) Shared Task 2023". We highlight our methodology for subtask 1 which deals with country-level dialect identification. Recognizin…

Dialect IdentificationMulti-class Classificationspeech-recognitionSpeech Recognition

ArTST: Arabic Text and Speech Transformer

2023-10-25 · Hawau Olamide Toyin, Amirbek Djanibekov, Ajinkya Kulkarni, Hanan Aldarmaki

We present ArTST, a pre-trained Arabic text and speech transformer for supporting open-source speech technologies for the Arabic language. The model architecture follows the unified-modal framework, SpeechT5, that was re…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Dialect Identificationspeech-recognition+5

GlotLID: Language Identification for Low-Resource Languages

2023-10-24 · Amir Hossein Kargaran, Ayyoob Imani, François Yvon, Hinrich Schütze

Several recent papers have published good solutions for language identification (LID) for about 300 high-resource and medium-resource languages. However, there is no LID available that (i) covers a wide range of low-reso…

Dialect IdentificationLanguage Identification

NADI 2023: The Fourth Nuanced Arabic Dialect Identification Shared Task

2023-10-24 · Muhammad Abdul-Mageed, AbdelRahim Elmadany, Chiyu Zhang, El Moatez Billah Nagoudi 외

We describe the findings of the fourth Nuanced Arabic Dialect Identification Shared Task (NADI 2023). The objective of NADI is to help advance state-of-the-art Arabic NLP by creating opportunities for teams of researcher…

Dialect IdentificationMachine Translationvalid
1–20 / 189 다음 →