paper-with-me

홈 › Papers

Automatic Token and Turn Level Language Identification for Code-Switched Text Dialog: An Analysis Across Language Pairs and Corpora

2018-07-01 · WS 2018 7 · Vikram Ramanarayanan, Robert Pugh

We examine the efficacy of various feature{--}learner combinations for language identification in different types of text-based code-switched interactions {--} human-human dialog, human-machine dialog as well as monolog {--} at both the token and turn levels. In order to examine the generalization of such methods across language pairs and datasets, we analyze 10 different datasets of code-switched text. We extract a variety of character- and word-based text features and pass them into multiple learners, including conditional random fields, logistic regressors and recurrent neural networks. We further examine the efficacy of novel character-level embedding and GloVe features in improving performance and observe that our best-performing text system significantly outperforms a majority vote baseline across language pairs and datasets.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationSpoken Language UnderstandingText Generation

Methods 이 논문이 사용한 방법론

GloVe GloVe Embeddings are a type of word embedding that encode the co-occurrence probability ratio between two words as vector differences. GloVe uses a weighted least squares…

Similar Papers 제목 키워드 기반

Slang Detection and Identification

2019-11-01 · CONLL 2019 11 · Zhengqi Pei, Zhewei Sun, Yang Xu

The prevalence of informal language such as slang presents challenges for natural language systems, particularly in the automatic discovery of flexible word usages. Previous work has explored slang in terms of dictionary…

SentenceSentiment Analysis

Automatic Arabic Dialect Identification Systems for Written Texts: A Survey

2020-09-26 · Maha J. Althobaiti

Arabic dialect identification is a specific task of natural language processing, aiming to automatically predict the Arabic dialect of a given text. Arabic dialect identification is the first step in various natural lang…

Dialect IdentificationMachine TranslationSentenceSpeech Synthesis+6

Improving low-resource ASR using bilingual fine-tuning with language identification: a cross-linguistic evaluation

2026-06-16 · Reihaneh Amooie, Yun Hao, Wietse de Vries, Jelske Dijkstra 외 arxiv

This study explores how bilingual fine-tuning affects automatic speech recognition (ASR) in low-resource languages. We evaluate this method across nine linguistically and geographically diverse language pairs, covering a…

Language IdentificationSpeech Recognition

Unified model for code-switching speech recognition and language identification based on a concatenated tokenizer

2023-06-14 · Kunal Dhawan, Dima Rekesh, Boris Ginsburg

Code-Switching (CS) multilingual Automatic Speech Recognition (ASR) models can transcribe speech containing two or more alternating languages during a conversation. This paper proposes (1) a new method for creating code-…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Identificationspeech-recognition+2

Don't Believe Everything You Read: Enhancing Summarization Interpretability through Automatic Identification of Hallucinations in Large Language Models

2023-12-22 · Priyesh Vakharia, Devavrat Joshi, Meenal Chavan, Dhananjay Sonawane 외

Large Language Models (LLMs) are adept at text manipulation -- tasks such as machine translation and text summarization. However, these models can also be prone to hallucination, which can be detrimental to the faithfuln…

HallucinationMachine TranslationText SummarizationTranslation