paper-with-me

Papers

Open-Set Language Identification

2017-07-16 · Shervin Malmasi

We present the first open-set language identification experiments using one-class classification. We first highlight the shortcomings of traditional feature extraction methods and propose a hashing-based feature vectorization approach as a solution. Using a dataset of 10 languages from different writing systems, we train a One- Class Support Vector Machine using only a monolingual corpus for each language. Each model is evaluated against a test set of data from all 10 languages and we achieve an average F-score of 0.99, highlighting the effectiveness of this approach for open-set language identification.

📄 PDF Abstract BibTeX arXiv:1707.04817

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationLanguage IdentificationOne-Class Classification

Similar Papers 제목 키워드 기반

Robust Open-Set Spoken Language Identification and the CU MultiLang Dataset

2023-08-29 · Mustafa Eyceoz, Justin Lee, Siddharth Pittie, Homayoon Beigi

Most state-of-the-art spoken language identification models are closed-set; in other words, they can only output a language label from the set of classes they were trained on. Open-set spoken language identification syst…

Language IdentificationSpoken language identification

VarClass: An Open-source Language Identification Tool for Language Varieties

2014-05-01 · LREC 2014 5 · Marcos Zampieri, Binyam Gebre

This paper presents VarClass, an open-source tool for language identification available both to be downloaded as well as through a graphical user-friendly interface. The main difference of VarClass in comparison to other…

Information RetrievalLanguage IdentificationMachine TranslationText Categorization

Modernizing Open-Set Speech Language Identification

2022-05-20 · Mustafa Eyceoz, Justin Lee, Homayoon Beigi

While most modern speech Language Identification methods are closed-set, we want to see if they can be modified and adapted for the open-set problem. When switching to the open-set problem, the solution gains the ability…

Language IdentificationSpeech Language Identification

A reproduction of Apple's bi-directional LSTM models for language identification in short strings

2021-02-11 · EACL 2021 2 · Mads Toftrup, Søren Asger Sørensen, Manuel R. Ciosici, Ira Assent

Language Identification is the task of identifying a document's language. For applications like automatic spell checker selection, language identification must use very short strings such as text message fragments. In th…

Language Identification

OpenLID-v3: Improving the Precision of Closely Related Language Identification -- An Experience Report

2026-02-13 · Mariia Fedorova, Nikolay Arefyev, Maja Buljan, Jindřich Helcl 외 arxiv

Language identification (LID) is an essential step in building high-quality multilingual datasets from web data. Existing LID tools (such as OpenLID or GlotLID) often struggle to identify closely related languages and to…

Language Identification