paper-with-me

Papers

Automatic Language Identification for Celtic Texts

2022-03-09 · Olha Dovbnia, Anna Wróblewska

Language identification is an important Natural Language Processing task. It has been thoroughly researched in the literature. However, some issues are still open. This work addresses the identification of the related low-resource languages on the example of the Celtic language family. This work's main goals were: (1) to collect the dataset of three Celtic languages; (2) to prepare a method to identify the languages from the Celtic family, i.e. to train a successful classification model; (3) to evaluate the influence of different feature extraction methods, and explore the applicability of the unsupervised models as a feature extraction technique; (4) to experiment with the unsupervised feature extraction on a reduced annotated set. We collected a new dataset including Irish, Scottish, Welsh and English records. We tested supervised models such as SVM and neural networks with traditional statistical features alongside the output of clustering, autoencoder, and topic modelling methods. The analysis showed that the unsupervised features could serve as a valuable extension to the n-gram feature vectors. It led to an improvement in performance for more entangled classes. The best model achieved a 98\% F1 score and 97\% MCC. The dense neural network consistently outperformed the SVM model. The low-resource languages are also challenging due to the scarcity of available annotated training data. This work evaluated the performance of the classifiers using the unsupervised feature extraction on the reduced labelled dataset to handle this issue. The results uncovered that the unsupervised feature vectors are more robust to the labelled set reduction. Therefore, they proved to help achieve comparable classification performance with much less labelled data.

📄 PDF Abstract BibTeX arXiv:2203.04831

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identification

Methods 이 논문이 사용한 방법론

SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

Neural Models for Predicting Celtic Mutations

2020-05-01 · LREC 2020 5 · Kevin Scannell

The Celtic languages share a common linguistic phenomenon known as initial mutations; these consist of pronunciation and spelling changes that occur at the beginning of some words, triggered in certain semantic or syntac…

Grammatical Error DetectionLanguage ModelingLanguage Modelling

Multilingual Abstract Meaning Representation for Celtic Languages

2022-06-01 · CLTW (LREC) 2022 6 · Johannes Heinecke, Anastasia Shimorina

Deep Semantic Parsing into Abstract Meaning Representation (AMR) graphs has reached a high quality with neural-based seq2seq approaches. However, the training corpus for AMR is only available for English. Several approac…

Abstract Meaning RepresentationSemantic Parsing

Automatic clustering of Celtic coins based on 3D point cloud pattern analysis

2020-05-12 · Sofiane Horache, Jean-Emmanuel Deschaud, François Goulette, Katherine Gruel 외

The recognition and clustering of coins which have been struck by the same die is of interest for archeological studies. Nowadays, this work can only be performed by experts and is very tedious. In this paper, we propose…

Clustering

Subsegmental language detection in Celtic language text

2014-08-01 · WS 2014 8 · Akshay Minocha, Francis Tyers
Language IdentificationLanguage ModellingMachine Translation

Proceedings of the Celtic Language Technology Workshop

2019-08-01 · WS 2019 8 ·