Small Languages, Big Data: Multilingual Computational Tools and Techniques for the Lexicography of Endangered Languages
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
A survey of methods to ease the development of highly multilingual text mining applications
Multilingual text processing is useful because the information content found in different languages is complementary, both regarding facts and opinions. While Information Extraction and other text mining software can, in…
ArticlesBetter as Generators Than Classifiers: Leveraging LLMs and Synthetic Data for Low-Resource Multilingual Classification
Large Language Models (LLMs) have demonstrated remarkable multilingual capabilities, making them promising tools in both high- and low-resource languages. One particularly valuable use case is generating synthetic sample…
Synthetic Data GenerationEmbedding structure matters: Comparing methods to adapt multilingual vocabularies to new languages
Pre-trained multilingual language models underpin a large portion of modern NLP tools outside of English. A strong baseline for specializing these models for specific languages is Language-Adaptive Pre-Training (LAPT). H…
XLM-EMO: Multilingual Emotion Prediction in Social Media Text
Detecting emotion in text allows social and computational scientists to study how people behave and react to online events. However, developing these tools for different languages requires data that is not always availab…
PredictionThe DReaM Corpus: A Multilingual Annotated Corpus of Grammars for the World's Languages
There exist as many as 7000 natural languages in the world, and a huge number of documents describing those languages have been produced over the years. Most of those documents are in paper format. Any attempts to use mo…