paper-with-me

홈 › Papers

Annotation Efficient Language Identification from Weak Labels

2020-11-01 · EMNLP (WNUT) 2020 11 · Shriphani Palakodety, Ashiqur KhudaBukhsh

India is home to several languages with more than 30m speakers. These languages exhibit significant presence on social media platforms. However, several of these widely-used languages are under-addressed by current Natural Language Processing (NLP) models and resources. User generated social media content in these languages is also typically authored in the Roman script as opposed to the traditional native script further contributing to resource scarcity. In this paper, we leverage a minimally supervised NLP technique to obtain weak language labels from a large-scale Indian social media corpus leading to a robust and annotation-efficient language-identification technique spanning nine Romanized Indian languages. In fast-spreading pandemic situations such as the current COVID-19 situation, information processing objectives might be heavily tilted towards under-served languages in densely populated regions. We release our models to facilitate downstream analyses in these low-resource languages. Experiments across multiple social media corpora demonstrate the model’s robustness and provide several interesting insights on Indian language usage patterns on social media. We release an annotated data set of 1,000 comments in ten Romanized languages as a social media evaluation benchmark.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language Identification

Similar Papers 제목 키워드 기반

Learning Person Re-identification Models from Videos with Weak Supervision

2020-07-21 · Xueping Wang, Sujoy Paul, Dripta S. Raychaudhuri, Min Liu 외

Most person re-identification methods, being supervised techniques, suffer from the burden of massive annotation requirement. Unsupervised methods overcome this need for labeled data, but perform poorly compared to the s…

Multiple Instance LearningPerson Re-IdentificationVideo-Based Person Re-Identification

Weakly-supervised diagnosis identification from Italian discharge letters

2024-10-19 · Vittorio Torri, Elisa Barbieri, Anna Cantarutti, Carlo Giaquinto 외

Objective: Recognizing diseases from discharge letters is crucial for cohort selection and epidemiological analyses, as this is the only type of data consistently produced across hospitals. This is a classic document cla…

Document Classificationtext-classificationText Classification

Weakly supervised discriminative feature learning with state information for person identification

2020-02-27 · CVPR 2020 6 · Hong-Xing Yu, Wei-Shi Zheng

Unsupervised learning of identity-discriminative visual feature is appealing in real-world tasks where manual labelling is costly. However, the images of an identity can be visually discrepant when images are taken under…

Face RecognitionPerson IdentificationPerson Re-IdentificationPseudo Label+2

Leveraging Vision-Language Models as Weak Annotators in Active Learning

2026-05-01 · Phuong Ngoc Nguyen, Kaito Shiku, Ryoma Bise, Seiichi Uchida 외 arxiv

Active learning aims to reduce annotation cost by selectively querying informative samples for supervision under a limited labeling budget. In this work, we investigate how vision-language models (VLMs) can be leveraged …

Active Learning

Weakly Supervised Training of Speaker Identification Models

2018-06-22 · Martin Karu, Tanel Alumäe

We propose an approach for training speaker identification models in a weakly supervised manner. We concentrate on the setting where the training data consists of a set of audio recordings and the speaker annotation is p…

speaker-diarizationSpeaker DiarizationSpeaker Identification