paper-with-me

Papers

Platforms for Non-speakers Annotating Names in Any Language

2018-07-01 · ACL 2018 7 · Ying Lin, Cash Costello, Boliang Zhang, Di Lu, Heng Ji, James Mayfield, Paul McNamee

We demonstrate two annotation platforms that allow an English speaker to annotate names for any language without knowing the language. These platforms provided high-quality {'}{`}silver standard{''} annotations for low-resource language name taggers (Zhang et al., 2017) that achieved state-of-the-art performance on two surprise languages (Oromo and Tigrinya) at LoreHLT20171 and ten languages at TAC-KBP EDL2017 (Ji et al., 2017). We discuss strengths and limitations and compare other methods of creating silver- and gold-standard annotations using native speakers. We will make our tools publicly available for research use.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SCALAR: A Part-of-speech Tagger for Identifiers

2025-04-23 · Christian D. Newman, Brandon Scholten, Sophia Testa, Joshua A. C. Behler 외

The paper presents the Source Code Analysis and Lexical Annotation Runtime (SCALAR), a tool specialized for mapping (annotating) source code identifier names to their corresponding part-of-speech tag sequence (grammar pa…

TAG

Dragonfly: Advances in Non-Speaker Annotation for Low Resource Languages

2020-05-01 · LREC 2020 5 · Cash Costello, Shelby Anderson, Caitlyn Bishop, James Mayfield 외

Dragonfly is an open source software tool that supports annotation of text in a low resource language by non-speakers of the language. Using semantic and contextual information, non-speakers of a language familiar with t…

VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency

2026-06-23 · Viet Hoang Pham, Tran Trung Nguyen, Bao Thu Ho, Phuong Tuan Dat 외 arxiv

Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity. Most large-scale datasets rely on facial…

Speaker Recognition

A Large-scale Dataset for Hate Speech Detection on Vietnamese Social Media Texts

2021-03-22 · Son T. Luu, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen

In recent years, Vietnam witnesses the mass development of social network users on different social platforms such as Facebook, Youtube, Instagram, and Tiktok. On social medias, hate speech has become a critical problem …

Hate Speech DetectionVietnamese Hate Speech DetectionVietnamese Social Media Text Processing

A high quality and phonetic balanced speech corpus for Vietnamese

2019-04-11 · Pham Ngoc Phuong, Quoc Truong Do, Luong Chi Mai

This paper presents a high quality Vietnamese speech corpus that can be used for analyzing Vietnamese speech characteristic as well as building speech synthesis models. The corpus consists of 5400 clean-speech utterances…

Speech SynthesisVocal Bursts Intensity Prediction