paper-with-me

홈 › Papers

Towards Augmenting Lexical Resources for Slang and African American English

2020-12-01 · VarDial (COLING) 2020 12 · Alyssa Hwang, William R. Frey, Kathleen McKeown

Researchers in natural language processing have developed large, robust resources for understanding formal Standard American English (SAE), but we lack similar resources for variations of English, such as slang and African American English (AAE). In this work, we use word embeddings and clustering algorithms to group semantically similar words in three datasets, two of which contain high incidence of slang and AAE. Since high-quality clusters would contain related words, we could also infer the meaning of an unfamiliar word based on the meanings of words clustered with it. After clustering, we compute precision and recall scores using WordNet and ConceptNet as gold standards and show that these scores are unimportant when the given resources do not fully represent slang and AAE. Amazon Mechanical Turk and expert evaluations show that clusters with low precision can still be considered high quality, and we propose the new Cluster Split Score as a metric for machine-generated clusters. These contributions emphasize the gap in natural language processing research for variations of English and motivate further work to close it.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringWord Embeddings

Methods 이 논문이 사용한 방법론

American 설명 없음

Similar Papers 제목 키워드 기반

Twitter Universal Dependency Parsing for African-American and Mainstream American English

2018-07-01 · ACL 2018 7 · Su Lin Blodgett, Johnny Wei, Brendan O{'}Connor

Due to the presence of both Twitter-specific conventions and non-standard and dialectal language, Twitter presents a significant parsing challenge to current dependency parsing tools. We broaden English dependency parsin…

Dependency ParsingInformation RetrievalLanguage IdentificationPart-Of-Speech Tagging+1

SLAyiNG: A Diverse and Community-validated Dataset of Queer Slang

2025-09-22 · Leonor Veloso, Lea Hirlimann, Philipp Wicke, Hinrich Schütze arxiv

Queer vernacular is rarely studied in NLP, despite advancements in resources and evaluation for other sociolects and informal language. Because of this, NLP systems often process queer language incorrectly, e.g., they mi…

Automatic Speech Recognition of African American English: Lexical and Contextual Effects

2025-06-07 · Hamid Mojarad, Kevin Tang

Automatic Speech Recognition (ASR) models often struggle with the phonetic, phonological, and morphosyntactic features found in African American English (AAE). This study focuses on two key AAE variables: Consonant Clust…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

Advancing Conversational AI with Shona Slang: A Dataset and Hybrid Model for Digital Inclusion

2025-09-10 · Happymore Masoka arxiv

African languages remain underrepresented in natural language processing (NLP), with most corpora limited to formal registers that fail to capture the vibrancy of everyday communication. This work addresses this gap for …

Intent Recognition

Multilingual offensive lexicon annotated with contextual information

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Online hate speech and offensive comments detection is not a trivial research problem since pragmatic (contextual) factors influence what is considered offensive. Moreover, offensive terms are hardly found in classical l…

Abusive Language