Common Ground, Diverse Roots: The Difficulty of Classifying Common Examples in Spanish Varieties
Variations in languages across geographic regions or cultures are crucial to address to avoid biases in NLP systems designed for culturally sensitive tasks, such as hate speech detection or dialog with conversational agents. In languages such as Spanish, where varieties can significantly overlap, many examples can be valid across them, which we refer to as common examples. Ignoring these examples may cause misclassifications, reducing model accuracy and fairness. Therefore, accounting for these common examples is essential to improve the robustness and representativeness of NLP systems trained on such data. In this work, we address this problem in the context of Spanish varieties. We use training dynamics to automatically detect common examples or errors in existing Spanish datasets. We demonstrate the efficacy of using predicted label confidence for our Datamaps \cite{swayamdipta-etal-2020-dataset} implementation for the identification of hard-to-classify examples, especially common examples, enhancing model performance in variety identification tasks. Additionally, we introduce a Cuban Spanish Variety Identification dataset with common examples annotations developed to facilitate more accurate detection of Cuban and Caribbean Spanish varieties. To our knowledge, this is the first dataset focused on identifying the Cuban, or any other Caribbean, Spanish variety.
Code (0)
등록된 구현이 없습니다.
Tasks
FairnessHate Speech DetectionvalidSimilar Papers 제목 키워드 기반
Accurate Tree Roots Positioning and Sizing over Undulated Ground Surfaces by Common Offset GPR Measurements
Tree roots detection is a popular application of the Ground-penetrating radar (GPR). Normally, the ground surface above the tree roots is assumed to be flat, and standard processing methods based on hyperbolic fitting ar…
GPRAl Qamus al Muhit, a Medieval Arabic Lexicon in LMF
This paper describes the conversion into LMF, a standard lexicographic digital format of {`}al-q{\=a}m{\=u}s al-muḥ{\=\i}ṭ, a Medieval Arabic lexicon. The lexicon is first described, then all the steps required for the c…
Slice-Connection Clustering Algorithm for Tree Roots Recognition in Noisy 3D GPR Data
3D mapping of tree roots is a popular ground-penetrating radar (GPR) application. In real field tests, the recognition of tree roots suffers due to noisey reflection patterns from subsurface targets that are not of inter…
ClusteringGPRThe Roots of Performance Disparity in Multilingual Language Models: Intrinsic Modeling Difficulty or Design Choices?
Multilingual language models (LMs) promise broader NLP access, yet current systems deliver uneven performance across the world's languages. This survey examines why these gaps persist and whether they reflect intrinsic l…
3D Plant Root Skeleton Detection and Extraction
Plant roots typically exhibit a highly complex and dense architecture, incorporating numerous slender lateral roots and branches, which significantly hinders the precise capture and modeling of the entire root system. Ad…