High-Resource Methodological Bias in Low-Resource Investigations
The central bottleneck for low-resource NLP is typically regarded to be the quantity of accessible data, overlooking the contribution of data quality. This is particularly seen in the development and evaluation of low-resource systems via down sampling of high-resource language data. In this work we investigate the validity of this approach, and we specifically focus on two well-known NLP tasks for our empirical investigations: POS-tagging and machine translation. We show that down sampling from a high-resource language results in datasets with different properties than the low-resource datasets, impacting the model performance for both POS-tagging and machine translation. Based on these results we conclude that naive down sampling of datasets results in a biased view of how well these systems work in a low-resource scenario.
Code (0)
등록된 구현이 없습니다.
Tasks
Machine TranslationPOSPOS TaggingTranslationVocal Bursts Intensity PredictionSimilar Papers 제목 키워드 기반
An Empirical Study on the Characteristics of Bias upon Context Length Variation for Bangla
Pretrained language models inherently exhibit various social biases, prompting a crucial examination of their social impact across various linguistic contexts due to their widespread usage. Previous studies have provided…
How do datasets, developers, and models affect biases in a low-resourced language?
Sociotechnical systems, such as language technologies, frequently exhibit identity-based biases. These biases exacerbate the experiences of historically marginalized communities and remain understudied in low-resource co…
Sentiment AnalysisExploring the Potential of Large Language Models for Improving Digital Forensic Investigation Efficiency
The ever-increasing workload of digital forensic labs raises concerns about law enforcement's ability to conduct both cyber-related and non-cyber-related investigations promptly. Consequently, this article explores the p…
Cross-Lingual Probing and Community-Grounded Analysis of Gender Bias in Low-Resource Bengali
Large Language Models (LLMs) have achieved significant success in recent years; yet, issues of intrinsic gender bias persist, especially in non-English languages. Although current research mostly emphasizes English, the …
Bias DetectionLLM as Prompter: Low-resource Inductive Reasoning on Arbitrary Knowledge Graphs
Knowledge Graph (KG) inductive reasoning, which aims to infer missing facts from new KGs that are not seen during training, has been widely adopted in various applications. One critical challenge of KG inductive reasonin…
Knowledge Graphs