Different Time, Different Language: Revisiting the Bias Against Non-Native Speakers in GPT Detectors
LLM-based assistants have been widely popularised after the release of ChatGPT. Concerns have been raised about their misuse in academia, given the difficulty of distinguishing between human-written and generated text. To combat this, automated techniques have been developed and shown to be effective, to some extent. However, prior work suggests that these methods often falsely flag essays from non-native speakers as generated, due to their low perplexity extracted from an LLM, which is supposedly a key feature of the detectors. We revisit these statements two years later, specifically in the Czech language setting. We show that the perplexity of texts from non-native speakers of Czech is not lower than that of native speakers. We further examine detectors from three separate families and find no systematic bias against non-native speakers. Finally, we demonstrate that contemporary detectors operate effectively without relying on perplexity.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Is There a One-Model-Fits-All Approach to Information Extraction? Revisiting Task Definition Biases
Definition bias is a negative phenomenon that can mislead models. Definition bias in information extraction appears not only across datasets from different domains but also within datasets sharing the same domain. We ide…
AllFrom If-Statements to ML Pipelines: Revisiting Bias in Code-Generation
Prior work evaluates code generation bias primarily through simple conditional statements, which represent only a narrow slice of real-world programming and reveal solely overt, explicitly encoded bias. We demonstrate th…
Code GenerationKnowledgeable or Educated Guess? Revisiting Language Models as Knowledge Bases
Previous literatures show that pre-trained masked language models (MLMs) such as BERT can achieve competitive factual knowledge extraction performance on some datasets, indicating that MLMs can potentially be a reliable …
Revisiting the Shape-Bias of Deep Learning for Dermoscopic Skin Lesion Classification
It is generally believed that the human visual system is biased towards the recognition of shapes rather than textures. This assumption has led to a growing body of work aiming to align deep models' decision-making proce…
ClassificationDecision MakingLesion ClassificationSkin Lesion ClassificationRevisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias
We characterize and study zero-shot abstractive summarization in Large Language Models (LLMs) by measuring position bias, which we propose as a general formulation of the more restrictive lead bias phenomenon studied pre…
Abstractive Text SummarizationDecoderPosition