paper-with-me

Papers

Spotlights and Blindspots: Evaluating Machine-Generated Text Detection

2026-04-17 · Kevin Stowe, Kailash Patil arxiv

With the rise of generative language models, machine-generated text detection has become a critical challenge. A wide variety of models is available, but inconsistent datasets, evaluation metrics, and assessment strategies obscure comparisons of model effectiveness. To address this, we evaluate 15 different detection models from six distinct systems, as well as seven trained models, across seven English-language textual test sets and three creative human-written datasets. We provide an empirical analysis of model performance, the influence of training and evaluation data, and the impact of key metrics. We find that no single system excels in all areas and nearly all are effective for certain tasks, and the representation of model performance is critically linked to dataset and metric choices. We find high variance in model ranks based on datasets and metrics, and overall poor performance on novel human-written texts in high-risk domains. Across datasets and metrics, we find that methodological choices that are often assumed or overlooked are essential for clearly and accurately reflecting model performance.

📄 PDF Abstract BibTeX arXiv:2604.16607

Code (0)

등록된 구현이 없습니다.

Tasks

Text Detection

Similar Papers 제목 키워드 기반

Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders

2025-06-24 · Matyas Bohacek, Thomas Fel, Maneesh Agrawala, Ekdeep Singh Lubana

Despite their impressive performance, generative image models trained on large-scale datasets frequently fail to produce images with seemingly simple concepts -- e.g., human hands or objects appearing in groups of four -…

Memorization

Towards a More Rigorous Science of Blindspot Discovery in Image Classification Models

2022-07-08 · Gregory Plumb, Nari Johnson, Ángel Alexander Cabrera, Ameet Talwalkar

A growing body of work studies Blindspot Discovery Methods ("BDM"s): methods that use an image embedding to find semantically meaningful (i.e., united by a human-understandable concept) subsets of the data where an image…

Dimensionality Reductionimage-classificationImage Classification

SlideAVSR: A Dataset of Paper Explanation Videos for Audio-Visual Speech Recognition

2024-01-18 · Hao Wang, Shuhei Kurita, Shuichiro Shimizu, Daisuke Kawahara

Audio-visual speech recognition (AVSR) is a multimodal extension of automatic speech recognition (ASR), using video as a complement to audio. In AVSR, considerable efforts have been directed at datasets for facial featur…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Image Comprehension+3

Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents

2025-09-13 · Ankan Mullick, Sombit Bose, Rounak Saha, Ayan Kumar Bhowmick 외 arxiv

In this paper, we introduce Spotlight, a novel paradigm for information extraction that produces concise, engaging narratives by highlighting the most compelling aspects of a document. Unlike traditional summaries, which…

Information Extraction

ESCAPE: Countering Systematic Errors from Machine's Blind Spots via Interactive Visual Analysis

2023-03-16 · Yongsu Ahn, Yu-Ru Lin, Panpan Xu, Zeng Dai

Classification models learn to generalize the associations between data samples and their target classes. However, researchers have increasingly observed that machine learning practice easily leads to systematic errors i…