Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers
As large-scale, pre-trained language models achieve human-level and superhuman accuracy on existing language understanding tasks, statistical bias in benchmark data and probing studies have recently called into question their true capabilities. For a more informative evaluation than accuracy on text classification tasks can offer, we propose evaluating systems through a novel measure of prediction coherence. We apply our framework to two existing language understanding benchmarks with different properties to demonstrate its versatility. Our experimental results show that this evaluation framework, although simple in ideas and implementation, is a quick, effective, and versatile measure to provide insight into the coherence of machines' predictions.
Code (1)
Tasks
text-classificationText ClassificationSimilar Papers 제목 키워드 기반
Iceberg: Enhancing HLS Modeling with Synthetic Data
Deep learning-based prediction models for High-Level Synthesis (HLS) of hardware designs often struggle to generalize. In this paper, we study how to close the generalizability gap of these models through pretraining on …
Data AugmentationHigh-Level SynthesisLanguage ModelingLanguage Modelling+2CME Iceberg Order Detection and Prediction
We propose a method for detection and prediction of native and synthetic iceberg orders on Chicago Mercantile Exchange. Native (managed by the exchange) icebergs are detected using discrepancies between the resting volum…
PredictionHybrid NARX-LLM for Greenland Iceberg Discharge: Prompt-Driven Residual Correction
Greenland iceberg discharge exhibits complex nonlinear dynamics with limited observability, challenging traditional predictive models. We present a Hybrid NARX-LLM framework that combines a nonlinear autoregressive model…
SARA: Stress Test Reasoning in Audio Deepfake Detection
Audio Language Models (ALMs) offer a promising shift towards explainable audio deepfake detections (ADD), moving beyond \textit{black-box} classifiers by providing transparency to their predictions via reasoning traces. …
Audio Deepfake DetectionDiscourse Coherence in the Wild: A Dataset, Evaluation and Methods
To date there has been very little work on assessing discourse coherence methods on real-world data. To address this, we present a new corpus of real-world texts (GCDC) as well as the first large-scale evaluation of lead…
Coherence Evaluation