On the Limits of Minimal Pairs in Contrastive Evaluation
Minimal sentence pairs are frequently used to analyze the behavior of language models. It is often assumed that model behavior on contrastive pairs is predictive of model behavior at large. We argue that two conditions are necessary for this assumption to hold: First, a tested hypothesis should be well-motivated, since experiments show that contrastive evaluation can lead to false positives. Secondly, test data should be chosen such as to minimize distributional discrepancy between evaluation time and deployment time. For a good approximation of deployment-time decoding, we recommend that minimal pairs are created based on machine-generated text, as opposed to human-written references. We present a contrastive evaluation suite for English-German MT that implements this recommendation.
Code (1)
Tasks
SentenceSimilar Papers 제목 키워드 기반
DMMG: Dual Min-Max Games for Self-Supervised Skeleton-Based Action Recognition
In this work, we propose a new Dual Min-Max Games (DMMG) based self-supervised skeleton action recognition method by augmenting unlabeled data in a contrastive learning framework. Our DMMG consists of a viewpoint variati…
Action RecognitionContrastive LearningData AugmentationSelf-supervised Skeleton-based Action Recognition+1Cross-Stream Contrastive Learning for Self-Supervised Skeleton-Based Action Recognition
Self-supervised skeleton-based action recognition enjoys a rapid growth along with the development of contrastive learning. The existing methods rely on imposing invariance to augmentations of 3D skeleton within a single…
Action RecognitionContrastive LearningRepresentation LearningSelf-supervised Skeleton-based Action Recognition+1Progressive Compositionality In Text-to-Image Generative Models
Despite the impressive text-to-image (T2I) synthesis capabilities of diffusion models, they often struggle to understand compositional relationships between objects and attributes, especially in complex settings. Existin…
AttributeContrastive LearningQuestion AnsweringVisual Question Answering+1Rel3D: A Minimally Contrastive Benchmark for Grounding Spatial Relations in 3D
Understanding spatial relations (e.g., "laptop on table") in visual input is important for both humans and robots. Existing datasets are insufficient as they lack large-scale, high-quality 3D ground truth information, wh…
RelationSpatial Relation RecognitionHybrid-Collaborative Augmentation and Contrastive Sample Adaptive-Differential Awareness for Robust Attributed Graph Clustering
Due to its powerful capability of self-supervised representation learning and clustering, contrastive attributed graph clustering (CAGC) has achieved great success, which mainly depends on effective data augmentation and…
Representation LearningContrastive LearningData AugmentationGraph Clustering