paper-with-me

Papers

PECO: Examining Single Sentence Label Leakage in Natural Language Inference Datasets through Progressive Evaluation of Cluster Outliers

2021-12-16 · Michael Saxon, Xinyi Wang, Wenda Xu, William Yang Wang

Building natural language inference (NLI) benchmarks that are both challenging for modern techniques, and free from shortcut biases is difficult. Chief among these biases is "single sentence label leakage," where annotator-introduced spurious correlations yield datasets where the logical relation between (premise, hypothesis) pairs can be accurately predicted from only a single sentence, something that should in principle be impossible. We demonstrate that despite efforts to reduce this leakage, it persists in modern datasets that have been introduced since its 2018 discovery. To enable future amelioration efforts, introduce a novel model-driven technique, the progressive evaluation of cluster outliers (PECO) which enables both the objective measurement of leakage, and the automated detection of subpopulations in the data which maximally exhibit it.

📄 PDF Abstract BibTeX arXiv:2112.09237

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language InferenceSentence

Similar Papers 제목 키워드 기반

PECOS: Prediction for Enormous and Correlated Output Spaces

2020-10-12 · Hsiang-Fu Yu, Kai Zhong, Jiong Zhang, Wei-Cheng Chang 외

Many large-scale applications amount to finding relevant results from an enormous output space of potential candidates. For example, finding the best matching product from a large catalog or suggesting related search phr…

Prediction

PECoP: Parameter Efficient Continual Pretraining for Action Quality Assessment

2023-11-11 · Amirhossein Dadashzadeh, Shuchao Duan, Alan Whone, Majid Mirmehdi

The limited availability of labelled data in Action Quality Assessment (AQA), has forced previous works to fine-tune their models pretrained on large-scale domain-general datasets. This common approach results in weak ge…

Action Quality AssessmentContinual PretrainingSelf-Supervised Learning

SpeContext: Enabling Efficient Long-context Reasoning with Speculative Context Sparsity in LLMs

2025-11-30 · Jiaming Xu, Jiayi Pan, Hanzhen Wang, Yongkang Zhou 외 arxiv

In this paper, we point out that the objective of the retrieval algorithms is to align with the LLM, which is similar to the objective of knowledge distillation in LLMs. We analyze the similarity in information focus bet…

Knowledge Distillation

Mol-PECO: a deep learning model to predict human olfactory perception from molecular structures

2023-05-21 · Mengji Zhang, Yusuke Hiki, Akira Funahashi, Tetsuya J. Kobayashi

While visual and auditory information conveyed by wavelength of light and frequency of sound have been decoded, predicting olfactory information encoded by the combination of odorants remains challenging due to the unkno…

molecular representationRetrieval

PECon: Contrastive Pretraining to Enhance Feature Alignment between CT and EHR Data for Improved Pulmonary Embolism Diagnosis

2023-08-27 · Santosh Sanjeev, Salwa K. Al Khatib, Mai A. Shaaban, Ibrahim Almakky 외

Previous deep learning efforts have focused on improving the performance of Pulmonary Embolism(PE) diagnosis from Computed Tomography (CT) scans using Convolutional Neural Networks (CNN). However, the features from CT sc…

Computed Tomography (CT)Contrastive LearningPulmonary Embolism Detection