Analysis and Mitigation of Dataset Artifacts in OpenAI GPT-3
With the recent release of public beta, we took full advantage of OpenAI's Models-as-a-Service (MaaS) offering of GPT-3 to analyze and mitigate dataset artifacts in a model that has one of the highest number of parameters. Recent studies in dataset artifacts and adversarial attacks suggest that state of the art (SoTA) NLP models are susceptible to spurious correlations in training datasets. We decided to investigate GPT-3 on dataset artifacts taking advantage of its large scale and task-agnostic pre-training. We began by verifying few-shot capabilities of GPT-3 in order to lay the groundwork for analysis. Furthermore, we employed our approach to fine-tuning for the natural language inference (NLI) task. Using SNLI as a baseline, we carried out several experiments with Adversarial NLI (ANLI) to evaluate the performance and robustness of GPT-3. Our findings suggest that using adversarial datasets could mitigate dataset artifacts in GPT-3 at a negligible overall performance cost.
Code (0)
등록된 구현이 없습니다.
Tasks
Natural Language InferenceMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies
Prominent AI companies are producing 'safety frameworks' as a type of voluntary self-governance. These statements purport to establish risk thresholds and safety procedures for the development and deployment of highly ca…
Temporal Score Analysis for Understanding and Correcting Diffusion Artifacts
Visual artifacts remain a persistent challenge in diffusion models, even with training on massive datasets. Current solutions primarily rely on supervised detectors, yet lack understanding of why these artifacts occur in…
DenoisingMulti-Scales Data Augmentation Approach In Natural Language Inference For Artifacts Mitigation And Pre-Trained Model Optimization
Machine learning models can reach high performance on benchmark natural language processing (NLP) datasets but fail in more challenging settings. We study this issue when a pre-trained model learns dataset artifacts in n…
Data AugmentationModel OptimizationNatural Language InferenceSentenceCan Multi-modal (reasoning) LLMs detect document manipulation?
Document fraud poses a significant threat to industries reliant on secure and verifiable documentation, necessitating robust detection mechanisms. This study investigates the efficacy of state-of-the-art multi-modal larg…
Zero-shot GeneralizationFraud DetectionDefending Large Language Models Against Attacks With Residual Stream Activation Analysis
The widespread adoption of Large Language Models (LLMs), exemplified by OpenAI's ChatGPT, brings to the forefront the imperative to defend against adversarial threats on these models. These attacks, which manipulate an L…