paper-with-me

홈 › Papers

Analysis and Mitigation of Dataset Artifacts in OpenAI GPT-3

2021-12-19 · Mingi Ryu

With the recent release of public beta, we took full advantage of OpenAI's Models-as-a-Service (MaaS) offering of GPT-3 to analyze and mitigate dataset artifacts in a model that has one of the highest number of parameters. Recent studies in dataset artifacts and adversarial attacks suggest that state of the art (SoTA) NLP models are susceptible to spurious correlations in training datasets. We decided to investigate GPT-3 on dataset artifacts taking advantage of its large scale and task-agnostic pre-training. We began by verifying few-shot capabilities of GPT-3 in order to lay the groundwork for analysis. Furthermore, we employed our approach to fine-tuning for the natural language inference (NLI) task. Using SNLI as a baseline, we carried out several experiments with Adversarial NLI (ANLI) to evaluate the performance and robustness of GPT-3. Our findings suggest that using adversarial datasets could mitigate dataset artifacts in GPT-3 at a negligible overall performance cost.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language Inference

Methods 이 논문이 사용한 방법론

{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…

Similar Papers 제목 키워드 기반

The 2025 OpenAI Preparedness Framework does not guarantee any AI risk mitigation practices: a proof-of-concept for affordance analyses of AI safety policies

2025-09-29 · Sam Coggins, Alexander K. Saeri, Katherine A. Daniell, Lorenn P. Ruster 외 arxiv

Prominent AI companies are producing 'safety frameworks' as a type of voluntary self-governance. These statements purport to establish risk thresholds and safety procedures for the development and deployment of highly ca…

Temporal Score Analysis for Understanding and Correcting Diffusion Artifacts

2025-03-20 · CVPR 2025 1 · Yu Cao, Zengqun Zhao, Ioannis Patras, Shaogang Gong

Visual artifacts remain a persistent challenge in diffusion models, even with training on massive datasets. Current solutions primarily rely on supervised detectors, yet lack understanding of why these artifacts occur in…

Denoising

Multi-Scales Data Augmentation Approach In Natural Language Inference For Artifacts Mitigation And Pre-Trained Model Optimization

2022-12-16 · Zhenyuan Lu

Machine learning models can reach high performance on benchmark natural language processing (NLP) datasets but fail in more challenging settings. We study this issue when a pre-trained model learns dataset artifacts in n…

Data AugmentationModel OptimizationNatural Language InferenceSentence

Can Multi-modal (reasoning) LLMs detect document manipulation?

2025-08-14 · Zisheng Liang, Kidus Zewde, Rudra Pratap Singh, Disha Patil 외 arxiv

Document fraud poses a significant threat to industries reliant on secure and verifiable documentation, necessitating robust detection mechanisms. This study investigates the efficacy of state-of-the-art multi-modal larg…

Zero-shot GeneralizationFraud Detection

Defending Large Language Models Against Attacks With Residual Stream Activation Analysis

2024-06-05 · Amelia Kawasaki, Andrew Davis, Houssam Abbas

The widespread adoption of Large Language Models (LLMs), exemplified by OpenAI's ChatGPT, brings to the forefront the imperative to defend against adversarial threats on these models. These attacks, which manipulate an L…