paper-with-me

홈 › Papers

Dropout Robustness and Cognitive Profiling of Transformer Models via Stochastic Inference

2026-03-18 · Antônio Junior Alves Caiado, Michael Hahsler arxiv

Transformer-based language models are widely deployed for reasoning, yet their behavior under inference-time stochasticity remains underexplored. While dropout is common during training, its inference-time effects via Monte Carlo sampling lack systematic evaluation across architectures, limiting understanding of model reliability in uncertainty-aware applications. This work analyzes dropout-induced variability across 19 transformer models using MC Dropout with 100 stochastic forward passes per sample. Dropout robustness is defined as maintaining high accuracy and stable predictions under stochastic inference, measured by standard deviation of per-run accuracies. A cognitive decomposition framework disentangles performance into memory and reasoning components. Experiments span five dropout configurations yielding 95 unique evaluations on 1,000 samples. Results reveal substantial architectural variation. Smaller models demonstrate perfect prediction stability while medium-sized models exhibit notable volatility. Mid-sized models achieve the best overall performance; larger models excel at memory tasks. Critically, 53% of models suffer severe accuracy degradation under baseline MC Dropout, with task-specialized models losing up to 24 percentage points, indicating unsuitability for uncertainty quantification in these architectures. Asymmetric effects emerge: high dropout reduces memory accuracy by 27 percentage points while reasoning degrades only 1 point, suggesting memory tasks rely on stable representations that dropout disrupts. 84% of models demonstrate memory-biased performance. This provides the first comprehensive MC Dropout benchmark for transformers, revealing dropout robustness is architecture-dependent and uncorrelated with scale. The cognitive profiling framework offers actionable guidance for model selection in uncertainty-aware applications.

📄 PDF Abstract BibTeX arXiv:2603.17811

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CognitiveTwin: Robust Multi-Modal Digital Twins for Predicting Cognitive Decline in Alzheimer's Disease

2026-04-24 · Bulent Soykan, Gulsah Hancerliogullari Koksalmis, Hsin-Hsiung Huang, Laura J. Brattain arxiv

Predicting individual cognitive decline in Alzheimer's disease (AD) is difficult due to the heterogeneity of disease progression. Reliable clinical tools require not only high accuracy but also fairness across demographi…

Explicit Dropout: Deterministic Regularization for Transformer Architectures

2026-04-22 · Vidhi Agrawal, Illia Oleksiienko, Alexandros Iosifidis arxiv

Dropout is a widely used regularization technique in deep learning, but its effects are typically realized through stochastic masking rather than explicit optimization objectives. We propose a deterministic formulation t…

Audio ClassificationImage ClassificationAction Detection

Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference

2026-09-04 · Mostafa Elhoushi, Alex Pretko, Nolan Dey, Bin Claire Zhang 외 arxiv

Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have …

Pyramid Adversarial Training Improves ViT Performance

2021-11-30 · CVPR 2022 1 · Charles Herrmann, Kyle Sargent, Lu Jiang, Ramin Zabih 외

Aggressive data augmentation is a key component of the strong generalization capabilities of Vision Transformer (ViT). One such data augmentation technique is adversarial training (AT); however, many prior works have sho…

Adversarial AttackData AugmentationDomain GeneralizationImage Classification

Deepfake Video Detection with Spatiotemporal Dropout Transformer

2022-07-14 · Daichi Zhang, Fanzhao Lin, Yingying Hua, Pengju Wang 외

While the abuse of deepfake technology has caused serious concerns recently, how to detect deepfake videos is still a challenge due to the high photo-realistic synthesis of each frame. Existing image-level approaches oft…

Data AugmentationFace Swapping