From Robustness to Improved Generalization and Calibration in Pre-trained Language Models
Enhancing generalization and uncertainty quantification in pre-trained language models (PLMs) is crucial for their effectiveness and reliability. Building on machine learning research that established the importance of robustness for improving generalization, we investigate the role of representation smoothness, achieved via Jacobian and Hessian regularization, in enhancing PLM performance. Although such regularization methods have proven effective in computer vision, their application in natural language processing (NLP), where PLM inputs are derived from a discrete domain, poses unique challenges. We introduce a novel two-phase regularization approach, JacHess, which minimizes the norms of the Jacobian and Hessian matrices within PLM intermediate representations relative to their inputs. Our evaluation using the GLUE benchmark demonstrates that JacHess significantly improves in-domain generalization and calibration in PLMs, outperforming unregularized fine-tuning and other similar regularization methods.
Code (0)
등록된 구현이 없습니다.
Tasks
Domain GeneralizationUncertainty QuantificationSimilar Papers 제목 키워드 기반
Gaussian Stochastic Weight Averaging for Bayesian Low-Rank Adaptation of Large Language Models
Fine-tuned Large Language Models (LLMs) often suffer from overconfidence and poor calibration, particularly when fine-tuned on small datasets. To address these challenges, we propose a simple combination of Low-Rank Adap…
Bayesian InferenceLanguage as Prior, Vision as Calibration: Metric Scale Recovery for Monocular Depth Estimation
Relative-depth foundation models transfer well, yet monocular metric depth remains ill-posed due to unidentifiable global scale and heightened domain-shift sensitivity. Under a frozen-backbone calibration setting, we rec…
Monocular Depth EstimationUncertaintyRAG: Span-Level Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation
We present UncertaintyRAG, a novel approach for long-context Retrieval-Augmented Generation (RAG) that utilizes Signal-to-Noise Ratio (SNR)-based span uncertainty to estimate similarity between text chunks. This span unc…
ChunkingLanguage ModelingLanguage ModellingLarge Language Model+3An Empirical Study on Robustness to Spurious Correlations using Pre-trained Language Models
Recent work has shown that pre-trained language models such as BERT improve robustness to spurious correlations in the dataset. Intrigued by these results, we find that the key to their success is generalization from a s…
DiversityMulti-Task LearningNatural Language InferenceParaphrase IdentificationFoundation Feature-Driven Online End-Effector Pose Estimation: A Marker-Free and Learning-Free Approach
Accurate transformation estimation between camera space and robot space is essential. Traditional methods using markers for hand-eye calibration require offline image collection, limiting their suitability for online sel…
6D Pose EstimationPose EstimationRobot Pose EstimationZero-shot Generalization