paper-with-me

홈 › Papers

From Robustness to Improved Generalization and Calibration in Pre-trained Language Models

2024-03-31 · Josip Jukić, Jan Šnajder

Enhancing generalization and uncertainty quantification in pre-trained language models (PLMs) is crucial for their effectiveness and reliability. Building on machine learning research that established the importance of robustness for improving generalization, we investigate the role of representation smoothness, achieved via Jacobian and Hessian regularization, in enhancing PLM performance. Although such regularization methods have proven effective in computer vision, their application in natural language processing (NLP), where PLM inputs are derived from a discrete domain, poses unique challenges. We introduce a novel two-phase regularization approach, JacHess, which minimizes the norms of the Jacobian and Hessian matrices within PLM intermediate representations relative to their inputs. Our evaluation using the GLUE benchmark demonstrates that JacHess significantly improves in-domain generalization and calibration in PLMs, outperforming unregularized fine-tuning and other similar regularization methods.

📄 PDF Abstract BibTeX arXiv:2404.00758

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationUncertainty Quantification

Similar Papers 제목 키워드 기반

Gaussian Stochastic Weight Averaging for Bayesian Low-Rank Adaptation of Large Language Models

2024-05-06 · Emre Onal, Klemens Flöge, Emma Caldwell, Arsen Sheverdin 외

Fine-tuned Large Language Models (LLMs) often suffer from overconfidence and poor calibration, particularly when fine-tuned on small datasets. To address these challenges, we propose a simple combination of Low-Rank Adap…

Bayesian Inference

Language as Prior, Vision as Calibration: Metric Scale Recovery for Monocular Depth Estimation

2026-01-04 · Mingxia Zhan, Li Zhang, Beibei Wang, Yingjie Wang 외 arxiv

Relative-depth foundation models transfer well, yet monocular metric depth remains ill-posed due to unidentifiable global scale and heightened domain-shift sensitivity. Under a frozen-backbone calibration setting, we rec…

Monocular Depth Estimation

UncertaintyRAG: Span-Level Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation

2024-10-03 · Zixuan Li, Jing Xiong, Fanghua Ye, Chuanyang Zheng 외

We present UncertaintyRAG, a novel approach for long-context Retrieval-Augmented Generation (RAG) that utilizes Signal-to-Noise Ratio (SNR)-based span uncertainty to estimate similarity between text chunks. This span unc…

ChunkingLanguage ModelingLanguage ModellingLarge Language Model+3

An Empirical Study on Robustness to Spurious Correlations using Pre-trained Language Models

2020-07-14 · Lifu Tu, Garima Lalwani, Spandana Gella, He He

Recent work has shown that pre-trained language models such as BERT improve robustness to spurious correlations in the dataset. Intrigued by these results, we find that the key to their success is generalization from a s…

DiversityMulti-Task LearningNatural Language InferenceParaphrase Identification

Foundation Feature-Driven Online End-Effector Pose Estimation: A Marker-Free and Learning-Free Approach

2025-03-18 · Tianshu Wu, Jiyao Zhang, Shiqian Liang, Zhengxiao Han 외

Accurate transformation estimation between camera space and robot space is essential. Traditional methods using markers for hand-eye calibration require offline image collection, limiting their suitability for online sel…

6D Pose EstimationPose EstimationRobot Pose EstimationZero-shot Generalization