paper-with-me

홈 › Papers

Integrating Task Specific Information into Pretrained Language Models for Low Resource Fine Tuning

2020-11-01 · Findings of the Association for Computational Linguistics 2020 · Rui Wang, Shijing Si, Guoyin Wang, Lei Zhang, Lawrence Carin, Ricardo Henao

Pretrained Language Models (PLMs) have improved the performance of natural language understanding in recent years. Such models are pretrained on large corpora, which encode the general prior knowledge of natural languages but are agnostic to information characteristic of downstream tasks. This often results in overfitting when fine-tuned with low resource datasets where task-specific information is limited. In this paper, we integrate label information as a task-specific prior into the self-attention component of pretrained BERT models. Experiments on several benchmarks and real-word datasets suggest that the proposed approach can largely improve the performance of pretrained models when fine-tuning with small datasets.

📄 PDF Abstract BibTeX

Code (1)

raywangwr/bert_label_embedding 공식 구현 pytorch

Tasks

Natural Language Understanding

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Linear Warmup With Linear Decay Linear Warmup With Linear Decay is a learning rate schedule in which we increase the learning rate linearly for $n$ updates and then linearly decay afterwards.
Residual Connection 설명 없음
WordPiece 설명 없음

Similar Papers 제목 키워드 기반

Contextual information integration for stance detection via cross-attention

2022-11-03 · Tilman Beck, Andreas Waldis, Iryna Gurevych

Stance detection deals with identifying an author's stance towards a target. Most existing stance detection models are limited because they do not consider relevant contextual information which allows for inferring the s…

Language ModelingLanguage ModellingStance Detection

Towards Table-to-Text Generation with Pretrained Language Model: A Table Structure Understanding and Text Deliberating Approach

2023-01-05 · Miao Chen, Xinjiang Lu, Tong Xu, Yanyan Li 외

Although remarkable progress on the neural table-to-text methods has been made, the generalization issues hinder the applicability of these models due to the limited source tables. Large-scale pretrained language models …

DecoderDescriptiveLanguage ModelingLanguage Modelling+2

When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language Models

2026-08-20 · Su Yan, Rakesh Iyer arxiv

Many real-world AI systems represent entities, behaviors, and structured information using discrete machine-native symbols rather than natural language. While these representations are compact and preserve task-relevant …

Sequential RecommendationStructured Prediction

X-Fusion: Introducing New Modality to Frozen Large Language Models

2025-04-29 · Sicheng Mo, Thao Nguyen, Xun Huang, Siddharth Srinivasan Iyer 외

We propose X-Fusion, a framework that extends pretrained Large Language Models (LLMs) for multimodal tasks while preserving their language capabilities. X-Fusion employs a dual-tower design with modality-specific weights…

Image to text

DLFR-VAE: Dynamic Latent Frame Rate VAE for Video Generation

2025-02-17 · Zhihang Yuan, Siyuan Wang, Rui Xie, Hanling Zhang 외

In this paper, we propose the Dynamic Latent Frame Rate VAE (DLFR-VAE), a training-free paradigm that can make use of adaptive temporal compression in latent space. While existing video generative models apply fixed comp…

Video Generation