paper-with-me

홈 › Papers

Improving Contextual Representation with Gloss Regularized Pre-training

2022-05-13 · Findings (NAACL) 2022 7 · Yu Lin, Zhecheng An, Peihao Wu, Zejun Ma

Though achieving impressive results on many NLP tasks, the BERT-like masked language models (MLM) encounter the discrepancy between pre-training and inference. In light of this gap, we investigate the contextual representation of pre-training and inference from the perspective of word probability distribution. We discover that BERT risks neglecting the contextual word similarity in pre-training. To tackle this issue, we propose an auxiliary gloss regularizer module to BERT pre-training (GR-BERT), to enhance word semantic similarity. By predicting masked words and aligning contextual embeddings to corresponding glosses simultaneously, the word similarity can be explicitly modeled. We design two architectures for GR-BERT and evaluate our model in downstream tasks. Experimental results show that the gloss regularizer benefits BERT in word-level and sentence-level semantic representation. The GR-BERT achieves new state-of-the-art in lexical substitution task and greatly promotes BERT sentence representation in both unsupervised and supervised STS tasks.

📄 PDF Abstract BibTeX arXiv:2205.06603

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SimilaritySemantic Textual SimilaritySentenceSTSWord Similarity

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
WordPiece 설명 없음
Weight Decay 설명 없음
Multi-Head Attention 설명 없음
Attention Dropout Attention Dropout is a type of dropout used in attention-based architectures, where elements are randomly dropped out of the…
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…

Similar Papers 제목 키워드 기반

Improving Contextual Representation with Gloss Regularized Pre-training

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Though achieving impressive results on many NLP tasks, the BERT-like masked language models (MLM) encounter the discrepancy between pre-training and inference. In light of this gap, we investigate the contextual represen…

Semantic SimilaritySemantic Textual SimilaritySentenceSTS+1

C2ST: Cross-Modal Contextualized Sequence Transduction for Continuous Sign Language Recognition

2023-01-01 · ICCV 2023 1 · Huaiwen Zhang, Zihang Guo, Yang Yang, Xin Liu 외

Continuous Sign Language Recognition (CSLR) aims to transcribe the signs of an untrimmed video into written words or glosses. The mainstream framework for CSLR consists of a spatial module for visual representation l…

Language ModellingRepresentation LearningSign Language Recognition

Vec2Gloss: definition modeling leveraging contextualized vectors with Wordnet gloss

2023-05-29 · Yu-Hsiang Tseng, Mao-Chang Ku, Wei-Ling Chen, Yu-Lin Chang 외

Contextualized embeddings are proven to be powerful tools in multiple NLP tasks. Nonetheless, challenges regarding their interpretability and capability to represent lexical semantics still remain. In this paper, we prop…

C${^2}$RL: Content and Context Representation Learning for Gloss-free Sign Language Translation and Retrieval

2024-08-19 · Zhigang Chen, Benjia Zhou, Yiqing Huang, Jun Wan 외

Sign Language Representation Learning (SLRL) is crucial for a range of sign language-related downstream tasks such as Sign Language Translation (SLT) and Sign Language Retrieval (SLRet). Recently, many gloss-based and gl…

Gloss-free Sign Language TranslationRepresentation LearningRhythmSign Language Retrieval+1

ViConBERT: Context-Gloss Aligned Vietnamese Word Embedding for Polysemous and Sense-Aware Representations

2025-11-15 · Khang T. Huynh, Dung H. Nguyen, Binh T. Nguyen arxiv

Recent advances in contextualized word embeddings have greatly improved semantic tasks such as Word Sense Disambiguation (WSD) and contextual similarity, but most progress has been limited to high-resource languages like…

Word Sense DisambiguationContrastive Learning