paper-with-me

홈 › Papers

LEA: Improving Sentence Similarity Robustness to Typos Using Lexical Attention Bias

2023-07-06 · Mario Almagro, Emilio Almazán, Diego Ortego, David Jiménez

Textual noise, such as typos or abbreviations, is a well-known issue that penalizes vanilla Transformers for most downstream tasks. We show that this is also the case for sentence similarity, a fundamental task in multiple domains, e.g. matching, retrieval or paraphrasing. Sentence similarity can be approached using cross-encoders, where the two sentences are concatenated in the input allowing the model to exploit the inter-relations between them. Previous works addressing the noise issue mainly rely on data augmentation strategies, showing improved robustness when dealing with corrupted samples that are similar to the ones used for training. However, all these methods still suffer from the token distribution shift induced by typos. In this work, we propose to tackle textual noise by equipping cross-encoders with a novel LExical-aware Attention module (LEA) that incorporates lexical similarities between words in both sentences. By using raw text similarities, our approach avoids the tokenization shift problem obtaining improved robustness. We demonstrate that the attention bias introduced by LEA helps cross-encoders to tackle complex scenarios with textual noise, specially in domains with short-text descriptions and limited context. Experiments using three popular Transformer encoders in five e-commerce datasets for product matching show that LEA consistently boosts performance under the presence of noise, while remaining competitive on the original (clean) splits. We also evaluate our approach in two datasets for textual entailment and paraphrasing showing that LEA is robust to typos in domains with longer sentences and more natural context. Additionally, we thoroughly analyze several design choices in our approach, providing insights about the impact of the decisions made and fostering future research in cross-encoders dealing with typos.

📄 PDF Abstract BibTeX arXiv:2307.02912

Code (1)

m-almagro-cadiz/lea 공식 구현 pytorch

Tasks

Data AugmentationNatural Language InferenceSentenceSentence Similarity

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Making Sentence Embeddings Robust to User-Generated Content

2024-03-25 · Lydia Nishimwe, Benoît Sagot, Rachel Bawden

NLP models have been known to perform poorly on user-generated content (UGC), mainly because it presents a lot of lexical variations and deviates from the standard texts on which most of these models were trained. In thi…

SentenceSentence EmbeddingSentence-EmbeddingSentence Embeddings

Robust Encodings: A Framework for Combating Adversarial Typos

2020-05-04 · ACL 2020 6 · Erik Jones, Robin Jia, aditi raghunathan, Percy Liang

Despite excellent performance on many tasks, NLP systems are easily fooled by small adversarial perturbations of inputs. Existing procedures to defend against such perturbations are either (i) heuristic in nature and sus…

Sentence

EdaCSC: Two Easy Data Augmentation Methods for Chinese Spelling Correction

2024-09-08 · Lei Sheng, Shuai-Shuai Xu

Chinese Spelling Correction (CSC) aims to detect and correct spelling errors in Chinese sentences caused by phonetic or visual similarities. While current CSC models integrate pinyin or glyph features and have shown sign…

Data AugmentationSpelling Correction

CharacterBERT and Self-Teaching for Improving the Robustness of Dense Retrievers on Queries with Typos

2022-04-01 · Shengyao Zhuang, Guido Zuccon

Current dense retrievers are not robust to out-of-domain and outlier queries, i.e. their effectiveness on these queries is much poorer than what one would expect. In this paper, we consider a specific instance of such qu…

Passage RetrievalRetrieval

Fast, Effective, and Self-Supervised: Transforming Masked Language Models into Universal Lexical and Sentence Encoders

2021-04-16 · EMNLP 2021 11 · Fangyu Liu, Ivan Vulić, Anna Korhonen, Nigel Collier

Pretrained Masked Language Models (MLMs) have revolutionised NLP in recent years. However, previous work has indicated that off-the-shelf MLMs are not effective as universal lexical or sentence encoders without further t…

Contrastive LearningCross-Lingual Semantic Textual SimilarityEntity LinkingSemantic Similarity+4