paper-with-me

홈 › Papers

Making Sentence Embeddings Robust to User-Generated Content

2024-03-25 · Lydia Nishimwe, Benoît Sagot, Rachel Bawden

NLP models have been known to perform poorly on user-generated content (UGC), mainly because it presents a lot of lexical variations and deviates from the standard texts on which most of these models were trained. In this work, we focus on the robustness of LASER, a sentence embedding model, to UGC data. We evaluate this robustness by LASER's ability to represent non-standard sentences and their standard counterparts close to each other in the embedding space. Inspired by previous works extending LASER to other languages and modalities, we propose RoLASER, a robust English encoder trained using a teacher-student approach to reduce the distances between the representations of standard and UGC sentences. We show that with training only on standard and synthetic UGC-like data, RoLASER significantly improves LASER's robustness to both natural and artificial UGC data by achieving up to 2x and 11x better scores. We also perform a fine-grained analysis on artificial UGC data and find that our model greatly outperforms LASER on its most challenging UGC phenomena such as keyboard typos and social media abbreviations. Evaluation on downstream tasks shows that RoLASER performs comparably to or better than LASER on standard data, while consistently outperforming it on UGC data.

📄 PDF Abstract BibTeX arXiv:2403.17220

Code (1)

lydianish/rolaser 공식 구현

Tasks

SentenceSentence EmbeddingSentence-EmbeddingSentence Embeddings

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Building a User-Generated Content North-African Arabizi Treebank: Tackling Hell

2020-07-01 · ACL 2020 6 · Djam{\'e} Seddah, Farah Essaidi, Amal Fethi, Matthieu Futeral 외

We introduce the first treebank for a romanized user-generated content variety of Algerian, a North-African Arabic dialect known for its frequent usage of code-switching. Made of 1500 sentences, fully annotated in morpho…

Dependency ParsingPOSPOS TaggingSentence+1

Sentence-level Privacy for Document Embeddings

2021-11-16 · ACL ARR November 2021 11 · Anonymous

User language data can contain highly sensitive personal content. As such, it is imperative to offer users a strong and interpretable privacy guarantee when learning from their data. In this work we propose SentDP, pure …

Language ModelingLanguage ModellingSentenceSentiment Analysis+1

Sentence-level Privacy for Document Embeddings

2022-05-10 · ACL 2022 5 · Casey Meehan, Khalil Mrini, Kamalika Chaudhuri

User language data can contain highly sensitive personal content. As such, it is imperative to offer users a strong and interpretable privacy guarantee when learning from their data. In this work, we propose SentDP: pure…

Language ModelingLanguage ModellingSentenceSentiment Analysis+1

BioReddit: Word Embeddings for User-Generated Biomedical NLP

2019-11-01 · WS 2019 11 · Marco Basaldella, Nigel Collier

Word embeddings, in their different shapes and iterations, have changed the natural language processing research landscape in the last years. The biomedical text processing field is no stranger to this revolution; howeve…

Word Embeddings

Mechanistic Decomposition of Sentence Representations

2025-06-04 · Matthieu Tehenan, Vikram Natarajan, Jonathan Michala, Milton Lin 외

Sentence embeddings are central to modern NLP and AI systems, yet little is known about their internal structure. While we can compare these embeddings using measures such as cosine similarity, the contributing features …

Dictionary LearningSentenceSentence EmbeddingSentence-Embedding+1