paper-with-me

Papers

WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations

2018-08-28 · NAACL 2019 6 · Mohammad Taher Pilehvar, Jose Camacho-Collados

By design, word embeddings are unable to model the dynamic nature of words' semantics, i.e., the property of words to correspond to potentially different meanings. To address this limitation, dozens of specialized meaning representation techniques such as sense or contextualized embeddings have been proposed. However, despite the popularity of research on this topic, very few evaluation benchmarks exist that specifically focus on the dynamic semantics of words. In this paper we show that existing models have surpassed the performance ceiling of the standard evaluation dataset for the purpose, i.e., Stanford Contextual Word Similarity, and highlight its shortcomings. To address the lack of a suitable benchmark, we put forward a large-scale Word in Context dataset, called WiC, based on annotations curated by experts, for generic evaluation of context-sensitive representations. WiC is released in https://pilehvar.github.io/wic/.

📄 PDF Abstract BibTeX arXiv:1808.09121

Code (0)

등록된 구현이 없습니다.

Tasks

Word EmbeddingsWord Sense DisambiguationWord Similarity

Similar Papers 제목 키워드 기반

Word2rate: training and evaluating multiple word embeddings as statistical transitions

2021-04-16 · Gary Phua, Shaowei Lin, Dario Poletti

Using pretrained word embeddings has been shown to be a very effective way in improving the performance of natural language processing tasks. In fact almost any natural language tasks that can be thought of has been impr…

Sentiment AnalysisTranslationWord Embeddings

See No Evil: Semantic Context-Aware Privacy Risk Detection for AR

2026-04-14 · Jialu Liu, Yao Li, Zhuoheng Li, Huining Li 외 arxiv

Augmented reality (AR) systems pose unique privacy risks due to their continuous capture of visual data. Existing AR privacy frameworks lack semantic understanding of visual content, limiting their effectiveness in detec…

Chinese WPLC: A Chinese Dataset for Evaluating Pretrained Language Models on Word Prediction Given Long-Range Context

2021-11-01 · EMNLP 2021 11 · Huibin Ge, Chenxi Sun, Deyi Xiong, Qun Liu

This paper presents a Chinese dataset for evaluating pretrained language models on Word Prediction given Long-term Context (Chinese WPLC). We propose both automatic and manual selection strategies tailored to Chinese to …

DiversityLanguage ModelingLanguage Modelling

CoSimLex: A Resource for Evaluating Graded Word Similarity in Context

2019-12-11 · LREC 2020 5 · Carlos Santos Armendariz, Matthew Purver, Matej Ulčar, Senja Pollak 외

State of the art natural language processing tools are built on context-dependent word embeddings, but no direct method for evaluating these representations currently exists. Standard tasks and datasets for intrinsic eva…

Word EmbeddingsWord Sense DisambiguationWord Similarity

A Comparison of Context-sensitive Models for Lexical Substitution

2019-05-01 · WS 2019 5 · Aina Gar{\'\i} Soler, Anne Cocos, Marianna Apidianaki, Chris Callison-Burch

Word embedding representations provide good estimates of word meaning and give state-of-the art performance in semantic tasks. Embedding approaches differ as to whether and how they account for the context surrounding a …

Word Embeddings