Exploring a Continuous and Flexible Representation of the Lexicon
We aim at showing that lexical descriptions based on multifactorial and continuous models can be used by linguists and lexicographers (and not only by machines) so long as they are provided with a way to efficiently navigate data collections. We propose to demonstrate such a system.
Code (0)
등록된 구현이 없습니다.
Tasks
NavigateSemantic Textual SimilaritySimilar Papers 제목 키워드 기반
Exploring Word Embeddings for Unsupervised Textual User-Generated Content Normalization
Text normalization techniques based on rules, lexicons or supervised training requiring large corpora are not scalable nor domain interchangeable, and this makes them unsuitable for normalizing user-generated content (UG…
Semantic SimilaritySemantic Textual SimilarityText NormalizationWord EmbeddingsAlleviating Overfitting for Polysemous Words for Word Representation Estimation Using Lexicons
Though there are some works on improving distributed word representations using lexicons, the improper overfitting of the words that have multiple meanings is a remaining issue deteriorating the learning when lexicons ar…
LexLIP: Lexicon-Bottlenecked Language-Image Pre-Training for Large-Scale Image-Text Retrieval
Image-text retrieval (ITR) is a task to retrieve the relevant images/texts, given the query from another modality. The conventional dense retrieval paradigm relies on encoding images and texts into dense representations …
Image-text RetrievalRetrievalText RetrievalFecTek: Enhancing Term Weight in Lexicon-Based Retrieval with Feature Context and Term-level Knowledge
Lexicon-based retrieval has gained siginificant popularity in text retrieval due to its efficient and robust performance. To further enhance performance of lexicon-based retrieval, researchers have been diligently incorp…
Contrastive LearningRetrievalText RetrievalDP-Parse: Finding Word Boundaries from Raw Speech with an Instance Lexicon
Finding word boundaries in continuous speech is challenging as there is little or no equivalent of a 'space' delimiter between words. Popular Bayesian non-parametric models for text segmentation use a Dirichlet process t…
Language ModelingLanguage ModellingSegmentationText Segmentation