paper-with-me

Papers

Word Embedding based Edit Distance

2018-10-25 · Yilin Niu, chao qiao, Hang Li, Minlie Huang

Text similarity calculation is a fundamental problem in natural language processing and related fields. In recent years, deep neural networks have been developed to perform the task and high performances have been achieved. The neural networks are usually trained with labeled data in supervised learning, and creation of labeled data is usually very costly. In this short paper, we address unsupervised learning for text similarity calculation. We propose a new method called Word Embedding based Edit Distance (WED), which incorporates word embedding into edit distance. Experiments on three benchmark datasets show WED outperforms state-of-the-art unsupervised methods including edit distance, TF-IDF based cosine, word embedding based cosine, Jaccard index, etc.

📄 PDF Abstract BibTeX arXiv:1810.10752

Code (0)

등록된 구현이 없습니다.

Tasks

text similarity

Similar Papers 제목 키워드 기반

Paradigm Clustering with Weighted Edit Distance

2021-08-01 · ACL (SIGMORPHON) 2021 8 · Andrew Gerlach, Adam Wiemerslage, Katharina Kann

This paper describes our system for the SIGMORPHON 2021 Shared Task on Unsupervised Morphological Paradigm Clustering, which asks participants to group inflected forms together according their underlying lemma without th…

ClusteringLEMMAWord Embeddings

Unsupervised Lemmatization as Embeddings-Based Word Clustering

2019-08-22 · Rudolf Rosa, Zdeněk Žabokrtský

We focus on the task of unsupervised lemmatization, i.e. grouping together inflected forms of one word under one label (a lemma) without the use of annotated training data. We propose to perform agglomerative clustering …

ClusteringLEMMALemmatization

A Theoretical Framework for Acoustic Neighbor Embeddings

2024-12-03 · Woojay Jeon

This paper provides a theoretical framework for interpreting acoustic neighbor embeddings, which are representations of the phonetic content of variable-width audio or text in a fixed-dimensional embedding space. A proba…

Clustering

Distance-to-Distance Ratio: A Similarity Measure for Sentences Based on Rate of Change in LLM Embeddings

2026-01-25 · Abdullah Qureshi, Kenneth Rice, Alexander Wolpert arxiv

A measure of similarity between text embeddings can be considered adequate only if it adheres to the human perception of similarity between texts. In this paper, we introduce the distance-to-distance ratio (DDR), a novel…

Similarity-Based Unsupervised Spelling Correction Using BioWordVec: Development and Usability Study of Bacterial Culture and Antimicrobial Susceptibility Reports

2021-02-22 · JMIR Medical Informatics 2021 2 · Taehyeong Kim, Sung Won Han, Minji Kang, Se Ha Lee 외

Background: Existing bacterial culture test results for infectious diseases are written in unrefined text, resulting in many problems, including typographical errors and stop words. Effective spelling correction process…

Cultural Vocal Bursts Intensity PredictionSpelling Correction