paper-with-me

Papers

Word Embeddings for the Construction Domain

2016-10-28 · Antoine J. -P. Tixier, Michalis Vazirgiannis, Matthew R. Hallowell

We introduce word vectors for the construction domain. Our vectors were obtained by running word2vec on an 11M-word corpus that we created from scratch by leveraging freely-accessible online sources of construction-related text. We first explore the embedding space and show that our vectors capture meaningful construction-specific concepts. We then evaluate the performance of our vectors against that of ones trained on a 100B-word corpus (Google News) within the framework of an injury report classification task. Without any parameter tuning, our embeddings give competitive results, and outperform the Google News vectors in many cases. Using a keyword-based compression of the reports also leads to a significant speed-up with only a limited loss in performance. We release our corpus and the data set we created for the classification task as publicly available, in the hope that they will be used by future studies for benchmarking and building on our work.

📄 PDF Abstract BibTeX arXiv:1610.09333

Code (1)

Tixierae/WECD 공식 구현

Tasks

BenchmarkingGeneral ClassificationWord Embeddings

Similar Papers 제목 키워드 기반

Expert Concept-Modeling Ground Truth Construction for Word Embeddings Evaluation in Concept-Focused Domains

2020-12-01 · COLING 2020 8 · Arianna Betti, Martin Reynaert, Thijs Ossenkoppele, Yvette Oortwijn 외

We present a novel, domain expert-controlled, replicable procedure for the construction of concept-modeling ground truths with the aim of evaluating the application of word embeddings. In particular, our method is design…

Embeddings EvaluationPhilosophyWord Embeddings

Embeddings models for Buddhist Sanskrit

2022-06-01 · LREC 2022 6 · Ligeia Lugli, Matej Martinc, Andraž Pelicon, Senja Pollak

The paper presents novel resources and experiments for Buddhist Sanskrit, broadly defined here including all the varieties of Sanskrit in which Buddhist texts have been transmitted. We release a novel corpus of Buddhist …

Semantic SimilaritySemantic Textual SimilarityTransfer LearningWord Similarity

Tiny Word Embeddings Using Globally Informed Reconstruction

2020-12-01 · COLING 2020 8 · Sora Ohashi, Mao Isogawa, Tomoyuki Kajiwara, Yuki Arase

We reduce the model size of pre-trained word embeddings by a factor of 200 while preserving its quality. Previous studies in this direction created a smaller word embedding model by reconstructing pre-trained word repres…

Word EmbeddingsWord Similarity

Subword-based Compact Reconstruction of Word Embeddings

2019-06-01 · NAACL 2019 6 · Shota Sasaki, Jun Suzuki, Kentaro Inui

The idea of subword-based word embeddings has been proposed in the literature, mainly for solving the out-of-vocabulary (OOV) word problem observed in standard word-based word embeddings. In this paper, we propose a meth…

Word Embeddings

Yseop at FinSim-3 Shared Task 2021: Specializing Financial Domain Learning with Phrase Representations

2021-08-21 · FinNLP 2021 8 · Hanna Abi Akl, Dominique Mariko, Hugues de Mazancourt

In this paper, we present our approaches for the FinSim-3 Shared Task 2021: Learning Semantic Similarities for the Financial Domain. The aim of this shared task is to correctly classify a list of given terms from the fin…

SentenceSentence EmbeddingsWord Embeddings