paper-with-me

Papers

Distributed Distributional Similarities of Google Books Over the Centuries

2014-05-01 · LREC 2014 5 · Martin Riedl, Richard Steuer, Chris Biemann

This paper introduces a distributional thesaurus and sense clusters computed on the complete Google Syntactic N-grams, which is extracted from Google Books, a very large corpus of digitized books published between 1520 and 2008. We show that a thesaurus computed on such a large text basis leads to much better results than using smaller corpora like Wikipedia. We also provide distributional thesauri for equal-sized time slices of the corpus. While distributional thesauri can be used as lexical resources in NLP tasks, comparing word similarities over time can unveil sense change of terms across different decades or centuries, and can serve as a resource for diachronic lexicography. Thesauri and clusters are available for download.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Graph Clustering

Similar Papers 제목 키워드 기반

A Web Interface for Diachronic Semantic Search in Spanish

2017-04-01 · EACL 2017 4 · Pablo Gamallo, Iv{\'a}n Rodr{\'\i}guez-Torres, Marcos Garcia

This article describes a semantic system which is based on distributional models obtained from a chronologically structured language resource, namely Google Books Syntactic Ngrams.The models were created using dependency…

Characterizing the Google Books corpus: Strong limits to inferences of socio-cultural and linguistic evolution

2015-01-05 · Eitan Adam Pechenick, Christopher M. Danforth, Peter Sheridan Dodds

It is tempting to treat frequency trends from the Google Books data sets as indicators of the "true" popularity of various words and phrases. Doing so allows us to draw quantitatively strong conclusions about the evoluti…

Articles

Dynamics of core of language vocabulary

2017-05-29 · Valery D. Solovyev, Vladimir V. Bochkarev, Anna V. Shevlyakova

Studies of the overall structure of vocabulary and its dynamics became possible due to creation of diachronic text corpora, especially Google Books Ngram. This article discusses the question of core change rate and the d…

Verifying Heaps' law using Google Books Ngram data

2016-12-29 · Vladimir V. Bochkarev, Eduard Yu. Lerner, Anna V. Shevlyakova

This article is devoted to the verification of the empirical Heaps law in European languages using Google Books Ngram corpus data. The connection between word distribution frequency and expected dependence of individual …

Text Generation

Google Books N-gram Corpus used as a Grammar Checker

2012-04-01 · WS 2012 4 · Rogelio Nazar, Irene Renau