paper-with-me

홈 › Papers

Publishing the Trove Newspaper Corpus

2016-05-01 · LREC 2016 5 · Steve Cassidy

The Trove Newspaper Corpus is derived from the National Library of Australia{'}s digital archive of newspaper text. The corpus is a snapshot of the NLA collection taken in 2015 to be made available for language research as part of the Alveo Virtual Laboratory and contains 143 million articles dating from 1806 to 2007. This paper describes the work we have done to make this large corpus available as a research collection, facilitating access to individual documents and enabling large scale processing of the newspaper text in a cloud-based environment.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

The eIdentity Text Exploration Workbench

2014-05-01 · LREC 2014 5 · Fritz Kliche, Andr{\'e} Blessing, Ulrich Heid, Jonathan Sonntag

We work on tools to explore text contents and metadata of newspaper articles as provided by news archives. Our tool components are being integrated into an {``}Exploration Workbench{''} for Digital Humanities researchers…

ArticlesInformation RetrievalNamed Entity Recognition (NER)Retrieval

Finding Names in Trove: Named Entity Recognition for Australian Historical Newspapers

2015-12-01 · ALTA 2015 12 · Sunghwan Mac Kim, Steve Cassidy
Clusteringnamed-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)+1

The IPR-cleared Corpus of Contemporary Written and Spoken Romanian Language

2016-05-01 · LREC 2016 5 · Dan Tufi{\textcommabelow{s}}, Verginica Barbu Mititelu, Elena Irimia, {\textcommabelow{S}}tefan Daniel Dumitrescu 외

The article describes the current status of a large national project, CoRoLa, aiming at building a reference corpus for the contemporary Romanian language. Unlike many other national corpora, CoRoLa contains only - IPR c…

LemmatizationPart-Of-Speech Tagging

Lexical and Statistical Analysis of Bangla Newspaper and Literature: A Corpus-Driven Study on Diversity, Readability, and NLP Adaptation

2025-12-15 · Pramit Bhattacharyya, Arnab Bhattacharya arxiv

In this paper, we present a comprehensive corpus-driven analysis of Bangla literary and newspaper texts to investigate their lexical diversity, structural complexity and readability. We undertook Vacaspati and IndicCorp,…

Which Factors Impact Engagement on News Articles on Facebook?

2019-10-31 · Marc Faddoul

Social media is increasingly being used as a news-platform. To reach their intended audience, newspapers need for their articles to be well ranked by Facebook's news-feed algorithm. The number of likes, shares and other …

Articles