paper-with-me

홈 › Papers

DLT-Corpus: A Large-Scale Text Collection for the Distributed Ledger Technology Domain

2026-02-25 · Walter Hernandez Cruz, Peter Devine, Nikhil Vadgama, Paolo Tasca, Jiahua Xu arxiv

We introduce DLT-Corpus, the largest domain-specific text collection for Distributed Ledger Technology (DLT) research to date: 2.98 billion tokens from 22.12 million documents spanning scientific literature (37,440 publications), United States Patent and Trademark Office (USPTO) patents (49,023 filings), and social media (22 million posts). Existing Natural Language Processing (NLP) resources for DLT focus narrowly on cryptocurrency price prediction and smart contracts, leaving domain-specific language underexplored despite the sector's ~$3 trillion market capitalization and rapid technological evolution. We demonstrate DLT-Corpus' utility by analyzing patterns of technology emergence and market-innovation correlations. Findings reveal that technologies first appear in our scientific literature subset before reaching patents and social media, following traditional technology transfer patterns. While social media sentiment remains overwhelmingly bullish even during crypto winters, scientific and patent activity grows less tied to short-term sentiment, tracking overall market expansion in a virtuous cycle in which research precedes and enables economic growth that, in turn, funds further innovation. We release the DLT-Corpus and companion artifacts: LedgerBERT (+23% over BERT-base on DLT-specific Named Entity Recognition (NER) task), a sentiment analysis dataset of 23,301 crypto news headlines and descriptions, tools, and code.

📄 PDF Abstract BibTeX arXiv:2602.22045

Code (0)

등록된 구현이 없습니다.

Tasks

Sentiment Analysis

Similar Papers 제목 키워드 기반

Collection of a corpus of Dutch SMS

2012-05-01 · LREC 2012 5 · Maaske Treurniet, Orph{\'e}e De Clercq, Henk van den Heuvel, Nelleke Oostdijk

In this paper we present the first freely available corpus of Dutch text messages containing data originating from the Netherlands and Flanders. This corpus has been collected in the framework of the SoNaR project and co…

Document retrieval and question answering in medical documents. A large-scale corpus challenge.

2017-09-01 · RANLP 2017 9 · Curea Eric

Whenever employed on large datasets, information retrieval works by isolating a subset of documents from the larger dataset and then proceeding with low-level processing of the text. This is usually carried out by means …

Document ClassificationGeneral ClassificationInformation RetrievalQuestion Answering+1

Publishing the Trove Newspaper Corpus

2016-05-01 · LREC 2016 5 · Steve Cassidy

The Trove Newspaper Corpus is derived from the National Library of Australia{'}s digital archive of newspaper text. The corpus is a snapshot of the NLA collection taken in 2015 to be made available for language research …

Articles

Infrastructure for Semantic Annotation in the Genomics Domain

2020-05-01 · LREC 2020 5 · Mahmoud El-Haj, Nathan Rutherford, Matthew Coole, Ignatius Ezeani 외

We describe a novel super-infrastructure for biomedical text mining which incorporates an end-to-end pipeline for the collection, annotation, storage, retrieval and analysis of biomedical and life sciences literature, co…

Retrieval

GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training

2026-04-01 · Jesse van Oort, Frank Brinkkemper, Erik de Graaf, Bram Vanroy 외 arxiv

We present the GPT-NL Public Corpus, the biggest permissively licensed corpus of Dutch language resources. The GPT-NL Public Corpus contains 21 Dutch-only collections totalling 36B preprocessed Dutch tokens not present i…