Unfinished Business: Construction and Maintenance of a Semantically Tagged Historical Parliamentary Corpus, UK Hansard from 1803 to the present day
Creating, curating and maintaining modern political corpora is becoming an ever more involved task. As interest from various social bodies and the general public in political discourse grows so too does the need to enrich such datasets with metadata and linguistic annotations. Beyond this, such corpora must be easy to browse and search for linguists, social scientists, digital humanists and the general public. We present our efforts to compile a linguistically annotated and semantically tagged version of the Hansard corpus from 1803 right up to the present day. This involves combining multiple sources of documents and transcripts. We describe our toolchain for tagging; using several existing tools that provide tokenisation, part-of-speech tagging and semantic annotations. We also provide an overview of our bespoke web-based search interface built on LexiDB. In conclusion, we examine the completed corpus by looking at four case studies including semantic categories made available by our toolchain.
Code (0)
등록된 구현이 없습니다.
Tasks
Part-Of-Speech TaggingSimilar Papers 제목 키워드 기반
Using Fluorescence Recovery After Photobleaching (FRAP) to study dynamics of the Structural Maintenance of Chromosome (SMC) complex in vivo
The SMC complex, MukBEF, is important for chromosome organization and segregation in Escherichia coli. Fluorescently tagged MukBEF forms distinct spots (or 'foci') in the cell, where it is thought to carry out most of it…
Unfinished Architectures: A Perspective from Artificial Intelligence
Unfinished buildings are a constant throughout the history of architecture and have given rise to intense debates on the opportuneness of their completion, in addition to offering alibis for theorizing about the composit…
Cost-Sensitive Learning for Predictive Maintenance
In predictive maintenance, model performance is usually assessed by means of precision, recall, and F1-score. However, employing the model with best performance, e.g. highest F1-score, does not necessarily result in mini…
Model SelectionCoconstructions in spoken data: UD annotation guidelines and first results
The paper proposes annotation guidelines for syntactic dependencies that span across speaker turns - including collaborative coconstructions proper, wh-question answers, and backchannels - in spoken language treebanks wi…
QurAna: Corpus of the Quran annotated with Pronominal Anaphora
This paper presents QurAna: a large corpus created from the original Quranic text, where personal pronouns are tagged with their antecedence. These antecedents are maintained as an ontological list of concepts, which hav…
Coreference ResolutionInformation RetrievalMachine TranslationQuestion Answering+1