paper-with-me

홈 › Papers

KwicKwocKwac, a tool for rapidly generating concordances and marking up a literary text

2024-10-08 · Sebastian Barzaghi, Francesco Paolucci, Francesca Tomasi, Fabio Vitali

This paper introduces KwicKwocKwac 1.0 (KwicKK), a web application designed to enhance the annotation and enrichment of digital texts in the humanities. KwicKK provides a user-friendly interface that enables scholars and researchers to perform semi-automatic markup of textual documents, facilitating the identification of relevant entities such as people, organizations, and locations. Key functionalities include the visualization of annotated texts using KeyWord in Context (KWIC), KeyWord Out Of Context (KWOC), and KeyWord After Context (KWAC) methodologies, alongside automatic disambiguation of generic references and integration with Wikidata for Linked Open Data connections. The application supports metadata input and offers multiple download formats, promoting accessibility and ease of use. Developed primarily for the National Edition of Aldo Moro's works, KwicKK aims to lower the technical barriers for users while fostering deeper engagement with digital scholarly resources. The architecture leverages contemporary web technologies, ensuring scalability and reliability. Future developments will explore user experience enhancements, collaborative features, and integration of additional data sources.

📄 PDF Abstract BibTeX arXiv:2410.06043

Code (1)

sanofrank/KwicKwocKwac 공식 구현

Similar Papers 제목 키워드 기반

Kunji : A Resource Management System for Higher Productivity in Computer Aided Translation Tools

2019-12-01 · ICON 2019 12 · Priyank Gupta, Manish Shrivastava, Dipti Misra Sharma, Rashid Ahmad

Complex NLP applications, such as machine translation systems, utilize various kinds of resources namely lexical, multiword, domain dictionaries, maps and rules etc. Similarly, translators working on Computer Aided Trans…

Machine TranslationManagementNERTranslation

VPS-GradeUp: Graded Decisions on Usage Patterns

2016-05-01 · LREC 2016 5 · V{\'\i}t Baisa, Silvie Cinkov{\'a}, Ema Krej{\v{c}}ov{\'a}, Anna Vernerov{\'a}

We present VPS-GradeUp ― a set of 11,400 graded human decisions on usage patterns of 29 English lexical verbs from the Pattern Dictionary of English Verbs by Patrick Hanks. The annotation contains, for each verb lemma,…

ClusteringLEMMA

Harmonization Benchmarking Tool for Neuroimaging Datasets

2022-11-15 · Tom Osika, Ebrahim Ebrahim, Martin Styner, Marc Niethammer 외

A major data pre-processing step for large, multi-site studies is to handle site effects by harmonizing data, generating a dataset that enables more powerful analyses and more robust algorithms. There is a wide variety o…

BenchmarkingDiffusion MRI

Concordance Comparison as a Means of Assembling Local Grammars

2026-05-12 · Juliana Pirovani, Elias de Oliveira, Eric Laporte arxiv

Named Entity Recognition for person names is an important but non-trivial task in information extraction. This article uses a tool that compares the concordances obtained from two local grammars (LG) and highlights the d…

Information Extraction

Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination

2025-03-06 · Simin Chen, Pranav Pusarla, Baishakhi Ray

The rapid evolution of code largelanguage models underscores the need for effective and transparent benchmarking of their reasoning capabilities. However, the current benchmarking approach heavily depends on publicly ava…

Benchmarking