A Library Perspective on Supervised Text Processing in Digital Libraries: An Investigation in the Biomedical Domain
Digital libraries that maintain extensive textual collections may want to further enrich their content for certain downstream applications, e.g., building knowledge graphs, semantic enrichment of documents, or implementing novel access paths. All of these applications require some text processing, either to identify relevant entities, extract semantic relationships between them, or to classify documents into some categories. However, implementing reliable, supervised workflows can become quite challenging for a digital library because suitable training data must be crafted, and reliable models must be trained. While many works focus on achieving the highest accuracy on some benchmarks, we tackle the problem from a digital library practitioner. In other words, we also consider trade-offs between accuracy and application costs, dive into training data generation through distant supervision and large language models such as ChatGPT, LLama, and Olmo, and discuss how to design final pipelines. Therefore, we focus on relation extraction and text classification, using the showcase of eight biomedical benchmarks.
Code (1)
Tasks
Knowledge GraphsRelation Extractiontext-classificationText ClassificationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
A Library Perspective on Nearly-Unsupervised Information Extraction Workflows in Digital Libraries
Information extraction can support novel and effective access paths for digital libraries. Nevertheless, designing reliable extraction workflows can be cost-intensive in practice. On the one hand, suitable extraction met…
digHolo : High-speed library for off-axis digital holography and Hermite-Gaussian decomposition
'digHolo' is a numerical library for processing batches of input off-axis digital holography interferograms and outputting the corresponding reconstructed fields. Optionally the library can perform a modal decomposition …
Bibliometrics, Information Retrieval and Natural Language Processing: Natural Synergies to Support Digital Library Research
Document classification methods
Information on different fields which are collected by users requires appropriate management and organization to be structured in a standard way and retrieved fast and more easily. Document classification is a convention…
ClassificationDocument ClassificationGeneral ClassificationManagementA systematic review of Hate Speech automatic detection using Natural Language Processing
With the multiplication of social media platforms, which offer anonymity, easy access and online community formation, and online debate, the issue of hate speech detection and tracking becomes a growing challenge to soci…
Deep LearningHate Speech Detection