Annif at SemEval-2025 Task 5: Traditional XMTC augmented by LLMs
This paper presents the Annif system in SemEval-2025 Task 5 (LLMs4Subjects), which focussed on subject indexing using large language models (LLMs). The task required creating subject predictions for bibliographic records from the bilingual TIBKAT database using the GND subject vocabulary. Our approach combines traditional natural language processing and machine learning techniques implemented in the Annif toolkit with innovative LLM-based methods for translation and synthetic data generation, and merging predictions from monolingual models. The system ranked first in the all-subjects category and second in the tib-core-subjects category in the quantitative evaluation, and fourth in qualitative evaluations. These findings demonstrate the potential of combining traditional XMTC algorithms with modern LLM techniques to improve the accuracy and efficiency of subject indexing in multilingual contexts.
Code (2)
Tasks
Synthetic Data GenerationSimilar Papers 제목 키워드 기반
Annif at the GermEval-2025 LLMs4Subjects Task: Traditional XMTC Augmented by Efficient LLMs
This paper presents the Annif system in the LLMs4Subjects shared task (Subtask 2) at GermEval-2025. The task required creating subject predictions for bibliographic records using large language models, with a special foc…
Synthetic Data GenerationComputational EfficiencyCorrelation Networks for Extreme Multi-label Text Classification
This paper develops the Correlation Networks (CorNet) architecture for the extreme multi-label text classification (XMTC) task, where the objective is to tag an input text sequence with the most relevant subset of labels…
ClassificationMulti Label Text ClassificationMulti-Label Text ClassificationTAG+2Adversarial Examples for Extreme Multilabel Text Classification
Extreme Multilabel Text Classification (XMTC) is a text classification problem in which, (i) the output space is extremely large, (ii) each data point may have multiple positive labels, and (iii) the data follows a stron…
ClassificationMultilabel Text ClassificationRecommendation Systemstext-classification+1Extreme Multi-label Text Classification with Multi-layer Experts
Extreme multi-label text classification (XMTC) is the task of tagging each document with the relevant labels from a very large space of predefined categories, which presents an open challenge in the recent development of…
ClassificationMulti Label Text ClassificationMulti-Label Text Classificationtext-classification+1AttentionXML: Label Tree-based Attention-Aware Deep Model for High-Performance Extreme Multi-Label Text Classification
Extreme multi-label text classification (XMTC) is an important problem in the era of big data, for tagging a given text with the most relevant multiple labels from an extremely large-scale label set. XMTC can be found in…
General ClassificationMulti-Label Text ClassificationNews AnnotationProduct Categorization+2