paper-with-me

홈 › Papers

Annif at SemEval-2025 Task 5: Traditional XMTC augmented by LLMs

2025-04-28 · Osma Suominen, Juho Inkinen, Mona Lehtinen

This paper presents the Annif system in SemEval-2025 Task 5 (LLMs4Subjects), which focussed on subject indexing using large language models (LLMs). The task required creating subject predictions for bibliographic records from the bilingual TIBKAT database using the GND subject vocabulary. Our approach combines traditional natural language processing and machine learning techniques implemented in the Annif toolkit with innovative LLM-based methods for translation and synthetic data generation, and merging predictions from monolingual models. The system ranked first in the all-subjects category and second in the tib-core-subjects category in the quantitative evaluation, and fourth in qualitative evaluations. These findings demonstrate the potential of combining traditional XMTC algorithms with modern LLM techniques to improve the accuracy and efficiency of subject indexing in multilingual contexts.

📄 PDF Abstract BibTeX arXiv:2504.19675

Code (2)

NatLibFi/Annif 공식 구현
natlibfi/annif-llms4subjects 공식 구현

Tasks

Synthetic Data Generation

Similar Papers 제목 키워드 기반

Annif at the GermEval-2025 LLMs4Subjects Task: Traditional XMTC Augmented by Efficient LLMs

2025-08-21 · Osma Suominen, Juho Inkinen, Mona Lehtinen arxiv

This paper presents the Annif system in the LLMs4Subjects shared task (Subtask 2) at GermEval-2025. The task required creating subject predictions for bibliographic records using large language models, with a special foc…

Synthetic Data GenerationComputational Efficiency

Correlation Networks for Extreme Multi-label Text Classification

2022-08-23 · Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining 2022 8 · Guangxu Xun, Kishlay Jha, Jianhui Sun, Aidong Zhang

This paper develops the Correlation Networks (CorNet) architecture for the extreme multi-label text classification (XMTC) task, where the objective is to tag an input text sequence with the most relevant subset of labels…

ClassificationMulti Label Text ClassificationMulti-Label Text ClassificationTAG+2

Adversarial Examples for Extreme Multilabel Text Classification

2021-12-14 · Mohammadreza Qaraei, Rohit Babbar

Extreme Multilabel Text Classification (XMTC) is a text classification problem in which, (i) the output space is extremely large, (ii) each data point may have multiple positive labels, and (iii) the data follows a stron…

ClassificationMultilabel Text ClassificationRecommendation Systemstext-classification+1

Extreme Multi-label Text Classification with Multi-layer Experts

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Extreme multi-label text classification (XMTC) is the task of tagging each document with the relevant labels from a very large space of predefined categories, which presents an open challenge in the recent development of…

ClassificationMulti Label Text ClassificationMulti-Label Text Classificationtext-classification+1

AttentionXML: Label Tree-based Attention-Aware Deep Model for High-Performance Extreme Multi-Label Text Classification

2018-11-01 · NeurIPS 2019 12 · Ronghui You, Zihan Zhang, Ziye Wang, Suyang Dai 외

Extreme multi-label text classification (XMTC) is an important problem in the era of big data, for tagging a given text with the most relevant multiple labels from an extremely large-scale label set. XMTC can be found in…

General ClassificationMulti-Label Text ClassificationNews AnnotationProduct Categorization+2