paper-with-me

Papers Topic Classification

“Topic Classification” 태그가 달린 논문 186편 · 필터 해제

On Multilingual Encoder Language Model Compression for Low-Resource Languages

2025-05-22 · Daniil Gurgurov, Michal Gregor, Josef van Genabith, Simon Ostermann

In this paper, we combine two-step knowledge distillation, structured pruning, truncation, and vocabulary trimming for extremely compressing multilingual encoder-only language models for low-resource languages. Our novel…

Knowledge DistillationLanguage ModelingLanguage ModellingModel Compression+5

A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings

2025-05-17 · Fitsum Gaim, Hoyun Song, Huije Lee, Changgeon Ko 외

Content moderation research has recently made significant advances, but still fails to serve the majority of the world's languages due to the lack of resources, leaving millions of vulnerable users to online hostility. T…

Abusive LanguageTopic Classification

Low-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline

2025-05-16 · Hrishit Madhavi, Jacob Cherian, Yuvraj Khamkar, Dhananjay Bhagat

This paper presents an end-to-end suite for multilingual information extraction and processing from image-based documents. The system uses Optical Character Recognition (Tesseract) to extract text in languages such as En…

Abstractive Text SummarizationLanguage ModelingLanguage ModellingLarge Language Model+5

A thorough benchmark of automatic text classification: From traditional approaches to large language models

2025-04-02 · Washington Cunha, Leonardo Rocha, Marcos André Gonçalves

Automatic text classification (ATC) has experienced remarkable advancements in the past decade, best exemplified by recent small and large language models (SLMs and LLMs), leveraged by Transformer architectures. Despite …

Sentiment Analysistext-classificationText ClassificationTopic Classification

Detection of Somali-written Fake News and Toxic Messages on the Social Media Using Transformer-based Language Models

2025-03-23 · Muhidin A. Mohamed, Shuab D. Ahmed, Yahye A. Isse, Hanad M. Mohamed 외

The fact that everyone with a social media account can create and share content, and the increasing public reliance on social media platforms as a news and information source bring about significant challenges such as mi…

Language ModelingLanguage ModellingMisinformationTopic Classification

A Statistical Theory of Contrastive Learning via Approximate Sufficient Statistics

2025-03-21 · Licong Lin, Song Mei

Contrastive learning -- a modern approach to extract useful representations from unlabeled data by training models to distinguish similar samples from dissimilar ones -- has driven significant progress in foundation mode…

Contrastive LearningData AugmentationregressionTopic Classification

Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models

2025-02-18 · Jonathan Bourne

Oscar Wilde said, "The difference between literature and journalism is that journalism is unreadable, and literature is not read." Unfortunately, The digitally archived journalism of Oscar Wilde's 19th century often has …

Image to textOptical Character RecognitionOptical Character Recognition (OCR)Topic Classification

Concept Navigation and Classification via Open-Source Large Language Model Processing

2025-02-07 · Maël Kubli

This paper presents a novel methodological framework for detecting and classifying latent constructs, including frames, narratives, and topics, from textual data using Open-Source Large Language Models (LLMs). The propos…

ArticlesLanguage ModelingLanguage ModellingLarge Language Model+1

Analyzing the Effect of Linguistic Similarity on Cross-Lingual Transfer: Tasks and Experimental Setups Matter

2025-01-24 · Verena Blaschke, Masha Fedzechkina, Maartje ter Hoeve

Cross-lingual transfer is a popular approach to increase the amount of training data for NLP tasks in a low-resource context. However, the best strategy to decide which cross-lingual data to include is unclear. Prior res…

Cross-Lingual TransferDependency ParsingPOSPOS Tagging+1

DISHONEST: Dissecting misInformation Spread using Homogeneous sOcial NEtworks and Semantic Topic classification

2024-12-12 · Caleb Stam, Emily Saldanha, Mahantesh Halappanavar, Anurag Acharya

The emergence of the COVID-19 pandemic resulted in a significant rise in the spread of misinformation on online platforms such as Twitter. Oftentimes this growth is blamed on the idea of the "echo chamber." However, the …

MisinformationTopic Classification

Evaluating Pixel Language Models on Non-Standardized Languages

2024-12-12 · Alberto Muñoz-Ortiz, Verena Blaschke, Barbara Plank

We explore the potential of pixel-based models for transfer learning from standard languages to dialects. These models convert text into images that are divided into patches, enabling a continuous vocabulary representati…

Dependency ParsingIntent DetectionPart-Of-Speech TaggingTopic Classification+1

LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification

2024-11-29 · Taja Kuzman, Nikola Ljubešić

With the ever-increasing number of news stories available online, classifying them by topic, regardless of the language they are written in, has become crucial for enhancing readers' access to relevant content. To addres…

ArticlesClassificationNews Classificationtext-classification+2

QuickCharNet: An Efficient URL Classification Framework for Enhanced Search Engine Optimization

2024-10-22 · IEEE Access 2024 10 · Fardin Rastakhiz, Mahdi Eftekhari, Sahar Vahdati

The uniform resource locator (URL) conveys essential information about a page’s topic, authority, and security, which significantly influences its ranking in search engine results. However, many existing URL classificati…

ClassificationEfficient Neural NetworkMarketingSpam detection+1

From Measurement Instruments to Data: Leveraging Theory-Driven Synthetic Training Data for Classifying Social Constructs

2024-10-16 · Lukas Birkenmaier, Matthias Roth, Indira Sen

Computational text classification is a challenging task, especially for multi-dimensional social constructs. Recently, there has been increasing discussion that synthetic training data could enhance classification by off…

Classificationtext-classificationText ClassificationTopic Classification

Inference and Verbalization Functions During In-Context Learning

2024-10-12 · Junyi Tao, Xiaoyin Chen, Nelson F. Liu

Large language models (LMs) are capable of in-context learning from a few demonstrations (example-label pairs) to solve new tasks during inference. Despite the intuitive importance of high-quality demonstrations, previou…

In-Context LearningNatural Language InferenceSentiment AnalysisTopic Classification

The Large Language Model GreekLegalRoBERTa

2024-10-10 · Vasileios Saketos, Despina-Athanasia Pantazi, Manolis Koubarakis

We develop four versions of GreekLegalRoBERTa, which are four large language models trained on Greek legal and nonlegal text. We show that our models surpass the performance of GreekLegalBERT, Greek- LegalBERT-v2, and Gr…

Language ModelingLanguage ModellingLarge Language Modelmodel+3

Language Model-Driven Data Pruning Enables Efficient Active Learning

2024-10-05 · Abdul Hameed Azeemi, Ihsan Ayyub Qazi, Agha Ali Raza

Active learning (AL) optimizes data labeling efficiency by selecting the most informative instances for annotation. A key component in this procedure is an acquisition function that guides the selection process and ident…

Active LearningLanguage ModelingLanguage ModellingSentiment Analysis+1

Multilingual Topic Classification in X: Dataset and Analysis

2024-10-04 · Dimosthenis Antypas, Asahi Ushio, Francesco Barbieri, Jose Camacho-Collados

In the dynamic realm of social media, diverse topics are discussed daily, transcending linguistic boundaries. However, the complexities of understanding and categorising this content across various languages remain an im…

ClassificationDiversityTopic Classification

GrEmLIn: A Repository of Green Baseline Embeddings for 87 Low-Resource Languages Injected with Multilingual Graph Knowledge

2024-09-26 · Daniil Gurgurov, Rishu Kumar, Simon Ostermann

Contextualized embeddings based on large language models (LLMs) are available for various languages, but their coverage is often limited for lower resourced languages. Using LLMs for such languages is often difficult due…

Natural Language InferenceSentiment AnalysisTopic ClassificationWord Embeddings+1

Optimal and efficient text counterfactuals using Graph Neural Networks

2024-08-04 · Dimitris Lymperopoulos, Maria Lymperaiou, Giorgos Filandrianos, Giorgos Stamou

As NLP models become increasingly integral to decision-making processes, the need for explainability and interpretability has become paramount. In this work, we propose a framework that achieves the aforementioned by gen…

counterfactualDecision MakingSentiment AnalysisSentiment Classification+1
1–20 / 186 다음 →