Papers Topic Classification
“Topic Classification” 태그가 달린 논문 186편 · 필터 해제
On Multilingual Encoder Language Model Compression for Low-Resource Languages
In this paper, we combine two-step knowledge distillation, structured pruning, truncation, and vocabulary trimming for extremely compressing multilingual encoder-only language models for low-resource languages. Our novel…
Knowledge DistillationLanguage ModelingLanguage ModellingModel Compression+5A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
Content moderation research has recently made significant advances, but still fails to serve the majority of the world's languages due to the lack of resources, leaving millions of vulnerable users to online hostility. T…
Abusive LanguageTopic ClassificationLow-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline
This paper presents an end-to-end suite for multilingual information extraction and processing from image-based documents. The system uses Optical Character Recognition (Tesseract) to extract text in languages such as En…
Abstractive Text SummarizationLanguage ModelingLanguage ModellingLarge Language Model+5A thorough benchmark of automatic text classification: From traditional approaches to large language models
Automatic text classification (ATC) has experienced remarkable advancements in the past decade, best exemplified by recent small and large language models (SLMs and LLMs), leveraged by Transformer architectures. Despite …
Sentiment Analysistext-classificationText ClassificationTopic ClassificationDetection of Somali-written Fake News and Toxic Messages on the Social Media Using Transformer-based Language Models
The fact that everyone with a social media account can create and share content, and the increasing public reliance on social media platforms as a news and information source bring about significant challenges such as mi…
Language ModelingLanguage ModellingMisinformationTopic ClassificationA Statistical Theory of Contrastive Learning via Approximate Sufficient Statistics
Contrastive learning -- a modern approach to extract useful representations from unlabeled data by training models to distinguish similar samples from dissimilar ones -- has driven significant progress in foundation mode…
Contrastive LearningData AugmentationregressionTopic ClassificationReading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models
Oscar Wilde said, "The difference between literature and journalism is that journalism is unreadable, and literature is not read." Unfortunately, The digitally archived journalism of Oscar Wilde's 19th century often has …
Image to textOptical Character RecognitionOptical Character Recognition (OCR)Topic ClassificationConcept Navigation and Classification via Open-Source Large Language Model Processing
This paper presents a novel methodological framework for detecting and classifying latent constructs, including frames, narratives, and topics, from textual data using Open-Source Large Language Models (LLMs). The propos…
ArticlesLanguage ModelingLanguage ModellingLarge Language Model+1Analyzing the Effect of Linguistic Similarity on Cross-Lingual Transfer: Tasks and Experimental Setups Matter
Cross-lingual transfer is a popular approach to increase the amount of training data for NLP tasks in a low-resource context. However, the best strategy to decide which cross-lingual data to include is unclear. Prior res…
Cross-Lingual TransferDependency ParsingPOSPOS Tagging+1DISHONEST: Dissecting misInformation Spread using Homogeneous sOcial NEtworks and Semantic Topic classification
The emergence of the COVID-19 pandemic resulted in a significant rise in the spread of misinformation on online platforms such as Twitter. Oftentimes this growth is blamed on the idea of the "echo chamber." However, the …
MisinformationTopic ClassificationEvaluating Pixel Language Models on Non-Standardized Languages
We explore the potential of pixel-based models for transfer learning from standard languages to dialects. These models convert text into images that are divided into patches, enabling a continuous vocabulary representati…
Dependency ParsingIntent DetectionPart-Of-Speech TaggingTopic Classification+1LLM Teacher-Student Framework for Text Classification With No Manually Annotated Data: A Case Study in IPTC News Topic Classification
With the ever-increasing number of news stories available online, classifying them by topic, regardless of the language they are written in, has become crucial for enhancing readers' access to relevant content. To addres…
ArticlesClassificationNews Classificationtext-classification+2QuickCharNet: An Efficient URL Classification Framework for Enhanced Search Engine Optimization
The uniform resource locator (URL) conveys essential information about a page’s topic, authority, and security, which significantly influences its ranking in search engine results. However, many existing URL classificati…
ClassificationEfficient Neural NetworkMarketingSpam detection+1From Measurement Instruments to Data: Leveraging Theory-Driven Synthetic Training Data for Classifying Social Constructs
Computational text classification is a challenging task, especially for multi-dimensional social constructs. Recently, there has been increasing discussion that synthetic training data could enhance classification by off…
Classificationtext-classificationText ClassificationTopic ClassificationInference and Verbalization Functions During In-Context Learning
Large language models (LMs) are capable of in-context learning from a few demonstrations (example-label pairs) to solve new tasks during inference. Despite the intuitive importance of high-quality demonstrations, previou…
In-Context LearningNatural Language InferenceSentiment AnalysisTopic ClassificationThe Large Language Model GreekLegalRoBERTa
We develop four versions of GreekLegalRoBERTa, which are four large language models trained on Greek legal and nonlegal text. We show that our models surpass the performance of GreekLegalBERT, Greek- LegalBERT-v2, and Gr…
Language ModelingLanguage ModellingLarge Language Modelmodel+3Language Model-Driven Data Pruning Enables Efficient Active Learning
Active learning (AL) optimizes data labeling efficiency by selecting the most informative instances for annotation. A key component in this procedure is an acquisition function that guides the selection process and ident…
Active LearningLanguage ModelingLanguage ModellingSentiment Analysis+1Multilingual Topic Classification in X: Dataset and Analysis
In the dynamic realm of social media, diverse topics are discussed daily, transcending linguistic boundaries. However, the complexities of understanding and categorising this content across various languages remain an im…
ClassificationDiversityTopic ClassificationGrEmLIn: A Repository of Green Baseline Embeddings for 87 Low-Resource Languages Injected with Multilingual Graph Knowledge
Contextualized embeddings based on large language models (LLMs) are available for various languages, but their coverage is often limited for lower resourced languages. Using LLMs for such languages is often difficult due…
Natural Language InferenceSentiment AnalysisTopic ClassificationWord Embeddings+1Optimal and efficient text counterfactuals using Graph Neural Networks
As NLP models become increasingly integral to decision-making processes, the need for explainability and interpretability has become paramount. In this work, we propose a framework that achieves the aforementioned by gen…
counterfactualDecision MakingSentiment AnalysisSentiment Classification+1