paper-with-me

Papers Topic Classification

“Topic Classification” 태그가 달린 논문 186편 · 필터 해제

Assessing In-context Learning and Fine-tuning for Topic Classification of German Web Data

2024-07-23 · Julian Schelb, Roberto Ulloa, Andreas Spitz

Researchers in the political and social sciences often rely on classification models to analyze trends in information consumption by examining browsing histories of millions of webpages. Automated scalable methods are ne…

Binary ClassificationIn-Context LearningTopic Classification

Automatic Classification of News Subjects in Broadcast News: Application to a Gender Bias Representation Analysis

2024-07-19 · Valentin Pelloin, Lena Dodson, Émile Chapuis, Nicolas Hervé 외

This paper introduces a computational framework designed to delineate gender distribution biases in topics covered by French TV and radio news. We transcribe a dataset of 11.7k hours, broadcasted in 2023 on 21 French cha…

Language ModelingLanguage ModellingLarge Language ModelTopic Classification

Multi-task Prompt Words Learning for Social Media Content Generation

2024-07-10 · Haochen Xue, Chong Zhang, Chengzhi Liu, Fangyu Wu 외

The rapid development of the Internet has profoundly changed human life. Humans are increasingly expressing themselves and interacting with others on social media platforms. However, although artificial intelligence tech…

Keyword ExtractionScene RecognitionSentiment AnalysisTopic Classification

STF: Sentence Transformer Fine-Tuning For Topic Categorization With Limited Data

2024-07-03 · Kheir Eddine Daouadi, Yaakoub Boualleg, Oussama Guehairia

Nowadays, topic classification from tweets attracts considerable research attention. Different classification systems have been suggested thanks to these research efforts. Nevertheless, they face major challenges owing t…

ClassificationSentenceTopic Classification

Retrieval Augmented Zero-Shot Text Classification

2024-06-21 · Tassallah Abdullahi, Ritambhara Singh, Carsten Eickhoff

Zero-shot text learning enables text classifiers to handle unseen classes efficiently, alleviating the need for task-specific training data. A simple approach often relies on comparing embeddings of query (text) to those…

ClassificationRetrievaltext-classificationText Classification+2

Newswire: A Large-Scale Structured Database of a Century of Historical News

2024-06-13 · Emily Silcock, Abhishek Arora, Luca D'Amico-Wong, Melissa Dell

In the U.S. historically, local newspapers drew their content largely from newswires like the Associated Press. Historians argue that newswires played a pivotal role in creating a national identity and shared understandi…

ArticlesEntity DisambiguationLanguage ModelingLanguage Modelling+1

Topic Classification of Case Law Using a Large Language Model and a New Taxonomy for UK Law: AI Insights into Summary Judgment

2024-05-21 · Holli Sargeant, Ahmed Izzidien, Felix Steffek

This paper addresses a critical gap in legal analytics by developing and applying a novel taxonomy for topic classification of summary judgment cases in the United Kingdom. Using a curated dataset of summary judgment cas…

Language ModelingLanguage ModellingLarge Language ModelTopic Classification

SynthesizRR: Generating Diverse Datasets with Retrieval Augmentation

2024-05-16 · Abhishek Divekar, Greg Durrett

It is often desirable to distill the capabilities of large language models (LLMs) into smaller student models due to compute and memory constraints. One way to do this for classification tasks is via dataset synthesis, w…

Bias DetectionDiversityHumor DetectionInformation Retrieval+6

InsightNet: Structured Insight Mining from Customer Feedback

2024-05-12 · Sandeep Sricharan Mukku, Manan Soni, Jitenkumar Rana, Chetan Aggarwal 외

We propose InsightNet, a novel approach for the automated extraction of structured insights from customer reviews. Our end-to-end machine learning framework is designed to overcome the limitations of current solutions, i…

Semantic SimilaritySemantic Textual SimilarityTopic Classification

Addressing Topic Granularity and Hallucination in Large Language Models for Topic Modelling

2024-05-01 · Yida Mu, Peizhen Bai, Kalina Bontcheva, Xingyi Song

Large language models (LLMs) with their strong zero-shot topic extraction capabilities offer an alternative to probabilistic topic modelling and closed-set topic classification approaches. As zero-shot topic extractors, …

HallucinationTopic Classification

What Drives Performance in Multilingual Language Models?

2024-04-29 · Sina Bagheri Nezhad, Ameeta Agrawal

This study investigates the factors influencing the performance of multilingual large language models (MLLMs) across diverse languages. We study 6 MLLMs, including masked language models, autoregressive models, and instr…

Cross-Lingual TransferMultilingual NLPTopic ClassificationTransfer Learning

L3Cube-MahaNews: News-based Short Text and Long Document Classification Datasets in Marathi

2024-04-28 · Saloni Mittal, Vidula Magdum, Omkar Dhekane, Sharayu Hiwarkhedkar 외

The availability of text or topic classification datasets in the low-resource Marathi language is limited, typically consisting of fewer than 4 target labels, with some achieving nearly perfect accuracy. In this work, we…

ArticlesDocument Classificationtext-classificationText Classification+1

Forget NLI, Use a Dictionary: Zero-Shot Topic Classification for Low-Resource Languages with Application to Luxembourgish

2024-04-05 · Fred Philippy, Shohreh Haddadan, Siwen Guo

In NLP, zero-shot classification (ZSC) is the task of assigning labels to textual data without any labeled examples for the target classes. A common method for ZSC is to fine-tune a language model on a Natural Language I…

Language ModellingNatural Language InferenceTopic Classificationzero-shot-classification+1

Few-Shot Cross-Lingual Transfer for Prompting Large Language Models in Low-Resource Languages

2024-03-09 · Christopher Toukmaji

Large pre-trained language models (PLMs) are at the forefront of advances in Natural Language Processing. One widespread use case of PLMs is "prompting" - or in-context learning - where a user provides a description of a…

Abstractive Text SummarizationCross-Lingual TransferIn-Context LearningMachine Translation+3

Zero-Shot Topic Classification of Column Headers: Leveraging LLMs for Metadata Enrichment

2024-03-01 · Margherita Martorana, Tobias Kuhn, Lise Stork, Jacco van Ossenbruggen

Traditional dataset retrieval systems rely on metadata for indexing, rather than on the underlying data values. However, high-quality metadata creation and enrichment often require manual annotations, which is a labour-i…

Retrievaltext-classificationText ClassificationTopic Classification

LexC-Gen: Generating Data for Extremely Low-Resource Languages with Large Language Models and Bilingual Lexicons

2024-02-21 · Zheng-Xin Yong, Cristina Menghini, Stephen H. Bach

Data scarcity in low-resource languages can be addressed with word-to-word translations from labeled task data in high-resource languages using bilingual lexicons. However, bilingual lexicons often have limited lexical o…

Sentiment AnalysisTopic ClassificationTranslationWord Translation

Prompt-Based Bias Calibration for Better Zero/Few-Shot Learning of Language Models

2024-02-15 · Kang He, Yinghan Long, Kaushik Roy

Prompt-based learning is susceptible to intrinsic bias present in pre-trained language models (LMs), leading to sub-optimal performance in prompt-based zero/few-shot settings. In this work, we propose a null-input prompt…

FairnessFew-Shot LearningIn-Context LearningLanguage Modeling+3

Advancing NLP Models with Strategic Text Augmentation: A Comprehensive Study of Augmentation Methods and Curriculum Strategies

2024-02-14 · Himmet Toprak Kesgin, Mehmet Fatih Amasyali

This study conducts a thorough evaluation of text augmentation techniques across a variety of datasets and natural language processing (NLP) tasks to address the lack of reliable, generalized evidence for these methods. …

Sentiment AnalysisText AugmentationTopic Classification

L3Cube-IndicNews: News-based Short Text and Long Document Classification Datasets in Indic Languages

2024-01-04 · Aishwarya Mirashi, Srushti Sonavane, Purva Lingayat, Tejas Padhiyar 외

In this work, we introduce L3Cube-IndicNews, a multilingual text classification corpus aimed at curating a high-quality dataset for Indian regional languages, with a specific focus on news headlines and articles. We have…

ArticlesClassificationDocument ClassificationMultilingual text classification+4

Iterative Mask Filling: An Effective Text Augmentation Method Using Masked Language Modeling

2024-01-03 · Himmet Toprak Kesgin, Mehmet Fatih Amasyali

Data augmentation is an effective technique for improving the performance of machine learning models. However, it has not been explored as extensively in natural language processing (NLP) as it has in computer vision. In…

Data Augmentationfill-maskFill MaskLanguage Modeling+5
← 이전 21–40 / 186 다음 →