paper-with-me

홈 › Papers

Pushing The Limit of LLM Capacity for Text Classification

2024-02-12 · Yazhou Zhang, Mengyao Wang, Chenyu Ren, Qiuchi Li, Prayag Tiwari, Benyou Wang, Jing Qin

The value of text classification's future research has encountered challenges and uncertainties, due to the extraordinary efficacy demonstrated by large language models (LLMs) across numerous downstream NLP tasks. In this era of open-ended language modeling, where task boundaries are gradually fading, an urgent question emerges: have we made significant advances in text classification under the full benefit of LLMs? To answer this question, we propose RGPT, an adaptive boosting framework tailored to produce a specialized text classification LLM by recurrently ensembling a pool of strong base learners. The base learners are constructed by adaptively adjusting the distribution of training samples and iteratively fine-tuning LLMs with them. Such base learners are then ensembled to be a specialized text classification LLM, by recurrently incorporating the historical predictions from the previous learners. Through a comprehensive empirical comparison, we show that RGPT significantly outperforms 8 SOTA PLMs and 7 SOTA LLMs on four benchmarks by 1.36% on average. Further evaluation experiments show a clear surpassing of RGPT over human classification.

📄 PDF Abstract BibTeX arXiv:2402.07470

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationLanguage ModelingLanguage Modellingtext-classificationText Classification

Methods 이 논문이 사용한 방법론

BASE 설명 없음

Similar Papers 제목 키워드 기반

A baseline revisited: Pushing the limits of multi-segment models for context-aware translation

2022-10-19 · Suvodeep Majumder, Stanislas Lauly, Maria Nadejde, Marcello Federico 외

This paper addresses the task of contextual translation using multi-segment models. Specifically we show that increasing model capacity further pushes the limits of this approach and that deeper models are more suited to…

Knowledge DistillationTranslation

AraMUS: Pushing the Limits of Data and Model Scale for Arabic Natural Language Processing

2023-06-11 · Asaad Alghamdi, Xinyu Duan, Wei Jiang, Zhenhai Wang 외

Developing monolingual large Pre-trained Language Models (PLMs) is shown to be very successful in handling different tasks in Natural Language Processing (NLP). In this work, we present AraMUS, the largest Arabic PLM wit…

Few-Shot Learning

Knowledge Distillation Layer that Lets the Student Decide

2023-09-06 · Ada Gorgun, Yeti Z. Gurbuz, A. Aydin Alatan

Typical technique in knowledge distillation (KD) is regularizing the learning of a limited capacity model (student) by pushing its responses to match a powerful model's (teacher). Albeit useful especially in the penultim…

Knowledge Distillation

Pushing the boundaries of audiovisual word recognition using Residual Networks and LSTMs

2018-11-03 · Themos Stafylakis, Muhammad Haris Khan, Georgios Tzimiropoulos

Visual and audiovisual speech recognition are witnessing a renaissance which is largely due to the advent of deep learning methods. In this paper, we present a deep learning architecture for lipreading and audiovisual wo…

Lipreadingspeech-recognitionSpeech Recognition

Visual Prompt Guided Unified Pushing Policy

2026-02-22 · Hieu Bui, Ziyan Gao, Yuya Hosoda, Joo-Ho Lee arxiv

As one of the simplest non-prehensile manipulation skills, pushing has been widely studied as an effective means to rearrange objects. Existing approaches, however, typically rely on multi-step push plans composed of pre…