paper-with-me

Papers Extreme Summarization

“Extreme Summarization” 태그가 달린 논문 24편 · 필터 해제

Bridging the Data Gap: Creating a Hindi Text Summarization Dataset from the English XSUM

2026-01-04 · Praveenkumar Katwe, RakeshChandra Balabantaray, Kaliprasad Vittala arxiv

Current advancements in Natural Language Processing (NLP) have largely favored resource-rich languages, leaving a significant gap in high-quality datasets for low-resource languages like Hindi. This scarcity is particula…

Extreme SummarizationText Summarization

BOE-XSUM: Extreme Summarization in Clear Language of Spanish Legal Decrees and Notifications

2025-09-29 · Andrés Fernández García, Javier de la Rosa, Julio Gonzalo, Roser Morante 외 arxiv

The ability to summarize long documents succinctly is increasingly important in daily life due to information overload, yet there is a notable lack of such summaries for Spanish documents in general, and in the legal dom…

Extreme Summarization

DiffSampling: Enhancing Diversity and Accuracy in Neural Text Generation

2025-02-19 · Giorgio Franceschelli, Mirco Musolesi

Despite their growing capabilities, language models still frequently reproduce content from their training data, generate repetitive text, and favor common grammatical patterns and vocabulary. A possible cause is the dec…

DiversityExtreme SummarizationMathText Generation

APEX$^2$: Adaptive and Extreme Summarization for Personalized Knowledge Graphs

2024-12-23 · Zihao Li, Dongqi Fu, Mengting Ai, Jingrui He

Knowledge graphs (KGs), which store an extensive number of relational facts, serve various applications. Recently, personalized knowledge graphs (PKGs) have emerged as a solution to optimize storage costs by customizing …

Extreme SummarizationKnowledge Graphs

Explainable News Summarization -- Analysis and mitigation of Disagreement Problem

2024-10-24 · Seema Aswani, Sujala D. Shetty

Explainable AI (XAI) techniques for text summarization provide valuable understanding of how the summaries are generated. Recent studies have highlighted a major challenge in this area, known as the disagreement problem.…

Extreme SummarizationNews SummarizationText Summarization

ROUGE-K: Do Your Summaries Have Keywords?

2024-03-08 · Sotaro Takeshita, Simone Paolo Ponzetto, Kai Eckert

Keywords, that is, content-relevant words in summaries play an important role in efficient information conveyance, making it critical to assess if system-generated summaries contain such informative words during evaluati…

Extreme Summarization

Improving Primary Healthcare Workflow Using Extreme Summarization of Scientific Literature Based on Generative AI

2023-07-24 · Gregor Stiglic, Leon Kopitar, Lucija Gosak, Primoz Kocbek 외

Primary care professionals struggle to keep up to date with the latest scientific literature critical in guiding evidence-based practice related to their daily work. To help solve the above-mentioned problem, we employed…

Extreme Summarization

Curriculum-guided Abstractive Summarization for Mental Health Online Posts

2023-02-02 · Sajad Sotudeh, Nazli Goharian, Hanieh Deilamsalehy, Franck Dernoncourt

Automatically generating short summaries from users' online mental health posts could save counselors' reading time and reduce their fatigue so that they can provide timely responses to those seeking help for improving t…

Abstractive Text SummarizationExtreme SummarizationSentence

Curriculum-Guided Abstractive Summarization

2023-02-02 · Sajad Sotudeh, Hanieh Deilamsalehy, Franck Dernoncourt, Nazli Goharian

Recent Transformer-based summarization models have provided a promising approach to abstractive summarization. They go beyond sentence selection and extractive strategies to deal with more complicated tasks such as novel…

Abstractive Text SummarizationDecoderExtreme SummarizationInformativeness+1

WikiDes: A Wikipedia-Based Dataset for Generating Short Descriptions from Paragraphs

2022-09-27 · Hoang Thang Ta, Abu Bakar Siddiqur Rahman, Navonil Majumder, Amir Hussain 외

As free online encyclopedias with massive volumes of content, Wikipedia and Wikidata are key to many Natural Language Processing (NLP) tasks, such as information retrieval, knowledge base building, machine translation, t…

ArticlesContrastive LearningExtreme SummarizationRetrieval+4

X-SCITLDR: Cross-Lingual Extreme Summarization of Scholarly Documents

2022-05-30 · Sotaro Takeshita, Tommaso Green, Niklas Friedrich, Kai Eckert 외

The number of scientific publications nowadays is rapidly increasing, causing information overload for researchers and making it hard for scholars to keep up to date with current trends and lines of work. Consequently, r…

Extreme SummarizationMachine TranslationText Summarization

CiteSum: Citation Text-guided Scientific Extreme Summarization and Domain Adaptation with Limited Supervision

2022-05-12 · Yuning Mao, Ming Zhong, Jiawei Han

Scientific extreme summarization (TLDR) aims to form ultra-short summaries of scientific papers. Previous efforts on curating scientific TLDR datasets failed to scale up due to the heavy human annotation and domain exper…

Domain AdaptationExtreme SummarizationHeadline Generation

IndicBART: A Pre-trained Model for Indic Natural Language Generation

2021-11-16 · ACL ARR November 2021 11 · Anonymous

We study pre-trained sequence-to-sequence model for a specific-language family with a focus on Indic languages. We present IndicBART, a multilingual, sequence-to-sequence pre-trained model focusing on 11 Indic languages…

Extreme SummarizationMachine TranslationNMTText Generation+2

TLDR9+: A Large Scale Resource for Extreme Summarization of Social Media Posts

2021-10-04 · EMNLP (newsum) 2021 11 · Sajad Sotudeh, Hanieh Deilamsalehy, Franck Dernoncourt, Nazli Goharian

Recent models in developing summarization systems consist of millions of parameters and the model performance is highly dependent on the abundance of training data. While most existing summarization corpora contain data …

Extreme SummarizationSentence

IndicBART: A Pre-trained Model for Indic Natural Language Generation

2021-09-07 · Findings (ACL) 2022 5 · Raj Dabre, Himani Shrotriya, Anoop Kunchukuttan, Ratish Puduppully 외

In this paper, we study pre-trained sequence-to-sequence models for a group of related languages, with a focus on Indic languages. We present IndicBART, a multilingual, sequence-to-sequence pre-trained model focusing on …

Extreme SummarizationMachine TranslationNMTText Generation+2

ByT5: Towards a token-free future with pre-trained byte-to-byte models

2021-05-28 · Linting Xue, Aditya Barua, Noah Constant, Rami Al-Rfou 외

Most widely-used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benef…

Cross-Lingual Natural Language InferenceCross-Lingual NERCross-Lingual Paraphrase IdentificationCross-Lingual Question Answering+2

Focus Attention: Promoting Faithfulness and Diversity in Summarization

2021-05-25 · ACL 2021 5 · Rahul Aralikatte, Shashi Narayan, Joshua Maynez, Sascha Rothe 외

Professional summaries are written with document-level information, such as the theme of the document, in mind. This is in contrast with most seq2seq decoders which simultaneously learn to focus on salient content, while…

DiversityExtreme Summarization

The GEM Benchmark: Natural Language Generation, its Evaluation and Metrics

2021-02-02 · ACL (GEM) 2021 8 · Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi 외

We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics. Measuring progress in NLG relies on a constantly evolving ecosystem of automated metrics, datasets, and human evalu…

Abstractive Text SummarizationCross-Lingual Abstractive SummarizationData-to-Text GenerationExtreme Summarization+4

Multi-XScience: A Large-scale Dataset for Extreme Multi-document Summarization of Scientific Articles

2020-10-27 · EMNLP 2020 11 · Yao Lu, Yue Dong, Laurent Charlin

Multi-document summarization is a challenging task for which there exists little large-scale datasets. We propose Multi-XScience, a large-scale multi-document summarization dataset created from scientific articles. Multi…

ArticlesDescriptiveDocument SummarizationExtreme Summarization+1

TLDR: Extreme Summarization of Scientific Documents

2020-04-30 · Findings of the Association for Computational Linguistics 2020 · Isabel Cachola, Kyle Lo, Arman Cohan, Daniel S. Weld

We introduce TLDR generation, a new form of extreme summarization, for scientific papers. TLDR generation involves high source compression and requires expert background knowledge and understanding of complex domain-spec…

Abstractive Text SummarizationExtreme Summarization
1–20 / 24 다음 →