Papers Extreme Summarization
“Extreme Summarization” 태그가 달린 논문 24편 · 필터 해제
Bridging the Data Gap: Creating a Hindi Text Summarization Dataset from the English XSUM
Current advancements in Natural Language Processing (NLP) have largely favored resource-rich languages, leaving a significant gap in high-quality datasets for low-resource languages like Hindi. This scarcity is particula…
Extreme SummarizationText SummarizationBOE-XSUM: Extreme Summarization in Clear Language of Spanish Legal Decrees and Notifications
The ability to summarize long documents succinctly is increasingly important in daily life due to information overload, yet there is a notable lack of such summaries for Spanish documents in general, and in the legal dom…
Extreme SummarizationDiffSampling: Enhancing Diversity and Accuracy in Neural Text Generation
Despite their growing capabilities, language models still frequently reproduce content from their training data, generate repetitive text, and favor common grammatical patterns and vocabulary. A possible cause is the dec…
DiversityExtreme SummarizationMathText GenerationAPEX$^2$: Adaptive and Extreme Summarization for Personalized Knowledge Graphs
Knowledge graphs (KGs), which store an extensive number of relational facts, serve various applications. Recently, personalized knowledge graphs (PKGs) have emerged as a solution to optimize storage costs by customizing …
Extreme SummarizationKnowledge GraphsExplainable News Summarization -- Analysis and mitigation of Disagreement Problem
Explainable AI (XAI) techniques for text summarization provide valuable understanding of how the summaries are generated. Recent studies have highlighted a major challenge in this area, known as the disagreement problem.…
Extreme SummarizationNews SummarizationText SummarizationROUGE-K: Do Your Summaries Have Keywords?
Keywords, that is, content-relevant words in summaries play an important role in efficient information conveyance, making it critical to assess if system-generated summaries contain such informative words during evaluati…
Extreme SummarizationImproving Primary Healthcare Workflow Using Extreme Summarization of Scientific Literature Based on Generative AI
Primary care professionals struggle to keep up to date with the latest scientific literature critical in guiding evidence-based practice related to their daily work. To help solve the above-mentioned problem, we employed…
Extreme SummarizationCurriculum-guided Abstractive Summarization for Mental Health Online Posts
Automatically generating short summaries from users' online mental health posts could save counselors' reading time and reduce their fatigue so that they can provide timely responses to those seeking help for improving t…
Abstractive Text SummarizationExtreme SummarizationSentenceCurriculum-Guided Abstractive Summarization
Recent Transformer-based summarization models have provided a promising approach to abstractive summarization. They go beyond sentence selection and extractive strategies to deal with more complicated tasks such as novel…
Abstractive Text SummarizationDecoderExtreme SummarizationInformativeness+1WikiDes: A Wikipedia-Based Dataset for Generating Short Descriptions from Paragraphs
As free online encyclopedias with massive volumes of content, Wikipedia and Wikidata are key to many Natural Language Processing (NLP) tasks, such as information retrieval, knowledge base building, machine translation, t…
ArticlesContrastive LearningExtreme SummarizationRetrieval+4X-SCITLDR: Cross-Lingual Extreme Summarization of Scholarly Documents
The number of scientific publications nowadays is rapidly increasing, causing information overload for researchers and making it hard for scholars to keep up to date with current trends and lines of work. Consequently, r…
Extreme SummarizationMachine TranslationText SummarizationCiteSum: Citation Text-guided Scientific Extreme Summarization and Domain Adaptation with Limited Supervision
Scientific extreme summarization (TLDR) aims to form ultra-short summaries of scientific papers. Previous efforts on curating scientific TLDR datasets failed to scale up due to the heavy human annotation and domain exper…
Domain AdaptationExtreme SummarizationHeadline GenerationIndicBART: A Pre-trained Model for Indic Natural Language Generation
We study pre-trained sequence-to-sequence model for a specific-language family with a focus on Indic languages. We present IndicBART, a multilingual, sequence-to-sequence pre-trained model focusing on 11 Indic languages…
Extreme SummarizationMachine TranslationNMTText Generation+2TLDR9+: A Large Scale Resource for Extreme Summarization of Social Media Posts
Recent models in developing summarization systems consist of millions of parameters and the model performance is highly dependent on the abundance of training data. While most existing summarization corpora contain data …
Extreme SummarizationSentenceIndicBART: A Pre-trained Model for Indic Natural Language Generation
In this paper, we study pre-trained sequence-to-sequence models for a group of related languages, with a focus on Indic languages. We present IndicBART, a multilingual, sequence-to-sequence pre-trained model focusing on …
Extreme SummarizationMachine TranslationNMTText Generation+2ByT5: Towards a token-free future with pre-trained byte-to-byte models
Most widely-used pre-trained language models operate on sequences of tokens corresponding to word or subword units. By comparison, token-free models that operate directly on raw text (bytes or characters) have many benef…
Cross-Lingual Natural Language InferenceCross-Lingual NERCross-Lingual Paraphrase IdentificationCross-Lingual Question Answering+2Focus Attention: Promoting Faithfulness and Diversity in Summarization
Professional summaries are written with document-level information, such as the theme of the document, in mind. This is in contrast with most seq2seq decoders which simultaneously learn to focus on salient content, while…
DiversityExtreme SummarizationThe GEM Benchmark: Natural Language Generation, its Evaluation and Metrics
We introduce GEM, a living benchmark for natural language Generation (NLG), its Evaluation, and Metrics. Measuring progress in NLG relies on a constantly evolving ecosystem of automated metrics, datasets, and human evalu…
Abstractive Text SummarizationCross-Lingual Abstractive SummarizationData-to-Text GenerationExtreme Summarization+4Multi-XScience: A Large-scale Dataset for Extreme Multi-document Summarization of Scientific Articles
Multi-document summarization is a challenging task for which there exists little large-scale datasets. We propose Multi-XScience, a large-scale multi-document summarization dataset created from scientific articles. Multi…
ArticlesDescriptiveDocument SummarizationExtreme Summarization+1TLDR: Extreme Summarization of Scientific Documents
We introduce TLDR generation, a new form of extreme summarization, for scientific papers. TLDR generation involves high source compression and requires expert background knowledge and understanding of complex domain-spec…
Abstractive Text SummarizationExtreme Summarization