Papers Sentence Compression
“Sentence Compression” 태그가 달린 논문 149편 · 필터 해제
Tiny Transformers Excel at Sentence Compression
It is staggering that words of the English language, which are on average represented by 5--6 bytes of ASCII, require as much as 24 kilobytes when served to large language models. We show that there is room for more info…
SentenceSentence CompressionvalidTarget-Aware Language Modeling via Granular Data Sampling
Language model pretraining generally targets a broad range of use cases and incorporates data from diverse sources. However, there are instances where we desire a model that excels in specific areas without markedly comp…
Language ModelingLanguage ModellingSentenceSentence CompressionInstructCMP: Length Control in Sentence Compression through Instruction-based Large Language Models
Extractive summarization can produce faithful summaries but often requires additional constraints such as a desired summary length. Traditional sentence compression models do not typically consider the constraints becaus…
Extractive SummarizationSentenceSentence CompressionAutoBreach: Universal and Adaptive Jailbreaking with Efficient Wordplay-Guided Optimization
Despite the widespread application of large language models (LLMs) across various tasks, recent studies indicate that they are susceptible to jailbreak attacks, which can render their defense mechanisms ineffective. Howe…
SentenceSentence CompressionFrom Lengthy to Lucid: A Systematic Literature Review on NLP Techniques for Taming Long Sentences
Long sentences have been a persistent issue in written communication for many years since they make it challenging for readers to grasp the main points or follow the initial intention of the writer. This survey, conducte…
SentenceSentence CompressionSurveySystematic Literature ReviewReconstruct Before Summarize: An Efficient Two-Step Framework for Condensing and Summarizing Meeting Transcripts
Meetings typically involve multiple participants and lengthy conversations, resulting in redundant and trivial content. To overcome these challenges, we propose a two-step framework, Reconstruct before Summarize (RbS), f…
Language ModellingMeeting SummarizationSentence CompressionImproving Factual Consistency in Summarization with Compression-Based Post-Editing
State-of-the-art summarization models still struggle to be factually consistent with the input text. A model-agnostic way to address this problem is post-editing the generated summaries. However, existing approaches typi…
InformativenessSentenceSentence CompressionA Simple Yet Effective Corpus Construction Method for Chinese Sentence Compression
Deletion-based sentence compression in the English language has made significant progress over the past few decades. However, there is a lack of large-scale and high-quality parallel corpus (i.e., (sentence, compression)…
SentenceSentence CompressionUnsupervised Abstractive Dialogue Summarization with Word Graphs and POV Conversion
We advance the state-of-the-art in unsupervised abstractive dialogue summarization by utilizing multi-sentence compression graphs. Starting from well-founded assumptions about word graphs, we present simple but reliable …
Abstractive Dialogue SummarizationRerankingSentenceSentence CompressionEfficient Unsupervised Sentence Compression by Fine-tuning Transformers with Reinforcement Learning
Sentence compression reduces the length of text by removing non-essential content while preserving important facts and grammaticality. Unsupervised objective driven methods for sentence compression can be used to create …
reinforcement-learningReinforcement Learning (RL)SentenceSentence Compression+1SOM-NCSCM : An Efficient Neural Chinese Sentence Compression Model Enhanced with Self-Organizing Map
Sentence Compression (SC), which aims to shorten sentences while retaining important words that express the essential meanings, has been studied for many years in many languages, especially in English. However, improveme…
Question AnsweringSentenceSentence CompressionvalidA Novel Metric for Evaluating Semantics Preservation
In this paper, we leverage pre-trained language models (PLMs) to precisely evaluate the semantics preservation of edition process on sentences. Our metric, Neighboring Distribution Divergence (NDD), evaluates the disturb…
Language ModelingLanguage ModellingPredicate DetectionSentence+2Leveraging Information Bottleneck for Scientific Document Summarization
This paper presents an unsupervised extractive approach to summarize scientific long documents based on the Information Bottleneck principle. Inspired by previous work which uses the Information Bottleneck principle for …
Document SummarizationLanguage ModelingLanguage ModellingScientific Document Summarization+2Contextualized Semantic Distance between Highly Overlapped Texts
Overlapping frequently occurs in paired texts in natural language processing tasks like text editing and semantic similarity evaluation. Better evaluation of the semantic distance between the overlapped sentences benefit…
Domain AdaptationLanguage ModelingLanguage ModellingMasked Language Modeling+7Cross-Register Projection for Headline Part of Speech Tagging
Part of speech (POS) tagging is a familiar NLP task. State of the art taggers routinely achieve token-level accuracies of over 97% on news body text, evidence that the problem is well understood. However, the register of…
Open Information ExtractionPart-Of-Speech TaggingPOSPOS Tagging+2RepSum: Unsupervised Dialogue Summarization based on Replacement Strategy
In the field of dialogue summarization, due to the lack of training data, it is often difficult for supervised summary generation methods to learn vital information from dialogue context with limited data. Several attemp…
Dialogue GenerationSentenceSentence CompressionEtat de l’art en compression multi-phrases pour la synthèse de documents (State-of-the-art of multi-sentence compression for document summarization)
La compression multi-phrases est utilisée dans différentes tâches de résumé (microblogs, opinions, réunions ou articles de presse). Leur objectif est de proposer une reformulation compressée et grammaticalement correcte …
ArticlesDocument SummarizationSentenceSentence CompressionNon-Autoregressive Text Generation with Pre-trained Language Models
Non-autoregressive generation (NAG) has recently attracted great attention due to its fast inference speed. However, the generation quality of existing NAG models still lags behind their autoregressive counterparts. In t…
Machine TranslationSentenceSentence CompressionText Generation+2With Measured Words: Simple Sentence Selection for Black-Box Optimization of Sentence Compression Algorithms
Sentence Compression is the task of generating a shorter, yet grammatical version of a given sentence, preserving the essence of the original sentence. This paper proposes a Black-Box Optimizer for Compression (B-BOC): g…
SentenceSentence CompressionEvaluation Discrepancy Discovery: A Sentence Compression Case-study
Reliable evaluation protocols are of utmost importance for reproducible NLP research. In this work, we show that sometimes neither metric nor conventional human evaluation is sufficient to draw conclusions about system p…
SentenceSentence Compression