Rule-Based Approaches to Atomic Sentence Extraction
Natural language often combines multiple ideas into complex sentences. Atomic sentence extraction, the task of decomposing complex sentences into simpler sentences that each express a single idea, improves performance in information retrieval, question answering, and automated reasoning systems. Previous work has formalized the "split-and-rephrase" task and established evaluation metrics, and machine learning approaches using large language models have improved extraction accuracy. However, these methods lack interpretability and provide limited insight into which linguistic structures cause extraction failures. Although some studies have explored dependency-based extraction of subject-verb-object triples and clauses, no principled analysis has examined which specific clause structures and dependencies lead to extraction difficulties. This study addresses this gap by analyzing how complex sentence structures, including relative clauses, adverbial clauses, coordination patterns, and passive constructions, affect the performance of rule-based atomic sentence extraction. Using the WikiSplit dataset, we implemented dependency-based extraction rules in spaCy, generated 100 gold=standard atomic sentence sets, and evaluated performance using ROUGE and BERTScore. The system achieved ROUGE-1 F1 = 0.6714, ROUGE-2 F1 = 0.478, ROUGE-L F1 = 0.650, and BERTScore F1 = 0.5898, indicating moderate-to-high lexical, structural, and semantic alignment. Challenging structures included relative clauses, appositions, coordinated predicates, adverbial clauses, and passive constructions. Overall, rule-based extraction is reasonably accurate but sensitive to syntactic complexity.
Code (0)
등록된 구현이 없습니다.
Tasks
Information RetrievalQuestion AnsweringSimilar Papers 제목 키워드 기반
A Sentence Simplification System for Improving Relation Extraction
In this demo paper, we present a text simplification approach that is directed at improving the performance of state-of-the-art Open Relation Extraction (RE) systems. As syntactically complex sentences often pose a chall…
RelationRelation ExtractionSentenceText SimplificationGUIDO: A Hybrid Approach to Guideline Discovery & Ordering from Natural Language Texts
Extracting workflow nets from textual descriptions can be used to simplify guidelines or formalize textual descriptions of formal processes like business processes and algorithms. The task of manually extracting processe…
Dependency ParsingModel extractionSentenceSpecificityAutomatic Extraction of Learner Errors in ESL Sentences Using Linguistically Enhanced Alignments
We propose a new method of automatically extracting learner errors from parallel English as a Second Language (ESL) sentences in an effort to regularise annotation formats and reduce inconsistencies. Specifically, given …
Grammatical Error CorrectionMachine TranslationSentenceNERO: A Neural Rule Grounding Framework for Label-Efficient Relation Extraction
Deep neural models for relation extraction tend to be less reliable when perfectly labeled data is limited, despite their success in label-sufficient scenarios. Instead of seeking more instance-level labels from human an…
RelationRelation ExtractionSentenceABCD: A Graph Framework to Convert Complex Sentences to a Covering Set of Simple Sentences
Atomic clauses are fundamental text units for understanding complex sentences. Identifying the atomic sentences within complex sentences is important for applications such as summarization, argument mining, discourse ana…
Argument MiningDecoderDiscourse ParsingEDIT Task+3