Papers Sentence segmentation
“Sentence segmentation” 태그가 달린 논문 67편 · 필터 해제
Human Genome Book: Words, Sentences and Paragraphs
Since the completion of the human genome sequencing project in 2001, significant progress has been made in areas such as gene regulation editing and protein structure prediction. However, given the vast amount of genomic…
Protein Structure PredictionSentence segmentationTransfer LearningSegmentation en phrases : ouvrez les guillemets sans perdre le fil
This paper presents a graph cascade for sentence segmentation of XML documents. Our proposal offers sentences inside sentences for cases introduced by quotation marks and hyphens, and also pays particular attention to si…
SentenceSentence segmentationSegment Any Text: A Universal Approach for Robust, Efficient and Adaptable Sentence Segmentation
Segmenting text into sentences plays an early and crucial role in many NLP systems. This is commonly achieved by using rule-based or statistical methods relying on lexical features such as punctuation. Although some rece…
parameter-efficient fine-tuningSentenceSentence segmentationOpera Graeca Adnotata: Building a 34M+ Token Multilayer Corpus for Ancient Greek
In this article, the beta version 0.1.0 of Opera Graeca Adnotata (OGA), the largest open-access multilayer corpus for Ancient Greek (AG) is presented. OGA consists of 1,687 literary works and 34M+ tokens coming from the …
LemmatizationSentenceSentence segmentationAscle: A Python Natural Language Processing Toolkit for Medical Text Generation
This study introduces Ascle, a pioneering natural language processing (NLP) toolkit designed for medical text generation. Ascle is tailored for biomedical researchers and healthcare professionals with an easy-to-use, all…
Machine TranslationQuestion AnsweringSentence segmentationText Generation+3KG-GPT: A General Framework for Reasoning on Knowledge Graphs Using Large Language Models
While large language models (LLMs) have made considerable advancements in understanding and generating unstructured text, their application in structured data remains underexplored. Particularly, using LLMs for complex r…
Fact VerificationKnowledge GraphsRetrievalSentence+1GujiBERT and GujiGPT: Construction of Intelligent Information Processing Foundation Language Models for Ancient Texts
In the context of the rapid development of large language models, we have meticulously trained and introduced the GujiBERT and GujiGPT language models, which are foundational models specifically designed for intelligent …
Model SelectionPart-Of-Speech TaggingSentenceSentence segmentationWhere's the Point? Self-Supervised Multilingual Punctuation-Agnostic Sentence Segmentation
Many NLP pipelines split text into sentences as one of the crucial preprocessing steps. Prior sentence segmentation tools either rely on punctuation or require a considerable amount of sentence-segmented training data: b…
Machine TranslationSegmentationSentenceSentence segmentationProsodic features improve sentence segmentation and parsing
Parsing spoken dialogue presents challenges that parsing text does not, including a lack of clear sentence boundaries. We know from previous work that prosody helps in parsing single sentences (Tran et al. 2018), but we …
SentenceSentence segmentationSentence Identification with BOS and EOS Label Combinations
The sentence is a fundamental unit in many NLP applications. Sentence segmentation is widely used as the first preprocessing task, where an input text is split into consecutive sentences considering the end of the senten…
SentenceSentence segmentationSLATE: A Sequence Labeling Approach for Task Extraction from Free-form Inked Content
We present SLATE, a sequence labeling approach for extracting tasks from free-form content such as digitally handwritten (or "inked") notes on a virtual whiteboard. Our approach allows us to create a single, low-latency …
FormSegmentationSentenceSentence segmentationLeConTra: A Learner Corpus of English-to-Dutch News Translation
We present LeConTra, a learner corpus consisting of English-to-Dutch news translations enriched with translation process data. Three students of a Master’s programme in Translation were asked to translate 50 different En…
SentenceSentence segmentationTranslationMidas Loop: A Prioritized Human-in-the-Loop Annotation for Large Scale Multilayer Data
Large scale annotation of rich multilayer corpus data is expensive and time consuming, motivating approaches that integrate high quality automatic tools with active learning in order to prioritize human labeling of hard …
Active LearningManagementSegmentationSentence+1Mukayese: Turkish NLP Strikes Back
Having sufficient resources for language X lifts it from the under-resourced languages class, but not necessarily from the under-researched class. In this paper, we address the problem of the absence of organized benchma…
BenchmarkingLanguage ModelingLanguage ModellingSentence+1Mukayese: Turkish NLP Strikes Back
Having sufficient resources for a language X lifts it from the $\textit{under-resourced}$ languages class, but does not necessarily lift it from the $\textit{under-researched}$ class. In this paper, we address the proble…
BenchmarkingLanguage ModelingLanguage ModellingSentence+1CUNI Systems in WMT21: Revisiting Backtranslation Techniques for English-Czech NMT
We describe our two NMT systems submitted to the WMT2021 shared task in English-Czech news translation: CUNI-DocTransformer (document-level CUBBITT) and CUNI-Marian-Baselines. We improve the former with a better sentence…
NMTSegmentationSentenceSentence segmentation+1Transformer-Encoder-GRU (T-E-GRU) for Chinese Sentiment Analysis on Chinese Comment Text
Chinese sentiment analysis (CSA) has always been one of the challenges in natural language processing due to its complexity and uncertainty. Transformer has succeeded in capturing semantic features, but it uses position …
Chinese Sentiment AnalysisPositionSentenceSentence segmentation+1A unified approach to sentence segmentation of punctuated text in many languages
The sentence is a fundamental unit of text processing. Yet sentences in the wild are commonly encountered not in isolation, but unsegmented within larger paragraphs and documents. Therefore, the first step in many NLP pi…
SentenceSentence segmentationThe Reading Machine: A Versatile Framework for Studying Incremental Parsing Strategies
The Reading Machine, is a parsing framework that takes as input raw text and performs six standard nlp tasks: tokenization, pos tagging, morphological analysis, lemmatization, dependency parsing and sentence segmentation…
Dependency ParsingLemmatizationMorphological AnalysisPOS+3