paper-with-me

Papers Sentence segmentation

“Sentence segmentation” 태그가 달린 논문 67편 · 필터 해제

Human Genome Book: Words, Sentences and Paragraphs

2025-01-23 · Wang Liang

Since the completion of the human genome sequencing project in 2001, significant progress has been made in areas such as gene regulation editing and protein structure prediction. However, given the vast amount of genomic…

Protein Structure PredictionSentence segmentationTransfer Learning

Segmentation en phrases : ouvrez les guillemets sans perdre le fil

2024-07-29 · Sandrine Ollinger, Denis Maurel

This paper presents a graph cascade for sentence segmentation of XML documents. Our proposal offers sentences inside sentences for cases introduced by quotation marks and hyphens, and also pays particular attention to si…

SentenceSentence segmentation

Segment Any Text: A Universal Approach for Robust, Efficient and Adaptable Sentence Segmentation

2024-06-24 · Markus Frohmann, Igor Sterner, Ivan Vulić, Benjamin Minixhofer 외

Segmenting text into sentences plays an early and crucial role in many NLP systems. This is commonly achieved by using rule-based or statistical methods relying on lexical features such as punctuation. Although some rece…

parameter-efficient fine-tuningSentenceSentence segmentation

Opera Graeca Adnotata: Building a 34M+ Token Multilayer Corpus for Ancient Greek

2024-03-31 · Giuseppe G. A. Celano

In this article, the beta version 0.1.0 of Opera Graeca Adnotata (OGA), the largest open-access multilayer corpus for Ancient Greek (AG) is presented. OGA consists of 1,687 literary works and 34M+ tokens coming from the …

LemmatizationSentenceSentence segmentation

Ascle: A Python Natural Language Processing Toolkit for Medical Text Generation

2023-11-28 · Rui Yang, Qingcheng Zeng, Keen You, Yujie Qiao 외

This study introduces Ascle, a pioneering natural language processing (NLP) toolkit designed for medical text generation. Ascle is tailored for biomedical researchers and healthcare professionals with an easy-to-use, all…

Machine TranslationQuestion AnsweringSentence segmentationText Generation+3

KG-GPT: A General Framework for Reasoning on Knowledge Graphs Using Large Language Models

2023-10-17 · Jiho Kim, Yeonsu Kwon, Yohan Jo, Edward Choi

While large language models (LLMs) have made considerable advancements in understanding and generating unstructured text, their application in structured data remains underexplored. Particularly, using LLMs for complex r…

Fact VerificationKnowledge GraphsRetrievalSentence+1

GujiBERT and GujiGPT: Construction of Intelligent Information Processing Foundation Language Models for Ancient Texts

2023-07-11 · Dongbo Wang, Chang Liu, Zhixiao Zhao, Si Shen 외

In the context of the rapid development of large language models, we have meticulously trained and introduced the GujiBERT and GujiGPT language models, which are foundational models specifically designed for intelligent …

Model SelectionPart-Of-Speech TaggingSentenceSentence segmentation

Where's the Point? Self-Supervised Multilingual Punctuation-Agnostic Sentence Segmentation

2023-05-30 · Benjamin Minixhofer, Jonas Pfeiffer, Ivan Vulić

Many NLP pipelines split text into sentences as one of the crucial preprocessing steps. Prior sentence segmentation tools either rely on punctuation or require a considerable amount of sentence-segmented training data: b…

Machine TranslationSegmentationSentenceSentence segmentation

Prosodic features improve sentence segmentation and parsing

2023-02-23 · Elizabeth Nielsen, Sharon Goldwater, Mark Steedman

Parsing spoken dialogue presents challenges that parsing text does not, including a lack of clear sentence boundaries. We know from previous work that prosody helps in parsing single sentences (Tran et al. 2018), but we …

SentenceSentence segmentation

Sentence Identification with BOS and EOS Label Combinations

2023-01-31 · Takuma Udagawa, Hiroshi Kanayama, Issei Yoshida

The sentence is a fundamental unit in many NLP applications. Sentence segmentation is widely used as the first preprocessing task, where an input text is split into consecutive sentences considering the end of the senten…

SentenceSentence segmentation

SLATE: A Sequence Labeling Approach for Task Extraction from Free-form Inked Content

2022-11-08 · Apurva Gandhi, Ryan Serrao, Biyi Fang, Gilbert Antonius 외

We present SLATE, a sequence labeling approach for extracting tasks from free-form content such as digitally handwritten (or "inked") notes on a virtual whiteboard. Our approach allows us to create a single, low-latency …

FormSegmentationSentenceSentence segmentation

LeConTra: A Learner Corpus of English-to-Dutch News Translation

2022-06-01 · LREC 2022 6 · Bram Vanroy, Lieve Macken

We present LeConTra, a learner corpus consisting of English-to-Dutch news translations enriched with translation process data. Three students of a Master’s programme in Translation were asked to translate 50 different En…

SentenceSentence segmentationTranslation

Midas Loop: A Prioritized Human-in-the-Loop Annotation for Large Scale Multilayer Data

2022-06-01 · LREC (LAW) 2022 6 · Luke Gessler, Lauren Levine, Amir Zeldes

Large scale annotation of rich multilayer corpus data is expensive and time consuming, motivating approaches that integrate high quality automatic tools with active learning in order to prioritize human labeling of hard …

Active LearningManagementSegmentationSentence+1

Mukayese: Turkish NLP Strikes Back

2022-03-02 · Findings (ACL) 2022 5 · Ali Safaya, Emirhan Kurtuluş, Arda Göktoğan, Deniz Yuret

Having sufficient resources for language X lifts it from the under-resourced languages class, but not necessarily from the under-researched class. In this paper, we address the problem of the absence of organized benchma…

BenchmarkingLanguage ModelingLanguage ModellingSentence+1

Mukayese: Turkish NLP Strikes Back

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Having sufficient resources for a language X lifts it from the $\textit{under-resourced}$ languages class, but does not necessarily lift it from the $\textit{under-researched}$ class. In this paper, we address the proble…

BenchmarkingLanguage ModelingLanguage ModellingSentence+1

CUNI Systems in WMT21: Revisiting Backtranslation Techniques for English-Czech NMT

2021-11-01 · WMT (EMNLP) 2021 11 · Petr Gebauer, Ondřej Bojar, Vojtěch Švandelík, Martin Popel

We describe our two NMT systems submitted to the WMT2021 shared task in English-Czech news translation: CUNI-DocTransformer (document-level CUBBITT) and CUNI-Marian-Baselines. We improve the former with a better sentence…

NMTSegmentationSentenceSentence segmentation+1

Transformer-Encoder-GRU (T-E-GRU) for Chinese Sentiment Analysis on Chinese Comment Text

2021-08-01 · Binlong Zhang, Wei Zhou

Chinese sentiment analysis (CSA) has always been one of the challenges in natural language processing due to its complexity and uncertainty. Transformer has succeeded in capturing semantic features, but it uses position …

Chinese Sentiment AnalysisPositionSentenceSentence segmentation+1

A unified approach to sentence segmentation of punctuated text in many languages

2021-08-01 · ACL 2021 5 · Rachel Wicks, Matt Post

The sentence is a fundamental unit of text processing. Yet sentences in the wild are commonly encountered not in isolation, but unsegmented within larger paragraphs and documents. Therefore, the first step in many NLP pi…

SentenceSentence segmentation

The Reading Machine: A Versatile Framework for Studying Incremental Parsing Strategies

2021-08-01 · ACL (IWPT) 2021 8 · Franck Dary, Alexis Nasr

The Reading Machine, is a parsing framework that takes as input raw text and performs six standard nlp tasks: tokenization, pos tagging, morphological analysis, lemmatization, dependency parsing and sentence segmentation…

Dependency ParsingLemmatizationMorphological AnalysisPOS+3

Better Chinese Sentence Segmentation with Reinforcement Learning

2021-08-01 · Findings (ACL) 2021 8 · Srivatsan Srinivasan, Chris Dyer
reinforcement-learningReinforcement LearningReinforcement Learning (RL)Sentence+1
1–20 / 67 다음 →