Papers Text Segmentation
“Text Segmentation” 태그가 달린 논문 124편 · 필터 해제
Self-supervised Character-to-Character Distillation for Text Recognition
When handling complicated text images (e.g., irregular structures, low resolution, heavy occlusion, and uneven illumination), existing supervised text recognition methods are data-hungry. Although these methods employ la…
Data AugmentationRepresentation LearningScene Text RecognitionSelf-Learning+5Toward Unifying Text Segmentation and Long Document Summarization
Text segmentation is important for signaling a document's structure. Without segmenting a long document into topically coherent sections, it is difficult for readers to comprehend the text, let alone find important infor…
ArticlesDocument SummarizationExtractive SummarizationSegmentation+3Structured Summarization: Unified Text Segmentation and Segment Labeling as a Generation Task
Text segmentation aims to divide text into contiguous, semantically coherent segments, while segment labeling deals with producing labels for each segment. Past work has shown success in tackling segmentation and labelin…
DecoderSegmentationText SegmentationOCR for TIFF Compressed Document Images Directly in Compressed Domain Using Text segmentation and Hidden Markov Model
In today's technological era, document images play an important and integral part in our day to day life, and specifically with the surge of Covid-19, digitally scanned documents have become key source of communication, …
Optical Character Recognition (OCR)Text SegmentationDP-Parse: Finding Word Boundaries from Raw Speech with an Instance Lexicon
Finding word boundaries in continuous speech is challenging as there is little or no equivalent of a 'space' delimiter between words. Popular Bayesian non-parametric models for text segmentation use a Dirichlet process t…
Language ModelingLanguage ModellingSegmentationText SegmentationUnsupervised Tokenization Learning
In the presented study, we discover that the so-called "transition freedom" metric appears superior for unsupervised tokenization purposes in comparison to statistical metrics such as mutual information and conditional p…
Text SegmentationTowards Deployable OCR models for Indic languages
Recognition of text on word or line images, without the need for sub-word segmentation has become the mainstream of research and development of text recognition for Indian languages. Modelling unsegmented sequences using…
Optical Character Recognition (OCR)SegmentationText SegmentationTopWORDS-Seg: Simultaneous Text Segmentation and Word Discovery for Open-Domain Chinese Texts via Bayesian Inference
Processing open-domain Chinese texts has been a critical bottleneck in computational linguistics for decades, partially because text segmentation and word discovery often entangle with each other in this challenging scen…
Bayesian InferenceSegmentationText SegmentationSelf-supervised Implicit Glyph Attention for Text Recognition
The attention mechanism has become the \emph{de facto} module in scene text recognition (STR) methods, due to its capability of extracting character-level representations. These methods can be summarized into implicit at…
Scene Text RecognitionText SegmentationFuzzy Segmentations of a String
This article discusses a particular case of the data clustering problem, where it is necessary to find groups of adjacent text segments of the appropriate length that match a fuzzy pattern represented as a sequence of fu…
ClusteringSegmentationText SegmentationBTS: A Bi-Lingual Benchmark for Text Segmentation in the Wild
As a prerequisite of many text-related tasks such as text erasing and text style transfer, text segmentation arouses more and more attention recently. Current researches mainly focus on only English characters and di…
SegmentationStyle TransferText SegmentationText Style TransferWeakly supervised discourse segmentation for multiparty oral conversations
Discourse segmentation, the first step of discourse analysis, has been shown to improve results for text summarization, translation and other NLP tasks. While segmentation models for written text tend to perform well, th…
Discourse SegmentationSegmentationText SegmentationText Summarization+1Transformer over Pre-trained Transformer for Neural Text Segmentation with Enhanced Topic Coherence
This paper proposes a transformer over transformer framework, called Transformer$^2$, to perform neural text segmentation. It consists of two components: bottom-level sentence encoders using pre-trained transformers, and…
SegmentationSentenceSentence EmbeddingsText SegmentationWenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition
In this paper, we present WenetSpeech, a multi-domain Mandarin corpus consisting of 10000+ hours high-quality labeled speech, 2400+ hours weakly labeled speech, and about 10000 hours unlabeled speech, with 22400+ hours i…
Label Error DetectionOptical Character RecognitionOptical Character Recognition (OCR)speech-recognition+2Learning Constraints and Descriptive Segmentation for Subevent Detection
Event mentions in text correspond to real-world events of varying degrees of granularity. The task of subevent detection aims to resolve this granularity issue, recognizing the membership of multi-granular events in even…
DescriptiveText SegmentationSefamerve ARGE at SemEval-2021 Task 5: Toxic Spans Detection Using Segmentation Based 1-D Convolutional Neural Network Model
This paper describes our contribution to SemEval-2021 Task 5: Toxic Spans Detection. Our approach considers toxic spans detection as a segmentation problem. The system, Waw-unet, consists of a 1-D convolutional neural ne…
SegmentationSemantic SegmentationText SegmentationToxic Spans DetectionKIT’s IWSLT 2021 Offline Speech Translation System
This paper describes KIT’submission to the IWSLT 2021 Offline Speech Translation Task. We describe a system in both cascaded condition and end-to-end condition. In the cascaded condition, we investigated different end-to…
Machine Translationspeech-recognitionSpeech RecognitionText Segmentation+1Acoustic Data-Driven Subword Modeling for End-to-End Speech Recognition
Subword units are commonly used for end-to-end automatic speech recognition (ASR), while a fully acoustic-oriented subword modeling approach is somewhat missing. We propose an acoustic data-driven subword modeling (ADSM)…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Segmentationspeech-recognition+2Neural Sequence Segmentation as Determining the Leftmost Segments
Prior methods to text segmentation are mostly at token level. Despite the adequacy, this nature limits their full potential to capture the long-term dependencies among segments. In this work, we propose a novel framework…
ChunkingPart-Of-Speech TaggingPOSPOS Tagging+2Training and Domain Adaptation for Supervised Text Segmentation
Unlike traditional unsupervised text segmentation methods, recent supervised segmentation models rely on Wikipedia as the source of large-scale segmentation supervision. These models have, however, predominantly been eva…
Domain AdaptationHierarchical Text SegmentationSegmentationText Segmentation