Papers Text Segmentation
“Text Segmentation” 태그가 달린 논문 124편 · 필터 해제
The impact of fine tuning in LLaMA on hallucinations for named entity extraction in legal documentation
The extraction of information about traffic accidents from legal documents is crucial for quantifying insurance company costs. Extracting entities such as percentages of physical and/or psychological disability and the i…
Text SegmentationBP-Seg: A graphical model approach to unsupervised and non-contiguous text segmentation using belief propagation
Text segmentation based on the semantic meaning of sentences is a fundamental task with broad utility in many downstream applications. In this paper, we propose a graphical model-based unsupervised learning approach, nam…
SegmentationText SegmentationBR-TaxQA-R: A Dataset for Question Answering with References for Brazilian Personal Income Tax Law, including case law
This paper presents BR-TaxQA-R, a novel dataset designed to support question answering with references in the context of Brazilian personal income tax law. The dataset contains 715 questions from the 2024 official Q\&A d…
Answer GenerationQuestion AnsweringRAGRetrieval-augmented Generation+1TSAL: Few-shot Text Segmentation Based on Attribute Learning
Recently supervised learning rapidly develops in scene text segmentation. However, the lack of high-quality datasets and the high cost of pixel annotation greatly limit the development of them. Considering the well-perfo…
AttributeFew-Shot LearningSegmentationText Detection+1POSTA: A Go-to Framework for Customized Artistic Poster Generation
Poster design is a critical medium for visual communication. Prior work has explored automatic poster design using deep learning techniques, but these approaches lack text accuracy, user customization, and aesthetic appe…
Text SegmentationA Context-Driven Training-Free Network for Lightweight Scene Text Segmentation and Recognition
Modern scene text recognition systems often depend on large end-to-end architectures that require extensive training and are prohibitively expensive for real-time scenarios. In such cases, the deployment of heavy models …
Scene Text RecognitionText DetectionText SegmentationControlText: Unlocking Controllable Fonts in Multilingual Text Rendering without Font Annotations
This work demonstrates that diffusion models can achieve font-controllable multilingual text rendering using just raw images without font label annotations. Visual text rendering remains a significant challenge. While re…
Text SegmentationSegmenting Text and Learning Their Rewards for Improved RLHF in Language Model
Reinforcement learning from human feedback (RLHF) has been widely adopted to align language models (LMs) with human preference. Prior RLHF works typically take a bandit formulation, which, though intuitive, ignores the s…
Language ModelingLanguage ModellingText SegmentationDocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning
Document image segmentation is crucial in document analysis and recognition but remains challenging due to the heterogeneity of document formats and diverse segmentation tasks. Existing methods often treat these task…
Document Layout AnalysisImage SegmentationSegmentationSemantic Segmentation+2Char-SAM: Turning Segment Anything Model into Scene Text Segmentation Annotator with Character-level Visual Prompts
The recent emergence of the Segment Anything Model (SAM) enables various domain-specific segmentation tasks to be tackled cost-effectively by using bounding boxes as prompts. However, in scene text segmentation, SAM can …
SegmentationText DetectionText SegmentationSegment-Level Diffusion: A Framework for Controllable Long-Form Generation with Diffusion Language Models
Diffusion models have shown promise in text generation but often struggle with generating long, coherent, and contextually accurate text. Token-level diffusion overlooks word-order dependencies and enforces short output …
Contrastive LearningDecoderFormRepresentation Learning+2ROSE: Revolutionizing Open-Set Dense Segmentation with Patch-Wise Perceptual Large Multimodal Model
Advances in CLIP and large multimodal models (LMMs) have enabled open-vocabulary and free-text segmentation, yet existing models still require predefined category prompts, limiting free-form category self-generation. Mos…
PredictionSegmentationText SegmentationRecent Trends in Linear Text Segmentation: a Survey
Linear Text Segmentation is the task of automatically tagging text documents with topic shifts, i.e. the places in the text where the topics change. A well-established area of research in Natural Language Processing, dra…
SegmentationSurveyText SegmentationEnhancing Question Answering Precision with Optimized Vector Retrieval and Instructions
Question-answering (QA) is an important application of Information Retrieval (IR) and language models, and the latest trend is toward pre-trained large neural networks with embedding parameters. Augmenting QA performance…
Document EmbeddingInformation RetrievalQuestion AnsweringRetrieval+2Meta-Chunking: Learning Text Segmentation and Semantic Completion via Logical Perception
While Retrieval-Augmented Generation (RAG) has emerged as a promising paradigm for boosting large language models (LLMs) in knowledge-intensive tasks, it often overlooks the crucial aspect of text chunking within its wor…
Binary ClassificationChunkingInstruction FollowingMulti-hop Question Answering+4EEG-Language Modeling for Pathology Detection
Multimodal language modeling has enabled breakthroughs for representation learning, yet remains unexplored in the realm of functional brain data for pathology detection. This paper pioneers EEG-language models (ELMs) tra…
Contrastive LearningEEGLanguage ModelingLanguage Modelling+6WAS: Dataset and Methods for Artistic Text Segmentation
Accurate text segmentation results are crucial for text-related generative tasks, such as text image generation, text editing, text removal, and text style transfer. Recently, some scene text segmentation methods have ma…
DecoderDiversityImage GenerationSegmentation+3EAFormer: Scene Text Segmentation with Edge-Aware Transformers
Scene text segmentation aims at cropping texts from scene images, which is usually used to help generative models edit or remove texts. The existing text segmentation methods tend to involve various text-related supervis…
DecoderSegmentationText SegmentationTocBERT: Medical Document Structure Extraction Using Bidirectional Transformers
Text segmentation holds paramount importance in the field of Natural Language Processing (NLP). It plays an important role in several NLP downstream tasks like information retrieval and document summarization. In this wo…
Document SummarizationHierarchical Text SegmentationInformation Retrievalnamed-entity-recognition+5Automating Easy Read Text Segmentation
Easy Read text is one of the main forms of access to information for people with reading difficulties. One of the key characteristics of this type of text is the requirement to split sentences into smaller grammatical se…
SegmentationText Segmentation