paper-with-me

Papers

Guidance-Based Prompt Data Augmentation in Specialized Domains for Named Entity Recognition

2024-07-26 · Hyeonseok Kang, Hyein Seo, Jeesu Jung, SangKeun Jung, Du-Seong Chang, Riwoo Chung

While the abundance of rich and vast datasets across numerous fields has facilitated the advancement of natural language processing, sectors in need of specialized data types continue to struggle with the challenge of finding quality data. Our study introduces a novel guidance data augmentation technique utilizing abstracted context and sentence structures to produce varied sentences while maintaining context-entity relationships, addressing data scarcity challenges. By fostering a closer relationship between context, sentence structure, and role of entities, our method enhances data augmentation's effectiveness. Consequently, by showcasing diversification in both entity-related vocabulary and overall sentence structure, and simultaneously improving the training performance of named entity recognition task.

📄 PDF Abstract BibTeX arXiv:2407.18442

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentationnamed-entity-recognitionNamed Entity RecognitionSentence

Similar Papers 제목 키워드 기반

The arText prototype: An automatic system for writing specialized texts

2017-04-01 · EACL 2017 4 · Iria da Cunha, M. Amor Montan{\'e}, Luis Hysa

This article describes an automatic system for writing specialized texts in Spanish. The arText prototype is a free online text editor that includes different types of linguistic information. It is designed for a variety…

SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection

2026-03-23 · Kexian Tang, Jiani Wang, Shaowen Wang, Kaifeng Lyu arxiv

While large language models (LLMs) are pretrained on massive amounts of data, their knowledge coverage remains incomplete in specialized, data-scarce domains, motivating extensive efforts to study synthetic data generati…

Synthetic Data GenerationData Augmentation

DALDALL: Data Augmentation for Lexical and Semantic Diverse in Legal Domain by leveraging LLM-Persona

2026-03-24 · Janghyeok Choi, Jaewon Lee, Sungzoon Cho arxiv

Data scarcity remains a persistent challenge in low-resource domains. While existing data augmentation methods leverage the generative capabilities of large language models (LLMs) to produce large volumes of synthetic da…

Information RetrievalData Augmentation

Role Prompting Guided Domain Adaptation with General Capability Preserve for Large Language Models

2024-03-05 · Rui Wang, Fei Mi, Yi Chen, Boyang Xue 외

The growing interest in Large Language Models (LLMs) for specialized applications has revealed a significant challenge: when tailored to specific domains, LLMs tend to experience catastrophic forgetting, compromising the…

Domain Adaptation

Task Oriented In-Domain Data Augmentation

2024-06-24 · Xiao Liang, Xinyu Hu, Simiao Zuo, Yeyun Gong 외

Large Language Models (LLMs) have shown superior performance in various applications and fields. To achieve better performance on specialized domains such as law and advertisement, LLMs are often continue pre-trained on …

Data AugmentationMath