paper-with-me

Papers

DUET: Detection Utilizing Enhancement for Text in Scanned or Captured Documents

2021-06-10 · Eun-Soo Jung, HyeongGwan Son, Kyusam Oh, Yongkeun Yun, Soonhwan Kwon, Min Soo Kim

We present a novel deep neural model for text detection in document images. For robust text detection in noisy scanned documents, the advantages of multi-task learning are adopted by adding an auxiliary task of text enhancement. Namely, our proposed model is designed to perform noise reduction and text region enhancement as well as text detection. Moreover, we enrich the training data for the model with synthesized document images that are fully labeled for text detection and enhancement, thus overcome the insufficiency of labeled document image data. For the effective exploitation of the synthetic and real data, the training process is separated in two phases. The first phase is training only synthetic data in a fully-supervised manner. Then real data with only detection labels are added in the second phase. The enhancement task for the real data is weakly-supervised with information from their detection labels. Our methods are demonstrated in a real document dataset with performances exceeding those of other text detection methods. Moreover, ablations are conducted and the results confirm the effectiveness of the synthetic data, auxiliary task, and weak-supervision. Whereas the existing text detection studies mostly focus on the text in scenes, our proposed method is optimized to the applications for the text in scanned documents.

📄 PDF Abstract BibTeX arXiv:2106.05542

Code (0)

등록된 구현이 없습니다.

Tasks

Multi-Task LearningText Detection

Similar Papers 제목 키워드 기반

Document Enhancement System Using Auto-encoders

2019-09-14 · NeurIPS Workshop Document_Intelligen 2019 12 · Mehrdad J. Gangeh, Sunil R. Tiyyagura, Sridhar V. Dasaratha, Hamid Motahari 외

The conversion of scanned documents to digital forms is performed using an Optical Character Recognition (OCR) software. This work focuses on improving the quality of scanned documents in order to improve the OCR output.…

DenoisingDocument EnhancementOptical Character RecognitionOptical Character Recognition (OCR)

VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format

2024-11-27 · Yueqian Wang, Xiaojun Meng, Yuxuan Wang, Jianxin Liang 외

Recent researches on video large language models (VideoLLM) predominantly focus on model architectures and training datasets, leaving the interaction format between the user and the model under-explored. In existing work…

Dense Video CaptioningGrounded Video Question AnsweringHighlight DetectionQuestion Answering+3

Word-Entity Duet Representations for Document Ranking

2017-06-20 · Chenyan Xiong, Jamie Callan, Tie-Yan Liu

This paper presents a word-entity duet framework for utilizing knowledge bases in ad-hoc retrieval. In this work, the query and documents are modeled by word-based representations and entity-based representations. Rankin…

Document RankingLearning-To-RankRetrieval

MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation

2025-08-23 · Prerit Gupta, Jason Alexander Fotso-Puepi, Zhengyuan Li, Jay Mehta 외 arxiv

We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comprises 620 minutes of high-quality motion…

Detection of Uncertainty in Exceedance of Threshold (DUET): An Adversarial Patch Localizer

2023-03-18 · Terence Jie Chua, Wenhan Yu, Jun Zhao

Development of defenses against physical world attacks such as adversarial patches is gaining traction within the research community. We contribute to the field of adversarial patch detection by introducing an uncertaint…

Self-Driving Cars