paper-with-me

홈 › Papers

Towards Unified Multi-granularity Text Detection with Interactive Attention

2024-05-30 · Xingyu Wan, Chengquan Zhang, Pengyuan Lyu, Sen Fan, Zihan Ni, Kun Yao, Errui Ding, Jingdong Wang

Existing OCR engines or document image analysis systems typically rely on training separate models for text detection in varying scenarios and granularities, leading to significant computational complexity and resource demands. In this paper, we introduce "Detect Any Text" (DAT), an advanced paradigm that seamlessly unifies scene text detection, layout analysis, and document page detection into a cohesive, end-to-end model. This design enables DAT to efficiently manage text instances at different granularities, including *word*, *line*, *paragraph* and *page*. A pivotal innovation in DAT is the across-granularity interactive attention module, which significantly enhances the representation learning of text instances at varying granularities by correlating structural information across different text queries. As a result, it enables the model to achieve mutually beneficial detection performances across multiple text granularities. Additionally, a prompt-based segmentation module refines detection outcomes for texts of arbitrary curvature and complex layouts, thereby improving DAT's accuracy and expanding its real-world applicability. Experimental results demonstrate that DAT achieves state-of-the-art performances across a variety of text-related benchmarks, including multi-oriented/arbitrarily-shaped scene text detection, document layout analysis and page detection tasks.

📄 PDF Abstract BibTeX arXiv:2405.19765

Code (0)

등록된 구현이 없습니다.

Tasks

Document Layout AnalysisOptical Character Recognition (OCR)Representation LearningScene Text DetectionText Detection

Similar Papers 제목 키워드 기반

Unified Coarse-to-Fine Alignment for Video-Text Retrieval

2023-09-18 · ICCV 2023 1 · Ziyang Wang, Yi-Lin Sung, Feng Cheng, Gedas Bertasius 외

The canonical approach to video-text retrieval leverages a coarse-grained or fine-grained alignment between visual and textual information. However, retrieving the correct video according to the text query is often chall…

RetrievalText RetrievalText to Video RetrievalVideo Retrieval+1

GraCo: Granularity-Controllable Interactive Segmentation

2024-05-01 · CVPR 2024 1 · Yian Zhao, Kehan Li, Zesen Cheng, Pengchong Qiao 외

Interactive Segmentation (IS) segments specific objects or parts in the image according to user input. Current IS pipelines fall into two categories: single-granularity output and multi-granularity output. The latter aim…

Interactive SegmentationSegmentation

Multi-granularity Interactive Attention Framework for Residual Hierarchical Pronunciation Assessment

2026-01-05 · Hong Han, Hao-Chen Pei, Zhao-Zheng Nie, Xin Luo 외 arxiv

Automatic pronunciation assessment plays a crucial role in computer-assisted pronunciation training systems. Due to the ability to perform multiple pronunciation tasks simultaneously, multi-aspect multi-granularity pronu…

IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring

2026-05-23 · Xinge Peng, Yiting Lu, Xin Li, Zhibo Chen arxiv

We present IQA-Spider, the first image quality assessment (IQA) framework that unifies reasoning, grounding, and referring into a single LMM-based framework for multi-granularity quality understanding. Existing LMM-based…

Image Quality AssessmentQuestion Answering

Towards Efficient and Robust Moment Retrieval System: A Unified Framework for Multi-Granularity Models and Temporal Reranking

2025-04-11 · Huu-Loc Tran, Tinh-Anh Nguyen-Nhu, Huu-Phong Phan-Nguyen, Tien-Huy Nguyen 외

Long-form video understanding presents significant challenges for interactive retrieval systems, as conventional methods struggle to process extensive video content efficiently. Existing approaches often rely on single m…

Moment RetrievalQuestion AnsweringRerankingRetrieval+2