paper-with-me

Papers

Text Region Multiple Information Perception Network for Scene Text Detection

2024-01-18 · Jinzhi Zheng, Libo Zhang, Yanjun Wu, Chen Zhao

Segmentation-based scene text detection algorithms can handle arbitrary shape scene texts and have strong robustness and adaptability, so it has attracted wide attention. Existing segmentation-based scene text detection algorithms usually only segment the pixels in the center region of the text, while ignoring other information of the text region, such as edge information, distance information, etc., thus limiting the detection accuracy of the algorithm for scene text. This paper proposes a plug-and-play module called the Region Multiple Information Perception Module (RMIPM) to enhance the detection performance of segmentation-based algorithms. Specifically, we design an improved module that can perceive various types of information about scene text regions, such as text foreground classification maps, distance maps, direction maps, etc. Experiments on MSRA-TD500 and TotalText datasets show that our method achieves comparable performance with current state-of-the-art algorithms.

📄 PDF Abstract BibTeX arXiv:2401.10017

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Text DetectionSegmentationText Detection

Similar Papers 제목 키워드 기반

Local semantic enhanced convnet for aerial scene recognition

2021-07-08 · IEEE Transactions on Image Processing 2021 7 · Qi Bi, Kun Qin, Han Zhang, Gui-Song Xia

Aerial scene recognition is challenging due to the complicated object distribution and spatial arrangement in a large-scale aerial image. Recent studies attempt to explore the local semantic representation capability of …

Aerial Scene ClassificationImage ClassificationScene ClassificationScene Recognition

Aligning where to see and what to tell: image caption with region-based attention and scene factorization

2015-06-20 · Junqi Jin, Kun fu, Runpeng Cui, Fei Sha 외

Recent progress on automatic generation of image captions has shown that it is possible to describe the most salient information conveyed by images with accurate and meaningful sentences. In this paper, we propose an ima…

Image Captioning

Contextually-rich human affect perception using multimodal scene information

2023-03-13 · Digbalay Bose, Rajat Hebbar, Krishna Somandepalli, Shrikanth Narayanan

The process of human affect understanding involves the ability to infer person specific emotional states from various sources including images, speech, and language. Affect perception from images has predominantly focuse…

Perceive What Matters: Relevance-Driven Scheduling for Multimodal Streaming Perception

2026-03-13 · Dingcheng Huang, Xiaotong Zhang, Kamal Youcef-Toumi arxiv

In modern human-robot collaboration (HRC) applications, multiple perception modules jointly extract visual, auditory, and contextual cues to achieve comprehensive scene understanding, enabling the robot to provide approp…

Scene Understanding

Sparse-Aware Vector Quantization for Bandwidth-Efficient Collaborative 3D Semantic Occupancy Prediction

2026-07-02 · Feng Li, Chaokun Zhang, Gong Chen arxiv

Collaborative perception extends single-agent perception by enabling multiple vehicles to exchange complementary perceptual information. However, it introduces an inherent trade-off between perception gain and communicat…