paper-with-me

홈 › Papers

Class-Aware Mask-Guided Feature Refinement for Scene Text Recognition

2024-02-21 · Mingkun Yang, Biao Yang, Minghui Liao, Yingying Zhu, Xiang Bai

Scene text recognition is a rapidly developing field that faces numerous challenges due to the complexity and diversity of scene text, including complex backgrounds, diverse fonts, flexible arrangements, and accidental occlusions. In this paper, we propose a novel approach called Class-Aware Mask-guided feature refinement (CAM) to address these challenges. Our approach introduces canonical class-aware glyph masks generated from a standard font to effectively suppress background and text style noise, thereby enhancing feature discrimination. Additionally, we design a feature alignment and fusion module to incorporate the canonical mask guidance for further feature refinement for text recognition. By enhancing the alignment between the canonical mask feature and the text feature, the module ensures more effective fusion, ultimately leading to improved recognition performance. We first evaluate CAM on six standard text recognition benchmarks to demonstrate its effectiveness. Furthermore, CAM exhibits superiority over the state-of-the-art method by an average performance gain of 4.1% across six more challenging datasets, despite utilizing a smaller model size. Our study highlights the importance of incorporating canonical mask guidance and aligned feature refinement techniques for robust scene text recognition. The code is available at https://github.com/MelosY/CAM.

📄 PDF Abstract BibTeX arXiv:2402.13643

Code (2)

melosy/cam 공식 구현 pytorch
topdu/openocr pytorch

Tasks

DiversityScene Text Recognition

Methods 이 논문이 사용한 방법론

CAM Class activation maps could be used to interpret the prediction decision made by the convolutional neural network (CNN). Image source: [Learning Deep Features for…

Similar Papers 제목 키워드 기반

RUFNet: Query-Guided Support Mask Refinement and Uncertainty Fusion based on Hybrid Mamba for Few-Shot Brain Tumor Segmentation

2026-07-06 · Dongyi He, Xiangkai Wang, Binbing Xu, Bin Jiang 외 arxiv

Few-shot brain tumor segmentation remains challenging due to noisy support masks, inter-patient variations between support and query images, and the lack of pixel-wise confidence estimation. This study proposes RUFNet, a…

Medical Image SegmentationBrain Tumor Segmentation

Enhancing Image Matting in Real-World Scenes with Mask-Guided Iterative Refinement

2025-02-24 · Rui Liu

Real-world image matting is essential for applications in content creation and augmented reality. However, it remains challenging due to the complex nature of scenes and the scarcity of high-quality datasets. To address …

Benchmarkingfeature selectionImage Matting

Explainability-Guided Defense: Attribution-Aware Model Refinement Against Adversarial Data Attacks

2026-01-02 · Longwei Wang, Mohammad Navid Nayyem, Abdullah Al Rakin, KC Santosh 외 arxiv

The growing reliance on deep learning models in safety-critical domains such as healthcare and autonomous navigation underscores the need for defenses that are both robust to adversarial perturbations and transparent in …

Adversarial Robustness

Adapting SAM to Nuclei Instance Segmentation and Classification via Cooperative Fine-Grained Refinement

2026-03-30 · Jingze Su, Tianle Zhu, Jiaxin Cai, Zhiyi Wang 외 arxiv

Nuclei instance segmentation is critical in computational pathology for cancer diagnosis and prognosis. Recently, the Segment Anything Model has demonstrated exceptional performance in various segmentation tasks, leverag…

parameter-efficient fine-tuningInstance Segmentation

Image Inpainting by End-to-End Cascaded Refinement with Mask Awareness

2021-04-28 · Manyu Zhu, Dongliang He, Xin Li, Chao Li 외

Inpainting arbitrary missing regions is challenging because learning valid features for various masked regions is nontrivial. Though U-shaped encoder-decoder frameworks have been witnessed to be successful, most of them …

DecoderImage Inpaintingvalid