paper-with-me

Papers

SPIN: Structure-Preserving Inner Offset Network for Scene Text Recognition

2020-05-27 · Chengwei Zhang, Yunlu Xu, Zhanzhan Cheng, ShiLiang Pu, Yi Niu, Fei Wu, Futai Zou

Arbitrary text appearance poses a great challenge in scene text recognition tasks. Existing works mostly handle with the problem in consideration of the shape distortion, including perspective distortions, line curvature or other style variations. Therefore, methods based on spatial transformers are extensively studied. However, chromatic difficulties in complex scenes have not been paid much attention on. In this work, we introduce a new learnable geometric-unrelated module, the Structure-Preserving Inner Offset Network (SPIN), which allows the color manipulation of source data within the network. This differentiable module can be inserted before any recognition architecture to ease the downstream tasks, giving neural networks the ability to actively transform input intensity rather than the existing spatial rectification. It can also serve as a complementary module to known spatial transformations and work in both independent and collaborative ways with them. Extensive experiments show that the use of SPIN results in a significant improvement on multiple text recognition benchmarks compared to the state-of-the-arts.

📄 PDF Abstract BibTeX arXiv:2005.13117

Code (3)

hikopensource/davar-lab-ocr 공식 구현 pytorch
PaddlePaddle/PaddleOCR paddle
smilelite/spin_paddle paddle

Tasks

Color ManipulationScene Text Recognition

Similar Papers 제목 키워드 기반

Structured adaptive and random spinners for fast machine learning computations

2016-10-19 · Mariusz Bojarski, Anna Choromanska, Krzysztof Choromanski, Francois Fagan 외

We consider an efficient computational framework for speeding up several machine learning algorithms with almost no loss of accuracy. The proposed framework relies on projections via structured matrices that we call Stru…

BIG-bench Machine LearningDimensionality ReductionQuantization

A Self-Rotating Tri-Rotor UAV for Field of View Expansion and Autonomous Flight

2026-03-30 · Xiaobin Zhou, Zihao Zheng, Aoxu Jin, Lei Qiang 외 arxiv

Unmanned Aerial Vehicles (UAVs) perception relies on onboard sensors like cameras and LiDAR, which are limited by the narrow field of view (FoV). We present Self-Perception INertial Navigation Enabled Rotorcraft (SPINNER…

A Sparsity Inducing Nuclear-Norm Estimator (SpINNEr) for Matrix-Variate Regression in Brain Connectivity Analysis

2020-01-30 · Damian Brzyski, Xixi Hu, Joaquin Goni, Beau Ances 외

Classical scalar-response regression methods treat covariates as a vector and estimate a corresponding vector of regression coefficients. In medical applications, however, regressors are often in a form of multi-dimensio…

Functional Connectivityregression

Identifying Machine-Paraphrased Plagiarism

2021-03-22 · Jan Philip Wahle, Terry Ruas, Tomáš Foltýnek, Norman Meuschke 외

Employing paraphrasing tools to conceal plagiarized text is a severe threat to academic integrity. To enable the detection of machine-paraphrased text, we evaluate the effectiveness of five pre-trained word embedding mod…

ArticlesText Matching

End-to-End Dexterous Grasp Learning from Single-View Point Clouds via a Multi-Object Scene Dataset

2026-03-16 · Tao Geng, Dapeng Yang, Ziwei Liu, Le Zhang 외 arxiv

Dexterous grasping in multi-object scene constitutes a fundamental challenge in robotic manipulation. Current mainstream grasping datasets predominantly focus on single-object scenarios and predefined grasp configuration…

Point Clouds