paper-with-me

홈 › Papers

STELA: A Real-Time Scene Text Detector with Learned Anchor

2019-09-17 · Linjie Deng, Yanxiang Gong, Xinchen Lu, Yi Lin, Zheng Ma, Mei Xie

To achieve high coverage of target boxes, a normal strategy of conventional one-stage anchor-based detectors is to utilize multiple priors at each spatial position, especially in scene text detection tasks. In this work, we present a simple and intuitive method for multi-oriented text detection where each location of feature maps only associates with one reference box. The idea is inspired from the twostage R-CNN framework that can estimate the location of objects with any shape by using learned proposals. The aim of our method is to integrate this mechanism into a onestage detector and employ the learned anchor which is obtained through a regression operation to replace the original one into the final predictions. Based on RetinaNet, our method achieves competitive performances on several public benchmarks with a totally real-time efficiency (26:5fps at 800p), which surpasses all of anchor-based scene text detectors. In addition, with less attention on anchor design, we believe our method is easy to be applied on other analogous detection tasks. The code will publicly available at https://github.com/xhzdeng/stela.

📄 PDF Abstract BibTeX arXiv:1909.07549

Code (1)

xhzdeng/stela 공식 구현 pytorch

Tasks

Scene Text DetectionText Detection

Methods 이 논문이 사용한 방법론

Focal Loss A Focal Loss function addresses class imbalance during training in tasks like object detection. Focal loss applies a modulating term to the cross entropy loss in order to…
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
FPN 설명 없음
RetinaNet RetinaNet is a one-stage object detection model that utilizes a focal loss function to address class imbalance during training.…
SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…
Max Pooling Max Pooling is a pooling operation that calculates the maximum value for patches of a feature map, and uses it to create a downsampled (pooled) feature map. It is usually…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
R-CNN R-CNN, or Regions with CNN Features, is an object detection model that uses high-capacity CNNs to bottom-up region proposals in order to localize and segment objects. It…

Similar Papers 제목 키워드 기반

A Linguistics-Aware LLM Watermarking via Syntactic Predictability

2025-10-10 · Shinwoo Park, Hyejin Park, Hyeseon An, Yo-Sub Han arxiv

As large language models (LLMs) continue to advance rapidly, reliable governance tools have become critical. Publicly verifiable watermarking is particularly essential for fostering a trustworthy AI ecosystem. A central …

STELAR: Spatio-temporal Tensor Factorization with Latent Epidemiological Regularization

2020-12-08 · Nikos Kargas, Cheng Qian, Nicholas D. Sidiropoulos, Cao Xiao 외

Accurate prediction of the transmission of epidemic diseases such as COVID-19 is crucial for implementing effective mitigation measures. In this work, we develop a tensor method to predict the evolution of epidemic trend…

AttributePrediction

STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision

2025-08-12 · Chen Li, Han Zhang, Zhantao Yang, Fangyi Chen 외 arxiv

Vision-language models (VLMs) have made significant strides in reasoning, yet they often struggle with complex multimodal tasks and tend to generate overly verbose outputs. A key limitation is their reliance on chain-of-…

Reinforcement Learning

Enhancing Scene Text Detectors with Realistic Text Image Synthesis Using Diffusion Models

2023-11-28 · Ling Fu, Zijie Wu, Yingying Zhu, Yuliang Liu 외

Scene text detection techniques have garnered significant attention due to their wide-ranging applications. However, existing methods have a high demand for training data, and obtaining accurate human annotations is labo…

Image GenerationScene Text DetectionText Detection

TextNeRF: A Novel Scene-Text Image Synthesis Method based on Neural Radiance Fields

2024-01-01 · CVPR 2024 1 · Jialei Cui, Jianwei Du, Wenzhuo LIU, Zhouhui Lian

Acquiring large-scale well-annotated datasets is essential for training robust scene text detectors yet the process is often resource-intensive and time-consuming. While some efforts have been made to explore the syn…

Image GenerationNeRFText Detection