paper-with-me

Papers

You Only Recognize Once: Towards Fast Video Text Spotting

2019-03-08 · Zhanzhan Cheng, Jing Lu, Yi Niu, ShiLiang Pu, Fei Wu, Shuigeng Zhou

Video text spotting is still an important research topic due to its various real-applications. Previous approaches usually fall into the four-staged pipeline: text detection in individual images, framewisely recognizing localized text regions, tracking text streams and generating final results with complicated post-processing skills, which might suffer from the huge computational cost as well as the interferences of low-quality text. In this paper, we propose a fast and robust video text spotting framework by only recognizing the localized text one-time instead of frame-wisely recognition. Specifically, we first obtain text regions in videos with a well-designed spatial-temporal detector. Then we concentrate on developing a novel text recommender for selecting the highest-quality text from text streams and only recognizing the selected ones. Here, the recommender assembles text tracking, quality scoring and recognition into an end-to-end trainable module, which not only avoids the interferences from low-quality text but also dramatically speeds up the video text spotting process. In addition, we collect a larger scale video text dataset (LSVTD) for promoting the video text spotting community, which contains 100 text videos from 22 different real-life scenarios. Extensive experiments on two public benchmarks show that our method greatly speeds up the recognition process averagely by 71 times compared with the frame-wise manner, and also achieves the remarkable state-of-the-art.

📄 PDF Abstract BibTeX arXiv:1903.03299

Code (1)

hikopensource/davar-lab-ocr 공식 구현 pytorch

Tasks

Text DetectionText Spotting

Similar Papers 제목 키워드 기반

Learning Order Parameters from Videos of Dynamical Phases for Skyrmions with Neural Networks

2020-12-02 · Weidi Wang, Zeyuan Wang, Yinghui Zhang, Bo Sun 외

The ability to recognize dynamical phenomena (e.g., dynamical phases) and dynamical processes in physical events from videos, then to abstract physical concepts and reveal physical laws, lies at the core of human intelli…

TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection

2026-08-18 · Yearang Lee, Ho-Joong Kim, Seong-Whan Lee arxiv

Zero-Shot Temporal Action Detection (ZSTAD) aims to lo- calize and recognize action instances from unseen action categories in untrimmed videos. Although existing meth- ods have shown effectiveness by advancing architect…

Action DetectionVideo Alignment

Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters

2025-09-23 · Pin-Yen Chiu, I-Sheng Fang, Jun-Cheng Chen arxiv

Recent advances in diffusion models have significantly improved image and video synthesis. In addition, several concept control methods have been proposed to enable fine-grained, continuous, and flexible control over fre…

Continuous Control

An Internal Clock Based Space-time Neural Network for Motion Speed Recognition

2020-01-28 · Junwen Luo, Jiaoyan Chen

In this work we present a novel internal clock based space-time neural network for motion speed recognition. The developed system has a spike train encoder, a Spiking Neural Network (SNN) with internal clocking behaviors…

SKiT: a Fast Key Information Video Transformer for Online Surgical Phase Recognition

2023-01-01 · ICCV 2023 1 · Yang Liu, Jiayu Huo, Jingjing Peng, Rachel Sparks 외

This paper introduces SKiT, a fast Key information Transformer for phase recognition of videos. Unlike previous methods that rely on complex models to capture long-term temporal information, SKiT accurately recognize…

Online surgical phase recognitionSurgical phase recognition