paper-with-me

Papers

Word Searching in Scene Image and Video Frame in Multi-Script Scenario using Dynamic Shape Coding

2017-08-18 · Partha Pratim Roy, Ayan Kumar Bhunia, Avirup Bhattacharyya, Umapada Pal

Retrieval of text information from natural scene images and video frames is a challenging task due to its inherent problems like complex character shapes, low resolution, background noise, etc. Available OCR systems often fail to retrieve such information in scene/video frames. Keyword spotting, an alternative way to retrieve information, performs efficient text searching in such scenarios. However, current word spotting techniques in scene/video images are script-specific and they are mainly developed for Latin script. This paper presents a novel word spotting framework using dynamic shape coding for text retrieval in natural scene image and video frames. The framework is designed to search query keyword from multiple scripts with the help of on-the-fly script-wise keyword generation for the corresponding script. We have used a two-stage word spotting approach using Hidden Markov Model (HMM) to detect the translated keyword in a given text line by identifying the script of the line. A novel unsupervised dynamic shape coding based scheme has been used to group similar shape characters to avoid confusion and to improve text alignment. Next, the hypotheses locations are verified to improve retrieval performance. To evaluate the proposed system for searching keyword from natural scene image and video frames, we have considered two popular Indic scripts such as Bangla (Bengali) and Devanagari along with English. Inspired by the zone-wise recognition approach in Indic scripts[1], zone-wise text information has been used to improve the traditional word spotting performance in Indic scripts. For our experiment, a dataset consisting of images of different scenes and video frames of English, Bangla and Devanagari scripts were considered. The results obtained showed the effectiveness of our proposed word spotting approach.

📄 PDF Abstract BibTeX arXiv:1708.05529

Code (0)

등록된 구현이 없습니다.

Tasks

Keyword SpottingOptical Character Recognition (OCR)RetrievalText Retrieval

Similar Papers 제목 키워드 기반

Date-Field Retrieval in Scene Image and Video Frames using Text Enhancement and Shape Coding

2017-07-21 · Partha Pratim Roy, Ayan Kumar Bhunia, Umapada Pal

Text recognition in scene image and video frames is difficult because of low resolution, blur, background noise, etc. Since traditional OCRs do not perform well in such images, information retrieval using keywords could …

Information RetrievalRetrieval

The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision

2019-04-26 · ICLR 2019 5 · Jiayuan Mao, Chuang Gan, Pushmeet Kohli, Joshua B. Tenenbaum 외

We propose the Neuro-Symbolic Concept Learner (NS-CL), a model that learns visual concepts, words, and semantic parsing of sentences without explicit supervision on any of them; instead, our model learns by simply lookin…

Image-text RetrievalObjectQuestion AnsweringRetrieval+4

Dynamic gesture retrieval: searching videos by human pose sequence

2020-06-13 · Cheng Zhang

The number of static human poses is limited, it is hard to retrieve the exact videos using one single pose as the clue. However, with a pose sequence or a dynamic gesture as the keyword, retrieving specific videos become…

Retrieval

Ensemble Video Object Cut in Highly Dynamic Scenes

2013-06-01 · CVPR 2013 6 · Xiaobo Ren, Tony X. Han, Zhihai He

We consider video object cut as an ensemble of framelevel background-foreground object classifiers which fuses information across frames and refine their segmentation results in a collaborative and iterative manner. Our …

Change DetectionObjectSegmentationSemantic Segmentation

Lyric Video Analysis Using Text Detection and Tracking

2020-06-21 · Shota Sakaguchi, Jun Kato, Masataka Goto, Seiichi Uchida

We attempt to recognize and track lyric words in lyric videos. Lyric video is a music video showing the lyric words of a song. The main characteristic of lyric videos is that the lyric words are shown at frames synchrono…

ClusteringDynamic Time WarpingText DetectionVideo Generation