paper-with-me

Papers

Transcribing Content from Structural Images with Spotlight Mechanism

2019-05-27 · Yu Yin, Zhenya Huang, Enhong Chen, Qi Liu, Fuzheng Zhang, Xing Xie, Guoping Hu

Transcribing content from structural images, e.g., writing notes from music scores, is a challenging task as not only the content objects should be recognized, but the internal structure should also be preserved. Existing image recognition methods mainly work on images with simple content (e.g., text lines with characters), but are not capable to identify ones with more complex content (e.g., structured symbols), which often follow a fine-grained grammar. To this end, in this paper, we propose a hierarchical Spotlight Transcribing Network (STN) framework followed by a two-stage "where-to-what" solution. Specifically, we first decide "where-to-look" through a novel spotlight mechanism to focus on different areas of the original image following its structure. Then, we decide "what-to-write" by developing a GRU based network with the spotlight areas for transcribing the content accordingly. Moreover, we propose two implementations on the basis of STN, i.e., STNM and STNR, where the spotlight movement follows the Markov property and Recurrent modeling, respectively. We also design a reinforcement method to refine the framework by self-improving the spotlight mechanism. We conduct extensive experiments on many structural image datasets, where the results clearly demonstrate the effectiveness of STN framework.

📄 PDF Abstract BibTeX arXiv:1905.10954

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

GRU A Gated Recurrent Unit, or GRU, is a type of recurrent neural network. It is similar to an LSTM, but only has two gates - a reset…

Similar Papers 제목 키워드 기반

AVR: Synergizing Foundation Models for Audio-Visual Humor Detection

2024-06-15 · Sarthak Sharma, Orchid Chetia Phukan, Drishti Singh, Arun Balaji Buduru 외

In this work, we present, AVR application for audio-visual humor detection. While humor detection has traditionally centered around textual analysis, recent advancements have spotlighted multimodal approaches. However, t…

Humor Detection

Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents

2025-09-13 · Ankan Mullick, Sombit Bose, Rounak Saha, Ayan Kumar Bhowmick 외 arxiv

In this paper, we introduce Spotlight, a novel paradigm for information extraction that produces concise, engaging narratives by highlighting the most compelling aspects of a document. Unlike traditional summaries, which…

Information Extraction

Active Light Modulation to Counter Manipulation of Speech Visual Content

2025-04-30 · Hadleigh Schwartz, Xiaofeng Yan, Charles J. Carver, Xia Zhou

High-profile speech videos are prime targets for falsification, owing to their accessibility and influence. This work proposes Spotlight, a low-overhead and unobtrusive system for protecting live speech videos from visua…

Spotlighter: Revisiting Prompt Tuning from a Representative Mining View

2025-08-31 · Yutong Gao, Maoyuan Shao, Xinyang Huang, Chuang Zhu 외 arxiv

CLIP's success has demonstrated that prompt tuning can achieve robust cross-modal semantic alignment for tasks ranging from open-domain recognition to fine-grained classification. However, redundant or weakly relevant fe…

Field typing for improved recognition on heterogeneous handwritten forms

2019-09-23 · Ciprian Tomoiaga, Paul Feng, Mathieu Salzmann, Patrick Jayet

Offline handwriting recognition has undergone continuous progress over the past decades. However, existing methods are typically benchmarked on free-form text datasets that are biased towards good-quality images and hand…

Handwriting Recognition