paper-with-me

홈 › Papers

Spotlight Text Detector: Spotlight on Candidate Regions Like a Camera

2024-09-25 · Xu Han, Junyu Gao, Chuang Yang, Yuan Yuan, Qi Wang

The irregular contour representation is one of the tough challenges in scene text detection. Although segmentation-based methods have achieved significant progress with the help of flexible pixel prediction, the overlap of geographically close texts hinders detecting them separately. To alleviate this problem, some shrink-based methods predict text kernels and expand them to restructure texts. However, the text kernel is an artificial object with incomplete semantic features that are prone to incorrect or missing detection. In addition, different from the general objects, the geometry features (aspect ratio, scale, and shape) of scene texts vary significantly, which makes it difficult to detect them accurately. To consider the above problems, we propose an effective spotlight text detector (STD), which consists of a spotlight calibration module (SCM) and a multivariate information extraction module (MIEM). The former concentrates efforts on the candidate kernel, like a camera focus on the target. It obtains candidate features through a mapping filter and calibrates them precisely to eliminate some false positive samples. The latter designs different shape schemes to explore multiple geometric features for scene texts. It helps extract various spatial relationships to improve the model's ability to recognize kernel regions. Ablation studies prove the effectiveness of the designed SCM and MIEM. Extensive experiments verify that our STD is superior to existing state-of-the-art methods on various datasets, including ICDAR2015, CTW1500, MSRA-TD500, and Total-Text.

📄 PDF Abstract BibTeX arXiv:2409.16820

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Text DetectionText Detection

Methods 이 논문이 사용한 방법론

Focus 설명 없음
STD The Spatial-Channel Token Distillation method is proposed to improve the spatial and channel mixing from a novel knowledge distillation (KD) perspective. To be specific, we…

Similar Papers 제목 키워드 기반

Spotlight-TTS: Spotlighting the Style via Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech

2025-05-27 · Nam-Gyu Kim, Deok-Hyeon Cho, Seung-bin Kim, Seong-Whan Lee

Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We pr…

Style Transfertext-to-speechText to Speech

Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech

2025-11-18 · Nam-Gyu Kim arxiv

Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We pr…

Style Transfer

The Spotlight: A General Method for Discovering Systematic Errors in Deep Learning Models

2021-07-01 · Greg d'Eon, Jason d'Eon, James R. Wright, Kevin Leyton-Brown

Supervised learning models often make systematic errors on rare subsets of the data. When these subsets correspond to explicit labels in the data (e.g., gender, race) such poor performance can be identified straightforwa…

Recommendation Systems

Transcribing Content from Structural Images with Spotlight Mechanism

2019-05-27 · Yu Yin, Zhenya Huang, Enhong Chen, Qi Liu 외

Transcribing content from structural images, e.g., writing notes from music scores, is a challenging task as not only the content objects should be recognized, but the internal structure should also be preserved. Existin…

Introducing Spotlight: A Novel Approach for Generating Captivating Key Information from Documents

2025-09-13 · Ankan Mullick, Sombit Bose, Rounak Saha, Ayan Kumar Bhowmick 외 arxiv

In this paper, we introduce Spotlight, a novel paradigm for information extraction that produces concise, engaging narratives by highlighting the most compelling aspects of a document. Unlike traditional summaries, which…

Information Extraction