paper-with-me

홈 › Papers

Runner-Up Solution to ECCV 2022 Challenge on Out of Vocabulary Scene Text Understanding: Cropped Word Recognition

2022-08-04 · Zhangzi Zhu, Yu Hao, Wenqing Zhang, Chuhui Xue, Song Bai

This report presents our 2nd place solution to ECCV 2022 challenge on Out-of-Vocabulary Scene Text Understanding (OOV-ST) : Cropped Word Recognition. This challenge is held in the context of ECCV 2022 workshop on Text in Everything (TiE), which aims to extract out-of-vocabulary words from natural scene images. In the competition, we first pre-train SCATTER on the synthetic datasets, then fine-tune the model on the training set with data augmentations. Meanwhile, two additional models are trained specifically for long and vertical texts. Finally, we combine the output from different models with different layers, different backbones, and different seeds as the final results. Our solution achieves a word accuracy of 59.45\% when considering out-of-vocabulary words only.

📄 PDF Abstract BibTeX arXiv:2208.02747

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

1st Place Solution to ECCV 2022 Challenge on Out of Vocabulary Scene Text Understanding: End-to-End Recognition of Out of Vocabulary Words

2022-09-01 · Zhangzi Zhu, Chuhui Xue, Yu Hao, Wenqing Zhang 외

Scene text recognition has attracted increasing interest in recent years due to its wide range of applications in multilingual translation, autonomous driving, etc. In this report, we describe our solution to the Out of …

Autonomous DrivingScene Text RecognitionTranslation

The Runner-up Solution for YouTube-VIS Long Video Challenge 2022

2022-11-18 · Junfeng Wu, Yi Jiang, Qihao Liu, Xiang Bai 외

This technical report describes our 2nd-place solution for the ECCV 2022 YouTube-VIS Long Video Challenge. We adopt the previously proposed online video instance segmentation method IDOL for this challenge. In addition, …

Contrastive LearningInstance SegmentationSemantic SegmentationVideo Instance Segmentation

DreamRunner: Fine-Grained Storytelling Video Generation with Retrieval-Augmented Motion Adaptation

2024-11-25 · Zun Wang, Jialu Li, Han Lin, Jaehong Yoon 외

Storytelling video generation (SVG) has recently emerged as a task to create long, multi-motion, multi-scene videos that consistently represent the story described in the input text script. SVG holds great potential for …

Large Language ModelMotion PlanningRetrievalTest-time Adaptation+2

ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving

2025-07-02 · Kai Chen, Ruiyuan Gao, Lanqing Hong, Hang Xu 외 arxiv

In this paper, we present details of the 1st W-CODA workshop, held in conjunction with the ECCV 2024. W-CODA aims to explore next-generation solutions for autonomous driving corner cases, empowered by state-of-the-art mu…

Scene UnderstandingAutonomous Driving

First Place Solution to the ECCV 2024 ROAD++ Challenge @ ROAD++ Spatiotemporal Agent Detection 2024

2024-10-30 · Tengfei Zhang, Heng Zhang, Ruyang Li, Qi Deng 외

This report presents our team's solutions for the Track 1 of the 2024 ECCV ROAD++ Challenge. The task of Track 1 is spatiotemporal agent detection, which aims to construct an "agent tube" for road agents in consecutive v…

Data Augmentationobject-detectionObject Detection