paper-with-me

홈 › Papers

Impression Network for Video Object Detection

2017-12-16 · Congrui Hetang, Hongwei Qin, Shaohui Liu, Junjie Yan

Video object detection is more challenging compared to image object detection. Previous works proved that applying object detector frame by frame is not only slow but also inaccurate. Visual clues get weakened by defocus and motion blur, causing failure on corresponding frames. Multi-frame feature fusion methods proved effective in improving the accuracy, but they dramatically sacrifice the speed. Feature propagation based methods proved effective in improving the speed, but they sacrifice the accuracy. So is it possible to improve speed and performance simultaneously? Inspired by how human utilize impression to recognize objects from blurry frames, we propose Impression Network that embodies a natural and efficient feature aggregation mechanism. In our framework, an impression feature is established by iteratively absorbing sparsely extracted frame features. The impression feature is propagated all the way down the video, helping enhance features of low-quality frames. This impression mechanism makes it possible to perform long-range multi-frame feature fusion among sparse keyframes with minimal overhead. It significantly improves per-frame detection baseline on ImageNet VID while being 3 times faster (20 fps). We hope Impression Network can provide a new perspective on video feature enhancement. Code will be made available.

📄 PDF Abstract BibTeX arXiv:1712.05896

Code (0)

등록된 구현이 없습니다.

Tasks

Objectobject-detectionObject DetectionVideo Object Detection

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Deep Impression: Audiovisual Deep Residual Networks for Multimodal Apparent Personality Trait Recognition

2016-09-16 · Yağmur Güçlütürk, Umut Güçlü, Marcel A. J. van Gerven, Rob Van Lier

Here, we develop an audiovisual deep residual network for multimodal apparent personality trait recognition. The network is trained end-to-end for predicting the Big Five personality traits of people from their videos. T…

Facial Expression RecognitionFeature EngineeringPersonality Trait Recognition

Evaluation of GPT-4 for chest X-ray impression generation: A reader study on performance and perception

2023-11-12 · Sebastian Ziegelmayer, Alexander W. Marka, Nicolas Lenhart, Nadja Nehls 외

The remarkable generative capabilities of multimodal foundation models are currently being explored for a variety of applications. Generating radiological impressions is a challenging task that could significantly reduce…

RLTP: Reinforcement Learning to Pace for Delayed Impression Modeling in Preloaded Ads

2023-02-06 · Penghui Wei, Yongqiang Chen, Shaoguo Liu, Liang Wang 외

To increase brand awareness, many advertisers conclude contracts with advertising platforms to purchase traffic and then deliver advertisements to target audiences. In a whole delivery period, advertisers usually desire …

reinforcement-learningReinforcement Learning (RL)

Voice Impression Control in Zero-Shot TTS

2025-06-06 · Keinichi Fujita, Shota Horiguchi, Yusuke Ijima

Para-/non-linguistic information in speech is pivotal in shaping the listeners' impression. Although zero-shot text-to-speech (TTS) has achieved high speaker fidelity, modulating subtle para-/non-linguistic information t…

Language ModelingLanguage ModellingLarge Language Modeltext-to-speech+1

Enhancing Traffic Scene Predictions with Generative Adversarial Networks

2019-09-24 · Peter König, Sandra Aigner, Marco Körner

We present a new two-stage pipeline for predicting frames of traffic scenes where relevant objects can still reliably be detected. Using a recent video prediction network, we first generate a sequence of future frames ba…

DeblurringImage Super-ResolutionImage-to-Image Translationobject-detection+6