paper-with-me

홈 › Papers

A Perspective on Deep Vision Performance with Standard Image and Video Codecs

2024-04-18 · Christoph Reich, Oliver Hahn, Daniel Cremers, Stefan Roth, Biplob Debnath

Resource-constrained hardware, such as edge devices or cell phones, often rely on cloud servers to provide the required computational resources for inference in deep vision models. However, transferring image and video data from an edge or mobile device to a cloud server requires coding to deal with network constraints. The use of standardized codecs, such as JPEG or H.264, is prevalent and required to ensure interoperability. This paper aims to examine the implications of employing standardized codecs within deep vision pipelines. We find that using JPEG and H.264 coding significantly deteriorates the accuracy across a broad range of vision tasks and models. For instance, strong compression rates reduce semantic segmentation accuracy by more than 80% in mIoU. In contrast to previous findings, our analysis extends beyond image and action classification to localization and dense prediction tasks, thus providing a more comprehensive perspective.

📄 PDF Abstract BibTeX arXiv:2404.12330

Code (0)

등록된 구현이 없습니다.

Tasks

Image ClassificationSemantic Segmentation

Similar Papers 제목 키워드 기반

Robust Re-Identification by Multiple Views Knowledge Distillation

2020-07-08 · ECCV 2020 8 · Angelo Porrello, Luca Bergamini, Simone Calderara

To achieve robustness in Re-Identification, standard methods leverage tracking information in a Video-To-Video fashion. However, these solutions face a large drop in performance for single image queries (e.g., Image-To-V…

Knowledge DistillationPerson Re-IdentificationVehicle Re-IdentificationVideo-Based Person Re-Identification

Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent Analytics

2020-01-10 · Ling-Yu Duan, Jiaying Liu, Wenhan Yang, Tiejun Huang 외

Video coding, which targets to compress and reconstruct the whole frame, and feature compression, which only preserves and transmits the most critical information, stand at two ends of the scale. That is, one is with com…

Feature CompressionVideo Compression

Beyond the Frame: Generating 360° Panoramic Videos from Perspective Videos

2025-04-10 · Rundong Luo, Matthew Wallingford, Ali Farhadi, Noah Snavely 외

360{\deg} videos have emerged as a promising medium to represent our dynamic visual world. Compared to the "tunnel vision" of standard cameras, their borderless field of view offers a more complete perspective of our sur…

Question AnsweringVideo GenerationVideo StabilizationVisual Question Answering

IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models

2026-04-01 · Dong-Jae Lee, Sunghyun Baek, Junmo Kim arxiv

Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rapidly with the number of visual tokens. Existing token pruning methods mitigate th…

Deep Video Codec Control for Vision Models

2023-08-30 · Christoph Reich, Biplob Debnath, Deep Patel, Tim Prangemeier 외

Standardized lossy video coding is at the core of almost all real-world video processing pipelines. Rate control is used to enable standard codecs to adapt to different network bandwidth conditions or storage constraints…

Optical Flow EstimationSemantic SegmentationVideo Compression