paper-with-me

홈 › Papers

DEEPEYE: A Compact and Accurate Video Comprehension at Terminal Devices Compressed with Quantization and Tensorization

2018-05-21 · Yuan Cheng, Guangya Li, Hai-Bao Chen, Sheldon X. -D. Tan, Hao Yu

As it requires a huge number of parameters when exposed to high dimensional inputs in video detection and classification, there is a grand challenge to develop a compact yet accurate video comprehension at terminal devices. Current works focus on optimizations of video detection and classification in a separated fashion. In this paper, we introduce a video comprehension (object detection and action recognition) system for terminal devices, namely DEEPEYE. Based on You Only Look Once (YOLO), we have developed an 8-bit quantization method when training YOLO; and also developed a tensorized-compression method of Recurrent Neural Network (RNN) composed of features extracted from YOLO. The developed quantization and tensorization can significantly compress the original network model yet with maintained accuracy. Using the challenging video datasets: MOMENTS and UCF11 as benchmarks, the results show that the proposed DEEPEYE achieves 3.994x model compression rate with only 0.47% mAP decreased; and 15,047x parameter reduction and 2.87x speed-up with 16.58% accuracy improvement.

📄 PDF Abstract BibTeX arXiv:1805.07935

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionGeneral ClassificationModel Compressionobject-detectionObject DetectionQuantizationTemporal Action Localization

Similar Papers 제목 키워드 기반

DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

2025-05-20 · Ziwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao 외

Large Vision-Language Models (VLMs) have shown strong capabilities in multimodal understanding and reasoning, yet they are primarily constrained by text-based reasoning processes. However, achieving seamless integration …

HallucinationMathematical ReasoningMultimodal Reasoningreinforcement-learning+2

DeepEyeNet: Adaptive Genetic Bayesian Algorithm Based Hybrid ConvNeXtTiny Framework For Multi-Feature Glaucoma Eye Diagnosis

2025-01-19 · Angshuman Roy, Anuvab Sen, Soumyajit Gupta, Soham Haldar 외

Glaucoma is a leading cause of irreversible blindness worldwide, emphasizing the critical need for early detection and intervention. In this paper, we present DeepEyeNet, a novel and comprehensive framework for automated…

Bayesian Optimization

DeepEyesV2: Toward Agentic Multimodal Model

2025-11-07 · Jack Hong, Chenxiao Zhao, ChengLin Zhu, Weiheng Lu 외 arxiv

Agentic multimodal models should not only comprehend text and images, but also actively invoke external tools, such as code execution environments and web search, and integrate these operations into reasoning. In this wo…

Reinforcement LearningMathematical ReasoningMultimodal Reasoning

Differentiable Grammars for Videos

2019-02-01 · AJ Piergiovanni, Anelia Angelova, Michael S. Ryoo

This paper proposes a novel algorithm which learns a formal regular grammar from real-world continuous data, such as videos. Learning latent terminals, non-terminals, and production rules directly from continuous data al…

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

2025-07-28 · Yuying Ge, Yixiao Ge, Chen Li, Teng Wang 외 arxiv

Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal models lack essential temporally-struct…

Video Question AnsweringReinforcement LearningVideo CaptioningVideo Grounding