paper-with-me

홈 › Papers

COEF-VQ: Cost-Efficient Video Quality Understanding through a Cascaded Multimodal LLM Framework

2024-12-11 · Xin Dong, Sen Jia, Hongyu Xiong

Recently, with the emergence of recent Multimodal Large Language Model (MLLM) technology, it has become possible to exploit its video understanding capability on different classification tasks. In practice, we face the difficulty of huge requirements for GPU resource if we need to deploy MLLMs online. In this paper, we propose COEF-VQ, a novel cascaded MLLM framework for better video quality understanding on TikTok. To this end, we first propose a MLLM fusing all visual, textual and audio signals, and then develop a cascade framework with a lightweight model as pre-filtering stage and MLLM as fine-consideration stage, significantly reducing the need for GPU resource, while retaining the performance demonstrated solely by MLLM. To demonstrate the effectiveness of COEF-VQ, we deployed this new framework onto the video management platform (VMP) at TikTok, and performed a series of detailed experiments on two in-house tasks related to video quality understanding. We show that COEF-VQ leads to substantial performance gains with limit resource consumption in these two tasks.

📄 PDF Abstract BibTeX arXiv:2412.10435

Code (0)

등록된 구현이 없습니다.

Tasks

GPULanguage ModelingLanguage ModellingLarge Language ModelMultimodal Large Language ModelVideo Understanding

Similar Papers 제목 키워드 기반

Video Quality Enhancement Using Deep Learning-Based Prediction Models for Quantized DCT Coefficients in MPEG I-frames

2020-10-09 · Antonio J G Busson, Paulo R C Mendes, Daniel de S Moraes, Álvaro M da Veiga 외

Recent works have successfully applied some types of Convolutional Neural Networks (CNNs) to reduce the noticeable distortion resulting from the lossy JPEG/MPEG compression technique. Most of them are built upon the proc…

Decoder

SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation

2022-11-22 · CVPR 2023 1 · Wenxuan Zhang, Xiaodong Cun, Xuan Wang, Yong Zhang 외

Generating talking head videos through a face image and a piece of speech audio still contains many challenges. ie, unnatural head movement, distorted expression, and identity modification. We argue that these issues are…

Image AnimationTalking Head Generation

Building Scalable Video Understanding Benchmarks through Sports

2023-01-17 · Aniket Agarwal, Alex Zhang, Karthik Narasimhan, Igor Gilitschenski 외

Existing benchmarks for evaluating long video understanding falls short on two critical aspects, either lacking in scale or quality of annotations. These limitations arise from the difficulty in collecting dense annotati…

Video Understanding

ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

2024-06-06 · Lin Chen, Xilin Wei, Jinsong Li, Xiaoyi Dong 외

We present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs) via dense and precise captions. The serie…

Video CaptioningVideo GenerationVideo UnderstandingWorld Knowledge

Reduced Reference Perceptual Quality Model and Application to Rate Control for 3D Point Cloud Compression

2020-11-25 · Qi Liu, Hui Yuan, Raouf Hamzaoui, Honglei Su 외

In rate-distortion optimization, the encoder settings are determined by maximizing a reconstruction quality measure subject to a constraint on the bit rate. One of the main challenges of this approach is to define a qual…

Quantization