paper-with-me

Papers

ILCAS: Imitation Learning-Based Configuration-Adaptive Streaming for Live Video Analytics with Cross-Camera Collaboration

2023-08-19 · Duo Wu, Dayou Zhang, Miao Zhang, Ruoyu Zhang, Fangxin Wang, Shuguang Cui

The high-accuracy and resource-intensive deep neural networks (DNNs) have been widely adopted by live video analytics (VA), where camera videos are streamed over the network to resource-rich edge/cloud servers for DNN inference. Common video encoding configurations (e.g., resolution and frame rate) have been identified with significant impacts on striking the balance between bandwidth consumption and inference accuracy and therefore their adaption scheme has been a focus of optimization. However, previous profiling-based solutions suffer from high profiling cost, while existing deep reinforcement learning (DRL) based solutions may achieve poor performance due to the usage of fixed reward function for training the agent, which fails to craft the application goals in various scenarios. In this paper, we propose ILCAS, the first imitation learning (IL) based configuration-adaptive VA streaming system. Unlike DRL-based solutions, ILCAS trains the agent with demonstrations collected from the expert which is designed as an offline optimal policy that solves the configuration adaption problem through dynamic programming. To tackle the challenge of video content dynamics, ILCAS derives motion feature maps based on motion vectors which allow ILCAS to visually ``perceive'' video content changes. Moreover, ILCAS incorporates a cross-camera collaboration scheme to exploit the spatio-temporal correlations of cameras for more proper configuration selection. Extensive experiments confirm the superiority of ILCAS compared with state-of-the-art solutions, with 2-20.9% improvement of mean accuracy and 19.9-85.3% reduction of chunk upload lag.

📄 PDF Abstract BibTeX arXiv:2308.10068

Code (0)

등록된 구현이 없습니다.

Tasks

Deep Reinforcement LearningImitation Learning

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

Open GOP Resolution Switching in HTTP Adaptive Streaming with VVC

2021-03-11 · Robert Skupin, Christian Bartnik, Adam Wieckowski, Yago Sanchez 외

The user experience in adaptive HTTP streaming relies on offering bitrate ladders with suitable operation points for all users and typically involves multiple resolutions. While open GOP coding structures are generally k…

Lookahead When It Matters: Adaptive Non-causal Transformers for Streaming Neural Transducers

2023-05-07 · Grant P. Strimel, Yi Xie, Brian King, Martin Radfar 외

Streaming speech recognition architectures are employed for low-latency, real-time applications. Such architectures are often characterized by their causality. Causal architectures emit tokens at each frame, relying only…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

TuningIQA: Fine-Grained Blind Image Quality Assessment for Livestreaming Camera Tuning

2025-08-25 · Xiangfei Sheng, Zhichao Duan, Xiaofeng Pan, Yipo Huang 외 arxiv

Livestreaming has become increasingly prevalent in modern visual communication, where automatic camera quality tuning is essential for delivering superior user Quality of Experience (QoE). Such tuning requires accurate b…

Image Quality Assessment

LiveStar: Live Streaming Assistant for Real-World Online Video Understanding

2025-11-07 · Zhenyu Yang, Kairui Zhang, Yuhang Hu, Bing Wang 외 arxiv

Despite significant progress in Video Large Language Models (Video-LLMs) for offline video understanding, existing online Video-LLMs typically struggle to simultaneously process continuous frame-by-frame inputs and deter…

Facial Attractiveness Prediction in Live Streaming: A New Benchmark and Multi-modal Method

2025-01-05 · Hui Li, Xiaoyu Ren, Hongjiu Yu, Huiyu Duan 외

Facial attractiveness prediction (FAP) has long been an important computer vision task, which could be widely applied in live streaming for facial retouching, content recommendation, etc. However, previous FAP datasets a…

Diversity