paper-with-me

Papers

Towards Modality Transferable Visual Information Representation with Optimal Model Compression

2020-08-13 · Rongqun Lin, Linwei Zhu, Shiqi Wang, Sam Kwong

Compactly representing the visual signals is of fundamental importance in various image/video-centered applications. Although numerous approaches were developed for improving the image and video coding performance by removing the redundancies within visual signals, much less work has been dedicated to the transformation of the visual signals to another well-established modality for better representation capability. In this paper, we propose a new scheme for visual signal representation that leverages the philosophy of transferable modality. In particular, the deep learning model, which characterizes and absorbs the statistics of the input scene with online training, could be efficiently represented in the sense of rate-utility optimization to serve as the enhancement layer in the bitstream. As such, the overall performance can be further guaranteed by optimizing the new modality incorporated. The proposed framework is implemented on the state-of-the-art video coding standard (i.e., versatile video coding), and significantly better representation capability has been observed based on extensive evaluations.

📄 PDF Abstract BibTeX arXiv:2008.05642

Code (0)

등록된 구현이 없습니다.

Tasks

Model CompressionPhilosophy

Similar Papers 제목 키워드 기반

Modality-Transferable Emotion Embeddings for Low-Resource Multimodal Emotion Recognition

2020-09-21 · Asian Chapter of the Association for Computational Linguistics 2020 · Wenliang Dai, Zihan Liu, Tiezheng Yu, Pascale Fung

Despite the recent achievements made in the multi-modal emotion recognition task, two problems still exist and have not been well investigated: 1) the relationship between different emotion categories are not utilized, w…

Emotion RecognitionMultimodal Emotion RecognitionWord Embeddings

Dual Modality Prompt Tuning for Vision-Language Pre-Trained Model

2022-08-17 · Yinghui Xing, Qirui Wu, De Cheng, Shizhou Zhang 외

With the emergence of large pre-trained vison-language model like CLIP, transferable representations can be adapted to a wide range of downstream tasks via prompt tuning. Prompt tuning tries to probe the beneficial infor…

General KnowledgeLanguage ModellingVisual Prompt Tuning

VT-CLIP: Enhancing Vision-Language Models with Visual-guided Texts

2021-12-04 · Longtian Qiu, Renrui Zhang, Ziyu Guo, Ziyao Zeng 외

Contrastive Language-Image Pre-training (CLIP) has drawn increasing attention recently for its transferable visual representation learning. However, due to the semantic gap within datasets, CLIP's pre-trained image-text …

Language ModellingRepresentation LearningZero-Shot Learning

Unsupervised Modality-Transferable Video Highlight Detection with Representation Activation Sequence Learning

2024-03-14 · Tingtian Li, Zixun Sun, Xinyu Xiao

Identifying highlight moments of raw video materials is crucial for improving the efficiency of editing videos that are pervasive on internet platforms. However, the extensive work of manually labeling footage has create…

Contrastive LearningHighlight Detection

Leveraging Modality-specific Representations for Audio-visual Speech Recognition via Reinforcement Learning

2022-12-10 · Chen Chen, Yuchen Hu, Qiang Zhang, Heqing Zou 외

Audio-visual speech recognition (AVSR) has gained remarkable success for ameliorating the noise-robustness of speech recognition. Mainstream methods focus on fusing audio and visual inputs to obtain modality-invariant re…

Audio-Visual Speech Recognitionreinforcement-learningReinforcement Learning (RL)speech-recognition+2