paper-with-me

Papers

Cross-modal Variational Auto-encoder for Content-based Micro-video Background Music Recommendation

2021-07-15 · Jing Yi, Yaochen Zhu, Jiayi Xie, Zhenzhong Chen

In this paper, we propose a cross-modal variational auto-encoder (CMVAE) for content-based micro-video background music recommendation. CMVAE is a hierarchical Bayesian generative model that matches relevant background music to a micro-video by projecting these two multimodal inputs into a shared low-dimensional latent space, where the alignment of two corresponding embeddings of a matched video-music pair is achieved by cross-generation. Moreover, the multimodal information is fused by the product-of-experts (PoE) principle, where the semantic information in visual and textual modalities of the micro-video are weighted according to their variance estimations such that the modality with a lower noise level is given more weights. Therefore, the micro-video latent variables contain less irrelevant information that results in a more robust model generalization. Furthermore, we establish a large-scale content-based micro-video background music recommendation dataset, TT-150k, composed of approximately 3,000 different background music clips associated to 150,000 micro-videos from different users. Extensive experiments on the established TT-150k dataset demonstrate the effectiveness of the proposed method. A qualitative assessment of CMVAE by visualizing some recommendation results is also included.

📄 PDF Abstract BibTeX arXiv:2107.07268

Code (0)

등록된 구현이 없습니다.

Tasks

Music Recommendation

Similar Papers 제목 키워드 기반

Cross-modal Variational Auto-encoder with Distributed Latent Spaces and Associators

2019-05-30 · Dae Ung Jo, ByeongJu Lee, Jongwon Choi, Haanju Yoo 외

In this paper, we propose a novel structure for a cross-modal data association, which is inspired by the recent research on the associative learning structure of the brain. We formulate the cross-modal association in Bay…

Bayesian Inference

Multimodal Transformer for Parallel Concatenated Variational Autoencoders

2022-10-28 · Stephen D. Liang, Jerry M. Mendel

In this paper, we propose a multimodal transformer using parallel concatenated architecture. Instead of using patches, we use column stripes for images in R, G, B channels as the transformer input. The column stripes kee…

Decoder

Disentangling Latent Hands for Image Synthesis and Pose Estimation

2018-12-03 · CVPR 2019 6 · Linlin Yang, Angela Yao

Hand image synthesis and pose estimation from RGB images are both highly challenging tasks due to the large discrepancy between factors of variation ranging from image background content to camera viewpoint. To better an…

Image GenerationPose Estimation

StyleMeUp: Towards Style-Agnostic Sketch-Based Image Retrieval

2021-03-29 · CVPR 2021 1 · Aneeshan Sain, Ayan Kumar Bhunia, Yongxin Yang, Tao Xiang 외

Sketch-based image retrieval (SBIR) is a cross-modal matching problem which is typically solved by learning a joint embedding space where the semantic content shared between photo and sketch modalities are preserved. How…

DisentanglementImage RetrievalMeta-LearningRetrieval+1

Translating Visual Art into Music

2019-09-03 · Maximilian Müller-Eberstein, Nanne van Noord

The Synesthetic Variational Autoencoder (SynVAE) introduced in this research is able to learn a consistent mapping between visual and auditive sensory modalities in the absence of paired datasets. A quantitative evaluati…

Translation