paper-with-me

홈 › Papers

Dense Multimodal Fusion for Hierarchically Joint Representation

2018-10-08 · Di Hu, Feiping Nie, Xuelong. Li

Multiple modalities can provide more valuable information than single one by describing the same contents in various ways. Hence, it is highly expected to learn effective joint representation by fusing the features of different modalities. However, previous methods mainly focus on fusing the shallow features or high-level representations generated by unimodal deep networks, which only capture part of the hierarchical correlations across modalities. In this paper, we propose to densely integrate the representations by greedily stacking multiple shared layers between different modality-specific networks, which is named as Dense Multimodal Fusion (DMF). The joint representations in different shared layers can capture the correlations in different levels, and the connection between shared layers also provides an efficient way to learn the dependence among hierarchical correlations. These two properties jointly contribute to the multiple learning paths in DMF, which results in faster convergence, lower training loss, and better performance. We evaluate our model on three typical multimodal learning tasks, including audiovisual speech recognition, cross-modal retrieval, and multimodal classification. The noticeable performance in the experiments demonstrates that our model can learn more effective joint representation.

📄 PDF Abstract BibTeX arXiv:1810.03414

Code (0)

등록된 구현이 없습니다.

Tasks

Cross-Modal RetrievalRetrievalspeech-recognitionSpeech Recognition

Similar Papers 제목 키워드 기반

Improving Multimodal Fusion with Hierarchical Mutual Information Maximization for Multimodal Sentiment Analysis

2021-09-01 · EMNLP 2021 11 · Wei Han, Hui Chen, Soujanya Poria

In multimodal sentiment analysis (MSA), the performance of a model highly depends on the quality of synthesized embeddings. These embeddings are generated from the upstream process called multimodal fusion, which aims to…

Multimodal Sentiment AnalysisSentiment Analysis

A Joint Sequence Fusion Model for Video Question Answering and Retrieval

2018-08-07 · ECCV 2018 9 · Youngjae Yu, Jongseok Kim, Gunhee Kim

We present an approach named JSFusion (Joint Sequence Fusion) that can measure semantic similarity between any pairs of multimodal sequence data (e.g. a video clip and a language sentence). Our multimodal matching networ…

DecoderMultiple-choiceQuestion AnsweringRetrieval+6

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation

2025-08-01 · Le Wang, Jun Wang, Chunyu Qiang, Feng Deng 외 arxiv

We present AudioGen-Omni - a unified approach based on multimodal diffusion transformers (MMDit), capable of generating high-fidelity audio, speech, and song coherently synchronized with the input video. AudioGen-Omni in…

Audio Generation

H-DenseUNet: Hybrid Densely Connected UNet for Liver and Tumor Segmentation from CT Volumes

2017-09-21 · Xiaomeng Li, Hao Chen, Xiaojuan Qi, Qi Dou 외

Liver cancer is one of the leading causes of cancer death. To assist doctors in hepatocellular carcinoma diagnosis and treatment planning, an accurate and automatic liver and tumor segmentation method is highly demanded …

Automatic Liver And Tumor SegmentationGPUImage SegmentationLesion Segmentation+4

H-DenseFormer: An Efficient Hybrid Densely Connected Transformer for Multimodal Tumor Segmentation

2023-07-04 · Jun Shi, Hongyu Kan, Shulan Ruan, Ziqi Zhu 외

Recently, deep learning methods have been widely used for tumor segmentation of multimodal medical images with promising results. However, most existing methods are limited by insufficient representational ability, speci…

Tumor Segmentation