Structure Optimization for Deep Multimodal Fusion Networks using Graph-Induced Kernels
A popular testbed for deep learning has been multimodal recognition of human activity or gesture involving diverse inputs such as video, audio, skeletal pose and depth images. Deep learning architectures have excelled on such problems due to their ability to combine modality representations at different levels of nonlinear feature extraction. However, designing an optimal architecture in which to fuse such learned representations has largely been a non-trivial human engineering effort. We treat fusion structure optimization as a hyper-parameter search and cast it as a discrete optimization problem under the Bayesian optimization framework. We propose a novel graph-induced kernel to compute structural similarities in the search space of tree-structured multimodal architectures and demonstrate its effectiveness using two challenging multimodal human activity recognition datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Activity RecognitionBayesian OptimizationDeep LearningHuman Activity RecognitionSimilar Papers 제목 키워드 기반
SMGFM: Spectral Multimodal Graph Pretraining for Multimodal-Attributed Graphs
Multimodal-attributed graphs (MAGs) couple graph topology with node semantics from text, images, and other modalities. Traditional graph learning contextualizes node semantics by coupling topology with node features. How…
Graph LearningModality-Guided Mixture of Graph Experts with Entropy-Triggered Routing for Multimodal Recommendation
Multimodal recommendation enhances ranking by integrating user-item interactions with item content, which is particularly effective under sparse feedback and long-tail distributions. However, multimodal signals are inher…
Multimodal RecommendationGraph LearningMultimodal Prediction based on Graph Representations
This paper proposes a learning model, based on rank-fusion graphs, for general applicability in multimodal prediction tasks, such as multimodal regression and image classification. Rank-fusion graphs encode information f…
image-classificationImage ClassificationPredictionRetrievalOptimization-Induced Graph Implicit Nonlinear Diffusion
Due to the over-smoothing issue, most existing graph neural networks can only capture limited dependencies with their inherently finite aggregation layers. To overcome this limitation, we propose a new kind of graph conv…
Multimodal Transformers are Hierarchical Modal-wise Heterogeneous Graphs
Multimodal Sentiment Analysis (MSA) is a rapidly developing field that integrates multimodal information to recognize sentiments, and existing models have made significant progress in this area. The central challenge in …
Multimodal Sentiment AnalysisSentiment Analysis