paper-with-me

홈 › Papers

Hierarchical Multimodal Metric Learning for Multimodal Classification

2017-07-01 · CVPR 2017 7 · Heng Zhang, Vishal M. Patel, Rama Chellappa

Multimodal classification arises in many computer vision tasks such as object classification and image retrieval. The idea is to utilize multiple sources (modalities) measuring the same instance to improve the overall performance compared to using a single source (modality). The varying characteristics exhibited by multiple modalities make it necessary to simultaneously learn the corresponding metrics. In this paper, we propose a multiple metrics learning algorithm for multimodal data. Metric of each modality is a product of two matrices: one matrix is modality specific, the other is enforced to be shared by all the modalities. The learned metrics can improve multimodal classification accuracy and experimental results on four datasets show that the proposed algorithm outperforms existing learning algorithms based on multiple metrics as well as other approaches tested on these datasets. Specifically, we report 95.0% object instance recognition accuracy, 89.2% object category recognition accuracy on the multi-view RGB-D dataset and 52.3% scene category recognition accuracy on SUN RGB-D dataset.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationImage RetrievalMetric LearningObjectRetrieval

Similar Papers 제목 키워드 기반

GLEAM: A Multimodal Imaging Dataset and HAMM for Glaucoma Classification

2026-03-13 · Jiao Wang, Chi Liu, Yiying Zhang, Hongchen Luo 외 arxiv

We propose glaucoma lesion evaluation and analysis with multimodal imaging (GLEAM), the first publicly available tri-modal glaucoma dataset comprising scanning laser ophthalmoscopy fundus images, circumpapillary OCT imag…

Representation Learning

Improving Multimodal Fusion with Hierarchical Mutual Information Maximization for Multimodal Sentiment Analysis

2021-09-01 · EMNLP 2021 11 · Wei Han, Hui Chen, Soujanya Poria

In multimodal sentiment analysis (MSA), the performance of a model highly depends on the quality of synthesized embeddings. These embeddings are generated from the upstream process called multimodal fusion, which aims to…

Multimodal Sentiment AnalysisSentiment Analysis

Multimodal Depression Classification Using Articulatory Coordination Features And Hierarchical Attention Based Text Embeddings

2022-02-13 · Nadee Seneviratne, Carol Espy-Wilson

Multimodal depression classification has gained immense popularity over the recent years. We develop a multimodal depression classification system using articulatory coordination features extracted from vocal tract varia…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

NeuroLingua: A Language-Inspired Hierarchical Framework for Multimodal Sleep Stage Classification Using EEG and EOG

2025-11-12 · Mahdi Samaee, Mehran Yazdi, Daniel Massicotte arxiv

Automated sleep stage classification from polysomnography remains limited by the lack of expressive temporal hierarchies, challenges in multimodal EEG and EOG fusion, and the limited interpretability of deep learning mod…

Causal Inference

Hierarchical Cross-Modality Semantic Correlation Learning Model for Multimodal Summarization

2021-12-16 · Litian Zhang, XiaoMing Zhang, Junshu Pan, Feiran Huang

Multimodal summarization with multimodal output (MSMO) generates a summary with both textual and visual content. Multimodal news report contains heterogeneous contents, which makes MSMO nontrivial. Moreover, it is observ…

Diversity