paper-with-me

홈 › Papers

Cross-Modality Attention with Semantic Graph Embedding for Multi-Label Classification

2019-12-17 · Renchun You, Zhiyao Guo, Lei Cui, Xiang Long, Yingze Bao, Shilei Wen

Multi-label image and video classification are fundamental yet challenging tasks in computer vision. The main challenges lie in capturing spatial or temporal dependencies between labels and discovering the locations of discriminative features for each class. In order to overcome these challenges, we propose to use cross-modality attention with semantic graph embedding for multi label classification. Based on the constructed label graph, we propose an adjacency-based similarity graph embedding method to learn semantic label embeddings, which explicitly exploit label relationships. Then our novel cross-modality attention maps are generated with the guidance of learned label embeddings. Experiments on two multi-label image classification datasets (MS-COCO and NUS-WIDE) show our method outperforms other existing state-of-the-arts. In addition, we validate our method on a large multi-label video classification dataset (YouTube-8M Segments) and the evaluation results demonstrate the generalization capability of our method.

📄 PDF Abstract BibTeX arXiv:1912.07872

Code (0)

등록된 구현이 없습니다.

Tasks

ClassificationGeneral ClassificationGraph Embeddingimage-classificationImage ClassificationMulti-Label ClassificationMUlTI-LABEL-ClASSIFICATIONMulti-Label Image ClassificationVideo Classification

Similar Papers 제목 키워드 기반

Bilateral Cross-Modality Graph Matching Attention for Feature Fusion in Visual Question Answering

2021-12-14 · JianJian Cao, Xiameng Qin, Sanyuan Zhao, Jianbing Shen

Answering semantically-complicated questions according to an image is challenging in Visual Question Answering (VQA) task. Although the image can be well represented by deep learning, the question is always simply embedd…

Graph MatchingQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Toward Effective Multimodal Graph Foundation Model: A Divide-and-Conquer Based Approach

2026-02-04 · Sicheng Liu, Xunkai Li, Daohan Su, Ru Zhang 외 arxiv

Graph Foundation Models (GFMs) have achieved remarkable success in generalizing across diverse domains. However, they mainly focus on Text-Attributed Graphs (TAGs), leaving Multimodal-Attributed Graphs (MAGs) largely unt…

Semantic Item Graph Enhancement for Multimodal Recommendation

2025-08-08 · Xiaoxiong Zhang, Xin Zhou, Zhiwei Zeng, Dusit Niyato 외 arxiv

Multimodal recommendation systems have attracted increasing attention for their improved performance by leveraging items' multimodal information. Prior methods often build modality-specific item-item semantic graphs from…

Multimodal RecommendationContrastive Learning

Exploiting Semantic Embedding and Visual Feature for Facial Action Unit Detection

2021-06-19 · CVPR 2021 1 · Huiyuan Yang, Lijun Yin, Yi Zhou, Jiuxiang Gu

Recent study on detecting facial action units (AU) has utilized auxiliary information (i.e., facial landmarks, relationship among AUs and expressions, web facial images, etc.), in order to improve the AU detection pe…

Action Unit DetectionFacial Action Unit DetectionSentence

GTP-4o: Modality-prompted Heterogeneous Graph Learning for Omni-modal Biomedical Representation

2024-07-08 · Chenxin Li, Xinyu Liu, Cheng Wang, Yifan Liu 외

Recent advances in learning multi-modal representation have witnessed the success in biomedical domains. While established techniques enable handling multi-modal information, the challenges are posed when extended to var…

BenchmarkingGraph EmbeddingGraph LearningHallucination