paper-with-me

Papers

MTANet: Multitask-Aware Network With Hierarchical Multimodal Fusion for RGB-T Urban Scene Understanding

2022-04-05 · journal 2022 4 · WuJie Zhou, Shaohua Dong, Jingsheng Lei, Lu Yu

Understanding urban scenes is a fundamental ability requirement for assisted driving and autonomous vehicles. Most of the available urban scene understanding methods use red-greenblue (RGB) images; however, their segmentation performances are prone to degradation under adverse lighting conditions. Recently, many effective artificial neural networks have been presented for urban scene understanding and have shown that incorporating RGB and thermal (RGB-T) images can improve segmentation accuracy even under unsatisfactory lighting conditions. However, the potential of multimodal feature fusion has not been fully exploited because operations such as simply concatenating the RGB and thermal features or averaging their maps have been adopted. To improve the fusion of multimodal features and the segmentation accuracy, we propose a multitask-aware network (MTANet) with hierarchical multimodal fusion (multiscale fusion strategy) for RGB-T urban scene understanding. We developed a hierarchical multimodal fusion module to enhance feature fusion and built a high-level semantic module to extract semantic information for merging with coarse features at various abstraction levels. Using the multilevel fusion module, we exploited low-, mid-, and high-level fusion to improve segmentation accuracy. The multitask module uses boundary, binary, and semantic supervision to optimize the MTANet parameters. Extensive experiments were performed on two benchmark RGB-T datasets to verify the improved performance of the proposed MTANet compared with state-of-the-art methods

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous VehiclesScene UnderstandingSegmentationThermal Image Segmentation

Similar Papers 제목 키워드 기반

A Multimodal-Multitask Framework with Cross-modal Relation and Hierarchical Interactive Attention for Semantic Comprehension

2025-08-22 · Mohammad Zia Ur Rehman, Devraj Raghuvanshi, Umang Jain, Shubhi Bansal 외 arxiv

A major challenge in multimodal learning is the presence of noise within individual modalities. This noise inherently affects the resulting multimodal representations, especially when these representations are obtained t…

SWAFN: Sentimental Words Aware Fusion Network for Multimodal Sentiment Analysis

2020-12-01 · COLING 2020 8 · Minping Chen, Xia Li

Multimodal sentiment analysis aims to predict sentiment of language text with the help of other modalities, such as vision and acoustic features. Previous studies focused on learning the joint representation of multiple …

Multimodal Sentiment AnalysisSentiment Analysis

Channel Exchanging Networks for Multimodal and Multitask Dense Image Prediction

2021-12-04 · Yikai Wang, Fuchun Sun, Wenbing Huang, Fengxiang He 외

Multimodal fusion and multitask learning are two vital topics in machine learning. Despite the fruitful progress, existing methods for both problems are still brittle to the same challenge -- it remains dilemmatic to int…

Semantic Segmentation

COMMA-DEER: COmmon-sense Aware Multimodal Multitask Approach for Detection of Emotion and Emotional Reasoning in Conversations

2022-10-01 · COLING 2022 10 · Soumitra Ghosh, Gopendra Vikram Singh, Asif Ekbal, Pushpak Bhattacharyya

Mental health is a critical component of the United Nations’ Sustainable Development Goals (SDGs), particularly Goal 3, which aims to provide “good health and well-being”. The present mental health treatment gap is exace…

Common Sense Reasoning

FlexCare: Leveraging Cross-Task Synergy for Flexible Multimodal Healthcare Prediction

2024-06-17 · Muhao Xu, Zhenfeng Zhu, Youru Li, Shuai Zheng 외

Multimodal electronic health record (EHR) data can offer a holistic assessment of a patient's health status, supporting various predictive healthcare tasks. Recently, several studies have embraced the multitask learning …