paper-with-me

홈 › Papers

xMTrans: Temporal Attentive Cross-Modality Fusion Transformer for Long-Term Traffic Prediction

2024-05-08 · Huy Quang Ung, Hao Niu, Minh-Son Dao, Shinya Wada, Atsunori Minamikawa

Traffic predictions play a crucial role in intelligent transportation systems. The rapid development of IoT devices allows us to collect different kinds of data with high correlations to traffic predictions, fostering the development of efficient multi-modal traffic prediction models. Until now, there are few studies focusing on utilizing advantages of multi-modal data for traffic predictions. In this paper, we introduce a novel temporal attentive cross-modality transformer model for long-term traffic predictions, namely xMTrans, with capability of exploring the temporal correlations between the data of two modalities: one target modality (for prediction, e.g., traffic congestion) and one support modality (e.g., people flow). We conducted extensive experiments to evaluate our proposed model on traffic congestion and taxi demand predictions using real-world datasets. The results showed the superiority of xMTrans against recent state-of-the-art methods on long-term traffic predictions. In addition, we also conducted a comprehensive ablation study to further analyze the effectiveness of each module in xMTrans.

📄 PDF Abstract BibTeX arXiv:2405.04841

Code (0)

등록된 구현이 없습니다.

Tasks

Traffic Prediction

Similar Papers 제목 키워드 기반

Continuous Emotion Recognition with Audio-visual Leader-follower Attentive Fusion

2021-07-02 · Su Zhang, Yi Ding, Ziquan Wei, Cuntai Guan

We propose an audio-visual spatial-temporal deep neural network with: (1) a visual block containing a pretrained 2D-CNN followed by a temporal convolutional network (TCN); (2) an aural block containing several parallel T…

Emotion Recognition

Representation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement

2022-05-07 · Bing Li, Jiaxin Chen, Dongming Zhang, Xiuguo Bao 외

Compressed video action recognition has recently drawn growing attention, since it remarkably reduces the storage and computational cost via replacing raw videos by sparsely sampled RGB frames and compressed motion cues …

Action RecognitionDenoisingRepresentation LearningTemporal Action Localization

Cross-Modality Attentive Feature Fusion for Object Detection in Multispectral Remote Sensing Imagery

2021-12-06 · Qingyun Fang, Zhaokui Wang

Cross-modality fusing complementary information of multispectral remote sensing image pairs can improve the perception ability of detection algorithms, making them more robust and reliable for a wider range of applicatio…

object-detectionObject Detection

Multi-channel Attentive Graph Convolutional Network With Sentiment Fusion For Multimodal Sentiment Analysis

2022-01-25 · Luwei Xiao, Xingjiao Wu, Wen Wu, Jing Yang 외

Nowadays, with the explosive growth of multimodal reviews on social media platforms, multimodal sentiment analysis has recently gained popularity because of its high relevance to these social media posts. Although most p…

Multimodal Sentiment AnalysisSentiment Analysis

Attentive Fusion Enhanced Audio-Visual Encoding for Transformer Based Robust Speech Recognition

2020-08-06 · Liangfa Wei, Jie Zhang, JunFeng Hou, Li-Rong Dai

Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strate…

Robust Speech Recognitionspeech-recognitionSpeech Recognition