xMTrans: Temporal Attentive Cross-Modality Fusion Transformer for Long-Term Traffic Prediction
Traffic predictions play a crucial role in intelligent transportation systems. The rapid development of IoT devices allows us to collect different kinds of data with high correlations to traffic predictions, fostering the development of efficient multi-modal traffic prediction models. Until now, there are few studies focusing on utilizing advantages of multi-modal data for traffic predictions. In this paper, we introduce a novel temporal attentive cross-modality transformer model for long-term traffic predictions, namely xMTrans, with capability of exploring the temporal correlations between the data of two modalities: one target modality (for prediction, e.g., traffic congestion) and one support modality (e.g., people flow). We conducted extensive experiments to evaluate our proposed model on traffic congestion and taxi demand predictions using real-world datasets. The results showed the superiority of xMTrans against recent state-of-the-art methods on long-term traffic predictions. In addition, we also conducted a comprehensive ablation study to further analyze the effectiveness of each module in xMTrans.
Code (0)
등록된 구현이 없습니다.
Tasks
Traffic PredictionSimilar Papers 제목 키워드 기반
Continuous Emotion Recognition with Audio-visual Leader-follower Attentive Fusion
We propose an audio-visual spatial-temporal deep neural network with: (1) a visual block containing a pretrained 2D-CNN followed by a temporal convolutional network (TCN); (2) an aural block containing several parallel T…
Emotion RecognitionRepresentation Learning for Compressed Video Action Recognition via Attentive Cross-modal Interaction with Motion Enhancement
Compressed video action recognition has recently drawn growing attention, since it remarkably reduces the storage and computational cost via replacing raw videos by sparsely sampled RGB frames and compressed motion cues …
Action RecognitionDenoisingRepresentation LearningTemporal Action LocalizationCross-Modality Attentive Feature Fusion for Object Detection in Multispectral Remote Sensing Imagery
Cross-modality fusing complementary information of multispectral remote sensing image pairs can improve the perception ability of detection algorithms, making them more robust and reliable for a wider range of applicatio…
object-detectionObject DetectionMulti-channel Attentive Graph Convolutional Network With Sentiment Fusion For Multimodal Sentiment Analysis
Nowadays, with the explosive growth of multimodal reviews on social media platforms, multimodal sentiment analysis has recently gained popularity because of its high relevance to these social media posts. Although most p…
Multimodal Sentiment AnalysisSentiment AnalysisAttentive Fusion Enhanced Audio-Visual Encoding for Transformer Based Robust Speech Recognition
Audio-visual information fusion enables a performance improvement in speech recognition performed in complex acoustic scenarios, e.g., noisy environments. It is required to explore an effective audio-visual fusion strate…
Robust Speech Recognitionspeech-recognitionSpeech Recognition