paper-with-me

Papers

MVX-ViT: Multimodal Collaborative Perception for 6G V2X Network Management Decisions Using Vision Transformer.

2024-08-30 · IEEE Open Journal of the Communications Society 2024 8 · Ghazi Gharsallah, Georges Kaddoum

Advancements in sixth-generation (6G) networks, coupled with the evolution of multimodal sensing in vehicle-to-everything (V2X) networks, have opened avenues for transformative research into multimodal-based artificial intelligence (AI) applications for wireless communication and network management. However, this promising research direction is often constrained by the limited availability of suitable datasets. In response, this paper introduces a comprehensive configurable co-simulation framework that integrates the state-of-the-art CARLA and Sionna simulators to generate a multimodal multi-view V2X (MVX) dataset. We present novel AI-based models to predict future line-of-sight (LoS) blockages and optimal beam direction as well as an innovative antenna position optimization (APO) solution, all of which are underpinned by the multimodal dataset MVX. Our framework capitalizes on collaborative perception and significantly enhances V2X communication by integrating LiDAR and wireless data. Thorough evaluations demonstrate that our collaborative perception approach outperforms traditional methods of both beam and blockage prediction in terms of accuracy and efficiency. Additionally, we evaluate the importance of infrastructural elements in V2X systems and conduct a computational study to illustrate that our framework is suitable for various operational scenarios and can be used as a digital twin solution. This work not only contributes to the field of V2X wireless communications by providing a versatile framework for network management but also sets the stage for future research on multi-sensor fusion in AI applications for V2X wireless communication environments to enhance the efficiency and resilience of future 6G networks

📄 PDF Abstract BibTeX

Code (1)

ghazigh/MVX 공식 구현

Tasks

Beam PredictionIntelligent CommunicationMultimodal Deep LearningSensor Fusion

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Residual Connection 설명 없음
Attention 설명 없음
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Multi-Head Attention 설명 없음
Vision Transformer The Vision Transformer, or ViT, is a model for image classification that employs a Transformer-like architecture over…

Similar Papers 제목 키워드 기반

MVX-ViT: Multimodal Collaborative Perception for 6G V2X Network Management Decisions Using Vision Transformer

2024-08-30 · IEEE Open Journal of the Communications Society 2024 8 · Ghazi Gharsalla, Georges Kaddoum

Advancements in sixth-generation (6G) networks, coupled with the evolution of multimodal sensing in vehicle-to-everything (V2X) networks, have opened avenues for transformative research into multimodal-based artificial i…

Integration of Mixture of Experts and Multimodal Generative AI in Internet of Vehicles: A Survey

2024-04-25 · Minrui Xu, Dusit Niyato, Jiawen Kang, Zehui Xiong 외

Generative AI (GAI) can enhance the cognitive, reasoning, and planning capabilities of intelligent modules in the Internet of Vehicles (IoV) by synthesizing augmented datasets, completing sensor data, and making sequenti…

Autonomous DrivingDecision MakingManagementMixture-of-Experts

Applying the Wizard-of-Oz Technique to Multimodal Human-Robot Dialogue

2017-03-10 · Matthew Marge, Claire Bonial, Brendan Byrne, Taylor Cassidy 외

Our overall program objective is to provide more natural ways for soldiers to interact and communicate with robots, much like how soldiers communicate with other soldiers today. We describe how the Wizard-of-Oz (WOz) met…

Dialogue ManagementManagementRobot Navigation

AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning

2025-11-14 · Jirong Zha, Yuxuan Fan, Tianyu Zhang, Geng Chen 외 arxiv

Multimodal Large Language Models (MLLMs) have shown promise in single-agent vision tasks, yet benchmarks for evaluating multi-agent collaborative perception remain scarce. This gap is critical, as multi-drone systems pro…

Scene Understanding

A Collaborative Multi-Modality Interaction for VLA-based End-to-End Autonomous Driving

2026-08-21 · Jingtao Sun, Xiaohai He, Yike Zhang, Dong Huang 외 arxiv

Vision-Language-Action (VLA) models have emerged as a powerful paradigm for end-to-end autonomous driving by jointly integrating perception, reasoning, and decision making within a unified multimodal framework. However, …

Visual Question AnsweringTrajectory PlanningAutonomous DrivingDecision Making