paper-with-me

Papers

Large Language Model-Driven Distributed Integrated Multimodal Sensing and Semantic Communications

2025-05-20 · Yubo Peng, Luping Xiang, Bingxin Zhang, Kun Yang

Traditional single-modal sensing systems-based solely on either radio frequency (RF) or visual data-struggle to cope with the demands of complex and dynamic environments. Furthermore, single-device systems are constrained by limited perspectives and insufficient spatial coverage, which impairs their effectiveness in urban or non-line-of-sight scenarios. To overcome these challenges, we propose a novel large language model (LLM)-driven distributed integrated multimodal sensing and semantic communication (LLM-DiSAC) framework. Specifically, our system consists of multiple collaborative sensing devices equipped with RF and camera modules, working together with an aggregation center to enhance sensing accuracy. First, on sensing devices, LLM-DiSAC develops an RF-vision fusion network (RVFN), which employs specialized feature extractors for RF and visual data, followed by a cross-attention module for effective multimodal integration. Second, a LLM-based semantic transmission network (LSTN) is proposed to enhance communication efficiency, where the LLM-based decoder leverages known channel parameters, such as transceiver distance and signal-to-noise ratio (SNR), to mitigate semantic distortion. Third, at the aggregation center, a transformer-based aggregation model (TRAM) with an adaptive aggregation attention mechanism is developed to fuse distributed features and enhance sensing accuracy. To preserve data privacy, a two-stage distributed learning strategy is introduced, allowing local model training at the device level and centralized aggregation model training using intermediate features. Finally, evaluations on a synthetic multi-view RF-visual dataset generated by the Genesis simulation engine show that LLM-DiSAC achieves a good performance.

📄 PDF Abstract BibTeX arXiv:2505.18194

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelSemantic Communication

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Concatenated Skip Connection A Concatenated Skip Connection is a type of skip connection that seeks to reuse features by concatenating them to new layers, allowing more information to be retained from…

Similar Papers 제목 키워드 기반

CM2: Multimodal Cultural Reasoning via an Integrated Multi-Agent Framework

2026-08-31 · Qi Li, Zhaojie Kang, Yingjie He, Zheng Lin 외 arxiv

Multimodal Large Language Models (MLLMs) have shown remarkable success in STEM domains, where progress is often driven by vertical, step-by-step deduction under relatively stable symbol systems. Their horizontal, interdi…

ERGeoBench:A Comprehensive Benchmark for Embodied Reasoning and Geo-localization in Multimodal Large Language Models

2026-05-29 · Kaiwen Xue, Tao Wei, Guoxin Zhang, Zhonghong Ou 외 arxiv

Multimodal large language models (MLLMs) have shown strong potential as embodied agents, yet embodied geo-localization remains underexplored due to the lack of fine-grained evaluation. We introduce ERGeoBench, a diagnost…

Common Sense ReasoningSpatial Reasoning

Data-driven and distributed governance of building facilities management using decentralized autonomous organization, digital twin, and large language models

2026-04-16 · Reachsak Ly, Alireza Shojaei, Xinghua Gao, Philip Agee 외 arxiv

While traditional AI and data-driven facilities management approaches have improved building operational efficiency, they remain constrained by centralized organizational structures that are vulnerable to cyber attacks, …

SIMAC: A Semantic-Driven Integrated Multimodal Sensing And Communication Framework

2025-03-11 · Yubo Peng, Luping Xiang, Kun Yang, Feibo Jiang 외

Traditional single-modality sensing faces limitations in accuracy and capability, and its decoupled implementation with communication systems increases latency in bandwidth-constrained environments. Additionally, single-…

Large Language ModelMulti-Task Learning

Aether Weaver: Multimodal Affective Narrative Co-Generation with Dynamic Scene Graphs

2025-07-29 · Saeed Ghorbani arxiv

We introduce Aether Weaver, a novel, integrated framework for multimodal narrative co-generation that overcomes limitations of sequential text-to-visual pipelines. Our system concurrently synthesizes textual narratives, …