paper-with-me

홈 › Papers

TokenCom: Vision-Language Model for Multimodal and Multitask Token Communications

2026-02-28 · Feibo Jiang, Siwei Tu, Li Dong, Xiaolong Li, Kezhi Wang, Cunhua Pan, Zhu Han, Jiangzhou Wang arxiv

Visual-Language Models (VLMs), with their strong capabilities in image and text understanding, offer a solid foundation for intelligent communications. However, their effectiveness is constrained by limited token granularity, overlong visual token sequences, and inadequate cross-modal alignment. To overcome these challenges, we propose TaiChi, a novel VLM framework designed for token communications. TaiChi adopts a dual-visual tokenizer architecture that processes both high- and low-resolution images to collaboratively capture pixel-level details and global conceptual features. A Bilateral Attention Network (BAN) is introduced to intelligently fuse multi-scale visual tokens, thereby enhancing visual understanding and producing compact visual tokens. In addition, a Kolmogorov Arnold Network (KAN)-based modality projector with learnable activation functions is employed to achieve precise nonlinear alignment from visual features to the text semantic space, thus minimizing information loss. Finally, TaiChi is integrated into a multimodal and multitask token communication system equipped with a joint VLM-channel coding scheme. Experimental results validate the superior performance of TaiChi, as well as the feasibility and effectiveness of the TaiChi-driven token communication system.

📄 PDF Abstract BibTeX arXiv:2603.00482

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Ada-TokenCom: Rate-Adaptive Token Communications via Large-Model-Driven Token Compression and Generation

2026-08-28 · Zijun Zhang, Li Qiao, Mahdi Boloursaz Mashhadi, Zhen Gao 외 arxiv

Token Communications (TokenCom) has recently emerged as a new paradigm in which tokens serve as unified units for communication and computation, enabling efficient multimodal semantic and goal-oriented transmission. In t…

Semantic Communication

Wireless TokenCom: RL-Based Tokenizer Agreement for Multi-User Wireless Token Communications

2026-02-12 · Farshad Zeinali, Mahdi Boloursaz Mashhadi, Dusit Niyato, Rahim Tafazolli arxiv

Token Communications (TokenCom) has recently emerged as an effective new paradigm, where tokens are the unified units of multimodal communications and computations, enabling efficient digital semantic- and goal-oriented …

Reinforcement Learning

Video TokenCom: Textual Intent-Guided Multi-Rate Video Token Communications with UEP-Based Adaptive Source-Channel Coding

2026-03-02 · Jingxuan Men, Mahdi Boloursaz Mashhadi, Ning Wang, Yi Ma 외 arxiv

Token Communication (TokenCom) is a new paradigm, motivated by the recent success of Large AI Models (LAMs) and Multimodal Large Language Models (MLLMs), where tokens serve as unified units of communication and computati…

Semantic Communication

TokenCompose: Text-to-Image Diffusion with Token-level Supervision

2023-12-06 · CVPR 2024 1 · ZiRui Wang, Zhizhou Sha, Zheng Ding, Yilin Wang 외

We present TokenCompose, a Latent Diffusion Model for text-to-image generation that achieves enhanced consistency between user-specified text prompts and model-generated images. Despite its tremendous success, the standa…

DenoisingImage GenerationObjectText to Image Generation+1

Look Less, Think Faster: Joint Token-Compute Adaptation for Multimodal LLMs

2026-07-22 · Pengcheng Wang, Zhiquan Wang, Jayoung Lee, Zhuoyan Xu 외 arxiv

Multimodal Large Language Models (MLLMs) have recently demonstrated strong performance across vision-language tasks. However, their high inference cost, arising from both the large number of input visual tokens and the h…