paper-with-me

홈 › Papers

Photon: Speedup Volume Understanding with Efficient Multimodal Large Language Models

2026-03-26 · Chengyu Fang, Heng Guo, Zheng Jiang, Chunming He, Xiu Li, Minfeng Xu arxiv

Multimodal large language models are promising for clinical visual question answering tasks, but scaling to 3D imaging is hindered by high computational costs. Prior methods often rely on 2D slices or fixed-length token compression, disrupting volumetric continuity and obscuring subtle findings. We present Photon, a framework that represents 3D medical volumes with token sequences of variable length. Photon introduces instruction-conditioned token scheduling and surrogate gradient propagation to adaptively reduce tokens during both training and inference, which lowers computational cost while mitigating the attention dilution caused by redundant tokens. It incorporates a custom backpropagation rule with gradient restoration to enable differentiable optimization despite discrete token drop. To stabilize token compression and ensure reliable use of visual evidence, Photon further applies regularization objectives that mitigate language-only bias and improve reliability. Experiments on diverse medical visual question answering tasks show that Photon achieves state-of-the-art accuracy while reducing resource usage and accelerating both training and inference.

📄 PDF Abstract BibTeX arXiv:2603.25155

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

Photon Field Networks for Dynamic Real-Time Volumetric Global Illumination

2023-04-14 · David Bauer, Qi Wu, Kwan-Liu Ma

Volume data is commonly found in many scientific disciplines, like medicine, physics, and biology. Experts rely on robust scientific visualization techniques to extract valuable insights from the data. Recent years have …

Data Visualization

Eulerian Single-Photon Vision

2023-01-01 · ICCV 2023 1 · Shantanu Gupta, Mohit Gupta

Single-photon sensors measure light signals at the finest possible resolution -- individual photons. These sensors introduce two major challenges in the form of strong Poisson noise and extremely large data acquisiti…

Edge DetectionImage ReconstructionMotion Estimation

PCNNA: A Photonic Convolutional Neural Network Accelerator

2018-07-23 · Armin Mehrabian, Yousra Al-Kabani, Volker J. Sorger, Tarek El-Ghazawi

Convolutional Neural Networks (CNN) have been the centerpiece of many applications including but not limited to computer vision, speech processing, and Natural Language Processing (NLP). However, the computationally expe…

fairDMS: Rapid Model Training by Data and Model Reuse

2022-04-20 · Ahsan Ali, Hemant Sharma, Rajkumar Kettimuthu, Peter Kenesei 외

Extracting actionable information rapidly from data produced by instruments such as the Linac Coherent Light Source (LCLS-II) and Advanced Photon Source Upgrade (APS-U) is becoming ever more challenging due to high (up t…

Information RetrievalmodelRetrieval

Single-photon Image Super-resolution via Self-supervised Learning

2023-03-03 · YiWei Chen, Chen Jiang, Yu Pan

Single-Photon Image Super-Resolution (SPISR) aims to recover a high-resolution volumetric photon counting cube from a noisy low-resolution one by computational imaging algorithms. In real-world scenarios, pairs of traini…

Image Super-ResolutionSelf-Supervised LearningSuper-Resolution