paper-with-me

홈 › Papers

GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations

2026-06-02 · Jonggwon Park, Seongeun Lee, Junhyun Park, Hannah Yun, Hyunwoong Kim, Sohyun Jeong, Hyewon Kang, Byungmu Yoon, Kyoyun Choi arxiv

Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workflows. However, this pairing reveals a mismatch in scale: each finding occupies only a small region of the image, yet supervision is provided only at the global image-report level. This poses a central challenge: prior approaches spread weight densely across all patches rather than concentrating on the sparse subset relevant to a given query. To address this, we present GLINT (Gated Language-Image alignmeNT), a framework that explicitly models this sparse correspondence. On the alignment side, we introduce Sparsely Gated Alignment, a novel architecture in which a sigmoid gate over a separate gate embedding space activates only the patches relevant to each textual query, enforcing explicit sparsity. On the representation side, we add Dense Feature Regularization, which anchors the trainable encoder's intermediate features to a frozen self-supervised learning (SSL) teacher, preserving the fine-grained patch features that the gate relies on. The same recipe applies to both 2D chest X-ray (CXR) and 3D chest computed tomography (CT), built with DINOv3 and V-JEPA 2.1, respectively. GLINT enables zero-shot classification, grounding, and segmentation from free-text queries, and to our knowledge is the first to demonstrate zero-shot segmentation on 3D CT volumes without mask supervision. Notably, the most pronounced gains arise on zero-shot grounding and segmentation, where sparse, query-specific localization is required, consistent with our design intent. In downstream evaluation, GLINT outperforms both SSL encoders and medical VLMs on classification, report generation, and segmentation.

📄 PDF Abstract BibTeX arXiv:2606.03180

Code (0)

등록된 구현이 없습니다.

Tasks

Self-Supervised Learning

Similar Papers 제목 키워드 기반

GLINT-RU: Gated Lightweight Intelligent Recurrent Units for Sequential Recommender Systems

2024-06-06 · Sheng Zhang, Maolin Wang, Wanyu Wang, Jingtong Gao 외

Transformer-based models have gained significant traction in sequential recommender systems (SRSs) for their ability to capture user-item interactions effectively. However, these models often suffer from high computation…

Recommendation SystemsSequential Recommendation

HMVLA: Hyperbolic Multimodal Fusion for Vision-Language-Action Models

2026-01-28 · Kun Wang, Xiao Feng, Mingcheng Qu, Tonghua Su arxiv

Vision Language Action (VLA) models have recently shown great potential in bridging multimodal perception with robotic control. However, existing methods often rely on direct fine-tuning of pre-trained Vision-Language Mo…

Real-time Image-based Lighting of Glints

2025-07-03 · Tom Kneiphof, Reinhard Klein arxiv

Image-based lighting is a widely used technique to reproduce shading under real-world lighting conditions, especially in real-time rendering applications. A particularly challenging scenario involves materials exhibiting…

Quadruped Parkour Learning: Sparsely Gated Mixture of Experts with Visual Input

2026-04-21 · Michael Ziegltrum, Jianhao Jiao, Tianhu Peng, Chengxu Zhou 외 arxiv

Robotic parkour provides a compelling benchmark for advancing locomotion over highly challenging terrain, including large discontinuities such as elevated steps. Recent approaches have demonstrated impressive capabilitie…

Computational Efficiency

Glinthawk: A Two-Tiered Architecture for Offline LLM Inference

2025-01-20 · Pouya Hamadanian, Sadjad Fouladi

We introduce Glinthawk, an architecture for offline Large Language Model (LLM) inference. By leveraging a two-tiered structure, Glinthawk optimizes the utilization of the high-end accelerators ("Tier 1") by offloading th…

CPULanguage ModelingLanguage ModellingLarge Language Model