paper-with-me

Papers

From Panel to Pixel: Zoom-In Vision-Language Pretraining from Biomedical Scientific Literature

2025-12-02 · Kun Yuan, Min Woo Sun, Zhen Chen, Alejandro Lozano, Xiangteng He, Shi Li, Nassir Navab, Xiaoxiao Sun, Nicolas Padoy, Serena Yeung-Levy arxiv

There is a growing interest in developing strong biomedical vision-language models. A popular approach to achieve robust representations is to use web-scale scientific data. However, current biomedical vision-language pretraining typically compresses rich scientific figures and text into coarse figure-level pairs, discarding the fine-grained correspondences that clinicians actually rely on when zooming into local structures. To tackle this issue, we introduce Panel2Patch, a novel data pipeline that mines hierarchical structure from existing biomedical scientific literature, i.e., multi-panel, marker-heavy figures and their surrounding text, and converts them into multi-granular supervision. Given scientific figures and captions, Panel2Patch parses layouts, panels, and visual markers, then constructs hierarchical aligned vision-language pairs at the figure, panel, and patch levels, preserving local semantics instead of treating each figure as a single data sample. Built on this hierarchical corpus, we develop a granularity-aware pretraining strategy that unifies heterogeneous objectives from coarse didactic descriptions to fine region-focused phrases. By applying Panel2Patch to only a small set of the literature figures, we extract far more effective supervision than prior pipelines, enabling substantially better performance with less pretraining data.

📄 PDF Abstract BibTeX arXiv:2512.02566

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MapReason-OSM: Can Vision-Language Models Make Graph-Verifiable Mobility Decisions from Street Maps ?

2026-06-21 · Srinivas Venkatanarayanan, Clement Pakkam Isaac arxiv

Vision-language models (VLMs) are increasingly used to read maps for logistics, delivery, and accessible navigation, where the output is an actionable decision (a route, a pin, a parking choice) that must respect the roa…

Effective End-to-End Vision Language Pretraining with Semantic Visual Loss

2023-01-18 · Xiaofeng Yang, Fayao Liu, Guosheng Lin

Current vision language pretraining models are dominated by methods using region visual features extracted from object detectors. Given their good performance, the extract-then-process pipeline significantly restricts th…

GPU

Accurate Gigapixel Crowd Counting by Iterative Zooming and Refinement

2023-05-16 · Arian Bakhtiarnia, Qi Zhang, Alexandros Iosifidis

The increasing prevalence of gigapixel resolutions has presented new challenges for crowd counting. Such resolutions are far beyond the memory and computation limits of current GPUs, and available deep neural network arc…

Crowd Counting

Zooming into Comics: Region-Aware RL Improves Fine-Grained Comic Understanding in Vision-Language Models

2025-11-09 · Yule Chen, Yufan Ren, Sabine Süsstrunk arxiv

Complex visual narratives, such as comics, present a significant challenge to Vision-Language Models (VLMs). Despite excelling on natural images, VLMs often struggle with stylized line art, onomatopoeia, and densely pack…

Reinforcement Learning

An Intelligent Pixel Replication Technique by Binary Decomposition for Digital Image Zooming

2014-05-13 · Kaeser M Sabrin, M Haider Ali

Image zooming is the process of enlarging the spatial resolution of a given digital image. We present a novel technique that intelligently modifies the classical pixel replication method for zooming. Our method decompose…