paper-with-me

Papers

When Visual Grounding Meets Gigapixel-level Large-scale Scenes: Benchmark and Approach

2024-01-01 · CVPR 2024 1 · Tao Ma, Bing Bai, Haozhe Lin, Heyuan Wang, Yu Wang, Lin Luo, Lu Fang

Visual grounding refers to the process of associating natural language expressions with corresponding regions within an image. Existing benchmarks for visual grounding primarily operate within small-scale scenes with a few objects. Nevertheless recent advances in imaging technology have enabled the acquisition of gigapixel-level images providing high-resolution details in large-scale scenes containing numerous objects. To bridge this gap between imaging and computer vision benchmarks and make grounding more practically valuable we introduce a novel dataset named GigaGrounding designed to challenge visual grounding models in gigapixel-level large-scale scenes. We extensively analyze and compare the dataset with existing benchmarks demonstrating that GigaGrounding presents unique challenges such as large-scale scene understanding gigapixel-level resolution significant variations in object scales and the "multi-hop expressions". Furthermore we introduced a simple yet effective grounding approach which employs a "glance-to-zoom-in" paradigm and exhibits enhanced capabilities for addressing the GigaGrounding task. The dataset is available at www.gigavision.ai.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingVisual Grounding

Similar Papers 제목 키워드 기반

Neural Image Compression for Gigapixel Histopathology Image Analysis

2018-11-07 · David Tellez, Geert Litjens, Jeroen van der Laak, Francesco Ciompi

We propose Neural Image Compression (NIC), a two-step method to build convolutional neural networks for gigapixel image analysis solely using weak image-level labels. First, gigapixel images are compressed using a neural…

Image CompressionMedical Diagnosis

PathoArgus: Advancing Evidence-Grounded Long-Context Visual Reasoning across Gigapixel Whole-Slide and Multi-Slide Case Contexts

2026-08-18 · Bowen Liu, Qixiang Zhang, Xiaomeng Li arxiv

Whole-slide pathology reasoning requires models to integrate gigapixel-scale visual evidence across complete case-linked slides, yet current question-answering benchmarks primarily measure final answer accuracy--a metric…

Visual Reasoning

Speed Up Object Detection on Gigapixel-Level Images With Patch Arrangement

2022-01-01 · CVPR 2022 1 · Jiahao Fan, Huabin Liu, Wenjie Yang, John See 외

With the appearance of super high-resolution (e.g., gigapixel-level) images, performing efficient object detection on such images becomes an important issue. Most existing works for efficient object detection on high…

Objectobject-detectionObject DetectionReal-Time Object Detection

PANDA: A Gigapixel-level Human-centric Video Dataset

2020-03-10 · CVPR 2020 6 · Xueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo 외

We present PANDA, the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis. The videos in PANDA were captured by a gigapixel camera and cover real-world scenes w…

4kAttributeHuman Detection

PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology

2025-12-19 · Fengchun Liu, Songhan Jiang, Linghan Cai, Ziyue Wang 외 arxiv

While Vision-Language Models (VLMs) have achieved notable progress in computational pathology (CPath), the gigapixel scale and spatial heterogeneity of Whole Slide Images (WSIs) continue to pose challenges for multimodal…

Instruction Following