paper-with-me

Papers

UR-Bench: A Benchmark for Multi-Hop Reasoning over Ultra-High-Resolution Images

2025-12-31 · Siqi Li, Xinyu Cai, Jianbiao Mei, Nianchen Deng, Pinlong Cai, Licheng Wen, Yufan Shen, Xuemeng Yang, Botian Shi, Yong Liu arxiv

Recent multimodal large language models (MLLMs) show strong capabilities in visual-language reasoning, yet their performance on ultra-high-resolution imagery remains largely unexplored. Existing visual question answering (VQA) benchmarks typically rely on medium-resolution data, offering limited visual complexity. To bridge this gap, we introduce Ultra-high-resolution Reasoning Benchmark (UR-Bench), a benchmark designed to evaluate the reasoning capabilities of MLLMs under extreme visual information. UR-Bench comprises two major categories, Humanistic Scenes and Natural Scenes, covering four subsets of ultra-high-resolution images with distinct spatial structures and data sources. Each subset contains images ranging from hundreds of megapixels to gigapixels, accompanied by questions organized into three levels, enabling evaluation of models' reasoning capabilities in ultra-high-resolution scenarios. We further propose an agent-based framework in which a language model performs reasoning by invoking external visual tools. In addition, we introduce Semantic Abstraction and Retrieval tools that enable more efficient processing of ultra-high-resolution images. We evaluate state-of-the-art models using both an end-to-end MLLMs and our agent-based framework, demonstrating the effectiveness of our framework.

📄 PDF Abstract BibTeX arXiv:2601.08748

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning

2026-06-04 · Gexin Huang, Yanting Yang, Myeongkyun Kang, Beidi Zhao 외 arxiv

Vision-language models (VLMs) excel on visual question answering and multimodal reasoning benchmarks. Yet their capability on ultra-resolution images - where critical evidence is tiny, subtle, spatially distant, or distr…

Visual Question AnsweringMultimodal ReasoningAnomaly DetectionVisual Reasoning

Advancing LLM Reasoning Generalists with Preference Trees

2024-04-02 · Lifan Yuan, Ganqu Cui, Hanbin Wang, Ning Ding 외

We introduce Eurus, a suite of large language models (LLMs) optimized for reasoning. Finetuned from Mistral-7B and CodeLlama-70B, Eurus models achieve state-of-the-art results among open-source models on a diverse set of…

BenchmarkingCode GenerationLogical Reasoning

ReXSonoVQA: A Video QA Benchmark for Procedure-Centric Ultrasound Understanding

2026-04-13 · Xucheng Wang, Xiaoman Zhang, Sung Eun Kim, Ankit Pal 외 arxiv

Ultrasound acquisition requires skilled probe manipulation and real-time adjustments. Vision-language models (VLMs) could enable autonomous ultrasound systems, but existing benchmarks evaluate only static images, not dyn…

Gemini: A Family of Highly Capable Multimodal Models

2023-12-19 · The Keyword 2023 12 · Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac 외

This report introduces a new family of multimodal models, Gemini, that exhibit remarkable capabilities across image, audio, video, and text understanding. The Gemini family consists of Ultra, Pro, and Nano sizes, suitabl…

1 Image, 2*2 StitchingArithmetic ReasoningCode GenerationImage Retrieval+5

Echo-α: Large Agentic Multimodal Reasoning Model for Ultrasound Interpretation

2026-04-30 · Jing Zhang, Wentao Jiang, Tao Huang, Zhiwei Wang 외 arxiv

Ultrasound interpretation requires both precise lesion localization and holistic clinical reasoning, yet existing methods typically excel at only one of these capabilities: specialized detectors offer strong localization…

Reinforcement LearningMultimodal Reasoning