paper-with-me

Papers Image Comprehension

“Image Comprehension” 태그가 달린 논문 49편 · 필터 해제

CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs

2025-05-30 · Ai Jian, Weijie Qiu, Xiaokun Wang, Peiyu Wang 외

Vision-Language Models (VLMs) have demonstrated remarkable progress in multimodal understanding, yet their capabilities for scientific reasoning remains inadequately assessed. Current multimodal benchmarks predominantly …

DiagnosticImage Comprehensionvalid

RGB-Th-Bench: A Dense benchmark for Visual-Thermal Understanding of Vision Language Models

2025-03-25 · Mehdi Moshtaghi, Siavash H. Khajavi, Joni Pajarinen

We introduce RGB-Th-Bench, the first benchmark designed to evaluate the ability of Vision-Language Models (VLMs) to comprehend RGB-Thermal image pairs. While VLMs have demonstrated remarkable progress in visual reasoning…

Image ComprehensionVisual Reasoning

RAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving

2025-03-18 · Yujin Wang, Quanfeng Liu, Zhengxin Jiang, Tianyi Wang 외

Accurately understanding and deciding high-level meta-actions is essential for ensuring reliable and safe autonomous driving systems. While vision-language models (VLMs) have shown significant potential in various autono…

Autonomous DrivingDecision MakingHallucinationImage Comprehension+3

CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation

2025-03-07 · Guanghao Zhang, Tao Zhong, Yan Xia, Zhelun Yu 외

While previous multimodal slow-thinking methods have demonstrated remarkable success in single-image understanding scenarios, their effectiveness becomes fundamentally constrained when extended to more complex multi-imag…

Image ComprehensionMemorization

New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration

2025-02-27 · Xuzheng Yang, Junzhuo Liu, Peng Wang, Guoqing Wang 외

Referring Expression Comprehension (REC) is a foundational cross-modal task that evaluates the interplay of language understanding, image comprehension, and language-to-image grounding. To advance this field, we introduc…

Image ComprehensionReferring ExpressionReferring Expression Comprehension

SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models

2025-02-18 · Xianfu Cheng, Wei zhang, Shiwei Zhang, Jian Yang 외

The increasing application of multi-modal large language models (MLLMs) across various sectors have spotlighted the essence of their output reliability and accuracy, particularly their ability to produce content grounded…

Image ComprehensionQuestion AnsweringText GenerationVisual Question Answering

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

2025-01-10 · You Li, Heyu Huang, Chi Chen, Kaiyu Huang 외

The recent advancement of Multimodal Large Language Models (MLLMs) has significantly improved their fine-grained perception of single images and general comprehension across multiple images. However, existing MLLMs still…

FormImage ComprehensionInstruction Following

RRHF-V: Ranking Responses to Mitigate Hallucinations in Multimodal Large Language Models with Human Feedback

2025-01-01 · Conference 2025 1 · Guoqing Chen, Fu Zhang, Jinghao Lin, Chenglong Lu 외

Multimodal large language models (MLLMs) demonstrate strong capabilities in multimodal understanding, reasoning, and interaction but still face the fundamental limitation of hallucinations, where they generate erroneous …

HallucinationImage ComprehensionImage Description

EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM

2024-12-12 · Zhuofan Zong, Dongzhi Jiang, Bingqi Ma, Guanglu Song 외

Significant achievements in personalization of diffusion models have been witnessed. Conventional tuning-free methods mostly encode multiple reference images by averaging their image embeddings as the injection condition…

Image ComprehensionImage GenerationInstruction FollowingLarge Language Model+2

RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts

2024-12-07 · Xu Liu, Zhouhui Lian

Remote Sensing Vision-Language Models (RS VLMs) have made much progress in the tasks of remote sensing (RS) image comprehension. While performing well in multi-modal reasoning and multi-turn conversations, the existing m…

Change DetectionImage ComprehensionInstruction FollowingLanguage Modeling+6

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation

2024-12-05 · CVPR 2025 1 · Yuying Ge, Yizhuo Li, Yixiao Ge, Ying Shan

In recent years, there has been a significant surge of interest in unifying image comprehension and generation within Large Language Models (LLMs). This growing interest has prompted us to explore extending this unificat…

Image ComprehensionRepresentation LearningText-to-Video GenerationVideo Generation

Survey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges

2024-12-04 · Minghao Shao, Abdul Basit, Ramesh Karri, Muhammad Shafique

Large Language Models (LLMs) represent a class of deep learning models adept at understanding natural language and generating coherent responses to various prompts or queries. These models far exceed the complexity of co…

Code GenerationImage ComprehensionLanguage ModelingLanguage Modelling+4

MMGenBench: Evaluating the Limits of LMMs from the Text-to-Image Generation Perspective

2024-11-21 · Hailang Huang, Yong Wang, Zixuan Huang, Huaqiu Li 외

Large Multimodal Models (LMMs) have demonstrated remarkable capabilities. While existing benchmarks for evaluating LMMs mainly focus on image comprehension, few works evaluate them from the image generation perspective. …

Image ComprehensionImage GenerationModel OptimizationText to Image Generation+1

CLIC: Contrastive Learning Framework for Unsupervised Image Complexity Representation

2024-11-19 · Shipeng Liu, Liang Zhao, Dengfeng Chen

As an essential visual attribute, image complexity affects human image comprehension and directly influences the performance of computer vision tasks. However, accurately assessing and quantifying image complexity faces …

AttributeContrastive LearningImage Comprehension

MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval

2024-11-13 · Yeong-Joon Ju, Ho-Joong Kim, Seong-Whan Lee

Recent multimodal retrieval methods have endowed text-based retrievers with multimodal capabilities by utilizing pre-training strategies for visual-text alignment. They often directly fuse the two modalities for cross-re…

Image ComprehensionInformation Retrievalmultimodal interactionRetrieval

Aquila: A Hierarchically Aligned Visual-Language Model for Enhanced Remote Sensing Image Comprehension

2024-11-09 · Kaixuan Lu, Ruiqian Zhang, Xiao Huang, Yuxing Xie

Recently, large vision language models (VLMs) have made significant strides in visual language capabilities through visual instruction tuning, showing great promise in the field of remote sensing image interpretation. Ho…

Image ComprehensionLanguage ModelingLanguage ModellingLarge Language Model

StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

2024-11-06 · Junming Lin, Zheng Fang, Chi Chen, Zihao Wan 외

The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focus primarily on offline video comprehensi…

Image ComprehensionStreaming video understandingVideo Understanding

Teach Multimodal LLMs to Comprehend Electrocardiographic Images

2024-10-21 · Ruoqi Liu, Yuelin Bai, Xiang Yue, Ping Zhang

The electrocardiogram (ECG) is an essential non-invasive diagnostic tool for assessing cardiac conditions. Existing automatic interpretation methods suffer from limited generalizability, focusing on a narrow range of car…

DiagnosticImage Comprehension

FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion

2024-10-16 · Jiacheng Ruan, Yebin Yang, Zehao Lin, Yuchen Feng 외

Benefiting from the revolutionary advances in large language models (LLMs) and foundational vision models, large vision-language models (LVLMs) have also made significant progress. However, current benchmarks focus on ta…

ArticlesImage Comprehension

FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension

2024-09-23 · Junzhuo Liu, Xuzheng Yang, Weiwei Li, Peng Wang

Referring Expression Comprehension (REC) is a crucial cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding. Consequently, it serves …

Image ComprehensionReferring ExpressionReferring Expression ComprehensionVisual Reasoning
1–20 / 49 다음 →