Papers Image Comprehension
“Image Comprehension” 태그가 달린 논문 49편 · 필터 해제
CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
Vision-Language Models (VLMs) have demonstrated remarkable progress in multimodal understanding, yet their capabilities for scientific reasoning remains inadequately assessed. Current multimodal benchmarks predominantly …
DiagnosticImage ComprehensionvalidRGB-Th-Bench: A Dense benchmark for Visual-Thermal Understanding of Vision Language Models
We introduce RGB-Th-Bench, the first benchmark designed to evaluate the ability of Vision-Language Models (VLMs) to comprehend RGB-Thermal image pairs. While VLMs have demonstrated remarkable progress in visual reasoning…
Image ComprehensionVisual ReasoningRAD: Retrieval-Augmented Decision-Making of Meta-Actions with Vision-Language Models in Autonomous Driving
Accurately understanding and deciding high-level meta-actions is essential for ensuring reliable and safe autonomous driving systems. While vision-language models (VLMs) have shown significant potential in various autono…
Autonomous DrivingDecision MakingHallucinationImage Comprehension+3CMMCoT: Enhancing Complex Multi-Image Comprehension via Multi-Modal Chain-of-Thought and Memory Augmentation
While previous multimodal slow-thinking methods have demonstrated remarkable success in single-image understanding scenarios, their effectiveness becomes fundamentally constrained when extended to more complex multi-imag…
Image ComprehensionMemorizationNew Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration
Referring Expression Comprehension (REC) is a foundational cross-modal task that evaluates the interplay of language understanding, image comprehension, and language-to-image grounding. To advance this field, we introduc…
Image ComprehensionReferring ExpressionReferring Expression ComprehensionSimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models
The increasing application of multi-modal large language models (MLLMs) across various sectors have spotlighted the essence of their output reliability and accuracy, particularly their ability to produce content grounded…
Image ComprehensionQuestion AnsweringText GenerationVisual Question AnsweringMigician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
The recent advancement of Multimodal Large Language Models (MLLMs) has significantly improved their fine-grained perception of single images and general comprehension across multiple images. However, existing MLLMs still…
FormImage ComprehensionInstruction FollowingRRHF-V: Ranking Responses to Mitigate Hallucinations in Multimodal Large Language Models with Human Feedback
Multimodal large language models (MLLMs) demonstrate strong capabilities in multimodal understanding, reasoning, and interaction but still face the fundamental limitation of hallucinations, where they generate erroneous …
HallucinationImage ComprehensionImage DescriptionEasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
Significant achievements in personalization of diffusion models have been witnessed. Conventional tuning-free methods mostly encode multiple reference images by averaging their image embeddings as the injection condition…
Image ComprehensionImage GenerationInstruction FollowingLarge Language Model+2RSUniVLM: A Unified Vision Language Model for Remote Sensing via Granularity-oriented Mixture of Experts
Remote Sensing Vision-Language Models (RS VLMs) have made much progress in the tasks of remote sensing (RS) image comprehension. While performing well in multi-modal reasoning and multi-turn conversations, the existing m…
Change DetectionImage ComprehensionInstruction FollowingLanguage Modeling+6Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation
In recent years, there has been a significant surge of interest in unifying image comprehension and generation within Large Language Models (LLMs). This growing interest has prompted us to explore extending this unificat…
Image ComprehensionRepresentation LearningText-to-Video GenerationVideo GenerationSurvey of different Large Language Model Architectures: Trends, Benchmarks, and Challenges
Large Language Models (LLMs) represent a class of deep learning models adept at understanding natural language and generating coherent responses to various prompts or queries. These models far exceed the complexity of co…
Code GenerationImage ComprehensionLanguage ModelingLanguage Modelling+4MMGenBench: Evaluating the Limits of LMMs from the Text-to-Image Generation Perspective
Large Multimodal Models (LMMs) have demonstrated remarkable capabilities. While existing benchmarks for evaluating LMMs mainly focus on image comprehension, few works evaluate them from the image generation perspective. …
Image ComprehensionImage GenerationModel OptimizationText to Image Generation+1CLIC: Contrastive Learning Framework for Unsupervised Image Complexity Representation
As an essential visual attribute, image complexity affects human image comprehension and directly influences the performance of computer vision tasks. However, accurately assessing and quantifying image complexity faces …
AttributeContrastive LearningImage ComprehensionMIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
Recent multimodal retrieval methods have endowed text-based retrievers with multimodal capabilities by utilizing pre-training strategies for visual-text alignment. They often directly fuse the two modalities for cross-re…
Image ComprehensionInformation Retrievalmultimodal interactionRetrievalAquila: A Hierarchically Aligned Visual-Language Model for Enhanced Remote Sensing Image Comprehension
Recently, large vision language models (VLMs) have made significant strides in visual language capabilities through visual instruction tuning, showing great promise in the field of remote sensing image interpretation. Ho…
Image ComprehensionLanguage ModelingLanguage ModellingLarge Language ModelStreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
The rapid development of Multimodal Large Language Models (MLLMs) has expanded their capabilities from image comprehension to video understanding. However, most of these MLLMs focus primarily on offline video comprehensi…
Image ComprehensionStreaming video understandingVideo UnderstandingTeach Multimodal LLMs to Comprehend Electrocardiographic Images
The electrocardiogram (ECG) is an essential non-invasive diagnostic tool for assessing cardiac conditions. Existing automatic interpretation methods suffer from limited generalizability, focusing on a narrow range of car…
DiagnosticImage ComprehensionFTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion
Benefiting from the revolutionary advances in large language models (LLMs) and foundational vision models, large vision-language models (LVLMs) have also made significant progress. However, current benchmarks focus on ta…
ArticlesImage ComprehensionFineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression Comprehension
Referring Expression Comprehension (REC) is a crucial cross-modal task that objectively evaluates the capabilities of language understanding, image comprehension, and language-to-image grounding. Consequently, it serves …
Image ComprehensionReferring ExpressionReferring Expression ComprehensionVisual Reasoning