paper-with-me

Papers Image Captioning

“Image Captioning” 태그가 달린 논문 2,086편 · 필터 해제

A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss

2026-09-01 · Suryaansh Jain, Rahasya Barkur, Vishal G, Ryan Rossi 외 hf

An image may be worth a thousand words, but most captioning models describe it in only a few. Modern vision-language models produce fluent high-level captions, yet routinely miss the attributes, counts, textures, materia…

Image Captioning

Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding

2026-08-27 · Gabriel Pirlogeanu, Dan Oneata, Horia Cucu, Herman Kamper arxiv

In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build models from such visually grounded speec…

Visual GroundingImage CaptioningKeyword Spotting

Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models

2026-08-26 · Raúl Vázquez, Aman Sinha, Chuyuan Li, Claudio Savelli 외 arxiv

In 2026, we held the fourth iteration of the SHROOM Shared Task series: SHROOM-Visions (\textbf{S}hared-task on \textbf{H}allucinations and \textbf{R}elated \textbf{O}bservable \textbf{O}vergeneration \textbf{M}istakes i…

Image CaptioningText Generation

Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

2026-08-21 · Haonan Jia, Shichao Dong, Zenghui Sun, Jiawen Zheng 외 arxiv

Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Language Models (LVLMs) to explore novel reasoning strategies. This limitation leads…

Reinforcement LearningImage Captioning

Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis

2026-08-21 · Yantao Li, Huanlin Gao, Fang Zhao, Chao Tan 외 arxiv

Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model verifies them in parallel. Its lossless guarantee has motivated a line of work th…

Visual ReasoningImage Captioning

Towards Clinically Faithful Medical Image Captioning via Enhanced Vision-Language Alignment

2026-08-20 · Yunseo Lee, Hyun Jun Kim, Heeseung Shin, Changwon Lim arxiv

Medical image captioning is a technique that accelerates early-stage diagnostic workflows and enhances the interpretability of medical diagnostic AI systems. However, unlike general image captioning, clinically reliable …

Image CaptioningType prediction

Clinically Structured Surrogate Rewards for Post-SFT Medical Image Captioning

2026-08-19 · Hyun Jun Kim, Heeseung Shin, Changwon Lim arxiv

Medical image captioning requires translating heterogeneous visual evidence into concise clinical descriptions, where errors in findings, assertion states, or anatomical relations can alter clinical meaning despite surfa…

Image Captioning

Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning

2026-08-02 · Zhiyue Liu, Wenkai Zhou, Jian Qin, Qipeng Jiang arxiv

Zero-shot image captioning aims to generate image descriptions without annotated image-text pairs. Recent approaches exploit text-to-image models to synthesize training data from text-only corpora, but most focus on impr…

Image Captioning

Towards Faithful Sentimental Image Captioning via Evidence-Aware Multi-Agent Reasoning

2026-07-28 · Tiecheng Cai, Zexian Yang, Chao Chen, Shanshan Lin 외 arxiv

Sentimental Image Captioning (SIC) requires balancing emotional expression with visual fidelity. Existing methods often struggle with this trade-off, leading to hallucinations due to insufficient local grounding and the …

Image Captioning

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions

2026-07-25 · Zhijiang Tang, Jiaxin Qi, Kaihua Tang, Yuhua Zheng 외 arxiv

Image captioning is a primary task in vision--language research, yet assessing how faithfully a caption preserves image semantics without relying on reference captions remains unsettled. Prevailing evaluations rely on hu…

Image Captioning

LEMUR 2: Unlocking Neural Network Diversity for AI

2026-07-07 · Tolgay Atinc Uzun, Waleed Khalid, Saif U Din, Sai Revanth Mulukuledu 외 arxiv

Existing NAS benchmarks (e.g., NAS-Bench, NATS-Bench) cover only narrow, task-specific regions of the architectural design space and lack cross-domain or deployment-aware evaluation. LEMUR 2 introduces a large-scale, ext…

Image Captioning

SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

2026-07-05 · Kai Tang, Jinhao You, Bohua Zhang, Yichen Guo 외 arxiv

Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answering. However, they remain susceptible to hallucinations, generating …

Visual Question AnsweringFeature EngineeringImage Captioning

Paying More Attention to Visual Tokens in Self-Evolving Large Multimodal Models

2026-06-25 · Shravan Venkatraman, Ritesh Thawkar, Omkar Thawakar, Rao Muhammad Anwer 외 arxiv

Recently, self-evolving large multimodal models (LMMs) have received attention for improving visual reasoning in a purely unsupervised setting. However, multi-role self-play and self-consistency reward schemes in existin…

Visual Question AnsweringImage CaptioningVisual Reasoning

VisChronos: Revolutionizing Image Captioning Through Real-Life Events

2026-06-23 · Phuc-Tan Nguyen, Hieu Nguyen, Minh-Triet Tran, Trung-Nghia Le arxiv

This paper aims to bridge the semantic gap between visual content and natural language understanding by leveraging historical events in the real world as a source of knowledge for caption generation. We propose VisChrono…

Natural Language UnderstandingDense CaptioningImage Captioning

Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications

2026-06-23 · Shayon Dasgupta, Avijit Dasgupta, C. V. Jawahar arxiv

Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified architecture enables strong performance on tasks like image caption…

Visual Question AnsweringImage Captioning

SAGE: An Expert-Annotated South Asian GI Endoscopy Dataset for Multimodal Learning and Hallucination Analysis

2026-06-20 · Niyoj Oli, Sachin Acharya, Sandesh Pokhrel, Sanjay Bhandari 외 arxiv

Gastrointestinal cancers represent a growing health burden in the South Asian region, driven largely by rapid changes in socio-economic conditions and lifestyle habits. However, early diagnosis remains limited by inadequ…

Multi-Label ClassificationMulti-class ClassificationVisual Question AnsweringImage Captioning

Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders

2026-06-19 · Sergio Lanza, Jae Hee Lee, Stefan Wermter arxiv

Vision Language Models (VLMs) have demonstrated impressive performance in tasks requiring joint understanding of images and text, such as image captioning and Visual Question Answering (VQA), but our understanding of the…

Visual Question AnsweringImage Captioning

MIRCaps: A Large-Scale Mixed-Domain Dataset with Image-Level and Region-Level Captions for Fine-Grained Vision-Language Learning

2026-06-19 · Arlindo Luciano Tulumba Roberto, Hyungjoon Kim arxiv

Despite recent progress in Vision-Language Models (VLMs), mixed-domain image-caption datasets for both general-purpose and CCTV-based video surveillance systems remain limited. To address this gap, we introduce a large-s…

Object DetectionImage Captioning

Hierarchical Multi-Modal Retrieval for Knowledge-Grounded News Image Captioning

2026-06-17 · Minh-Loi Nguyen, Xuan-Vu Le, Long-Bao Nguyen, Hoang-Bach Ngo 외 arxiv

Traditional image captioning methods often struggle to generate comprehensive, context-rich descriptions, especially for details not directly observable from visual cues. To overcome this, we propose a novel retrieval-au…

Image Captioning

CIAN: Multi-Stage Framework for Event-Enriched Image Captioning via Retrieval-Augmented Generation

2026-06-16 · Trinh Thi Thu Hien, Trung-Nghia Le arxiv

Event-enriched image captioning describes not only visible content but also the broader context of events, including timing, location, and participants, capabilities missing in most pixel-bound models. We propose the Con…

Image Captioning
1–20 / 2,086 다음 →