paper-with-me

Papers Caption Generation

“Caption Generation” 태그가 달린 논문 310편 · 필터 해제

GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning

2025-07-09 · S M Taslim Uddin Raju, Md. Milon Islam, Md Rezwanul Haque, Hamdi Altaheri 외

Microscopic assessment of histopathology images is vital for accurate cancer diagnosis and treatment. Whole Slide Image (WSI) classification and captioning have become crucial tasks in computer-aided pathology. However, …

Caption GenerationClusteringGraph Neural Networkimage-classification+3

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

2025-06-30 · Xiangtai Li, Tao Zhang, Yanwei Li, Haobo Yuan 외

Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack the ground locations and relations for vi…

Caption GenerationObjectVisual Grounding

SonicVerse: Multi-Task Learning for Music Feature-Informed Captioning

2025-06-18 · Anuradha Chopra, Abhinaba Roy, Dorien Herremans

Detailed captions that accurately reflect the characteristics of a music piece can enrich music databases and drive forward research in music AI. This paper introduces a multi-task music captioning model, SonicVerse, tha…

Caption GenerationDescriptiveKey DetectionLarge Language Model+2

EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits

2025-06-11 · Ron Yosef, Moran Yanuka, Yonatan Bitton, Dani Lischinski

Text-guided image editing, fueled by recent advancements in generative AI, is becoming increasingly widespread. This trend highlights the need for a comprehensive framework to verify text-guided edits and assess their qu…

Artifact DetectionCaption GenerationCommon Sense Reasoningtext-guided-image-editing

Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation

2025-06-03 · Israa A. Albadarneh, Bassam H. Hammo, Omar S. Al-Kadi

Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly im…

Caption GenerationImage CaptioningScene UnderstandingSurvey

FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion

2025-06-01 · Shunian Chen, Xinyuan Xie, Zheshu Chen, Liyan Zhao 외

High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and contextual accuracy, primarily due to their…

Audio captioningCaption GenerationInstruction FollowingLarge Language Model

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

2025-05-29 · Shi-Xue Zhang, Hongfa Wang, Duojun Huang, Xin Li 외

Video captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Although large vision-language models (VLMs) …

Caption GenerationLanguage ModelingLanguage ModellingLarge Language Model+2

NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-ID

2025-05-26 · Shihao Li, Chenglong Li, Aihua Zheng, Andong Lu 외

Multi-modal object re-identification (ReID) aims to extract identity features across heterogeneous spectral modalities to enable accurate recognition and retrieval in complex real-world scenarios. However, most existing …

AttributeCaption GenerationDescriptiveMixture-of-Experts+1

GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance

2025-05-25 · Mohammad Mahdi Moradi, Sudhir Mudur

Knowledge-Based Visual Question Answering (KB-VQA) methods focus on tasks that demand reasoning with information extending beyond the explicit content depicted in the image. Early methods relied on explicit knowledge bas…

Caption GenerationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Temporal Object Captioning for Street Scene Videos from LiDAR Tracks

2025-05-22 · Vignesh Gopinathan, Urs Zimmermann, Michael Arnold, Matthias Rottmann

Video captioning models have seen notable advancements in recent years, especially with regard to their ability to capture temporal information. While many research efforts have focused on architectural advancements, suc…

Caption GenerationVideo Captioning

LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts

2025-05-20 · Qifeng Cai, Hao Liang, Hejun Dong, Meiyi Qiang 외

Long videos contain a vast amount of information, making video-text retrieval an essential and challenging task in multimodal learning. However, existing benchmarks suffer from limited video duration, low-quality caption…

Caption GenerationRetrievalText RetrievalVideo Retrieval+2

Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives

2025-05-20 · Xingxing Weng, Chao Pang, Gui-Song Xia

Vision-language modeling (VLM) aims to bridge the information gap between images and natural language. Under the new paradigm of first pre-training on massive image-text pairs and then fine-tuning on task-specific data, …

Caption GenerationContrastive LearningGeneral KnowledgeImage Generation+2

VideoMultiAgents: A Multi-Agent Framework for Video Question Answering

2025-04-25 · Noriyuki Kugo, Xiang Li, Zixin Li, Ashish Gupta 외

Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feedi…

Caption GenerationEgoSchemaMultimodal ReasoningQuestion Answering+4

TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation

2025-04-24 · Ling You, Wenxuan Huang, Xinni Xie, Xiangyi Wei 외

Soccer is a globally popular sporting event, typically characterized by long matches and distinctive highlight moments. Recent advances in Multimodal Large Language Models (MLLMs) offer promising capabilities in temporal…

Caption GenerationDense Video CaptioningLanguage ModelingLanguage Modelling+6

Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training

2025-04-17 · Xinsong Zhang, Yarong Zeng, Xinting Huang, Hu Hu 외

In recent years, the field of vision-language model pre-training has experienced rapid advancements, driven primarily by the continuous enhancement of textual capabilities in large language models. However, existing trai…

Caption GenerationHallucinationLanguage ModelingLanguage Modelling

3D CoCa: Contrastive Learners are 3D Captioners

2025-04-13 · Ting Huang, Zeyu Zhang, Yemin Wang, Hao Tang

3D captioning, which aims to describe the content of 3D scenes in natural language, remains highly challenging due to the inherent sparsity of point clouds and weak cross-modal alignment in existing methods. To address t…

3D dense captioningCaption Generationcross-modal alignmentDecoder+2

Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention

2025-04-03 · Jiuniu Wang, Wenjia Xu, Qingzhong Wang, Antoni B. Chan

Recent advances in image captioning have focused on enhancing accuracy by substantially increasing the dataset and model size. While conventional captioning models exhibit high performance on established metrics such as …

Caption GenerationContrastive LearningImage Captioning

Identifying Multi-modal Knowledge Neurons in Pretrained Transformers via Two-stage Filtering

2025-03-29 · Yugen Sato, Tomohiro Takagi

Recent advances in large language models (LLMs) have led to the development of multimodal LLMs (MLLMs) in the fields of natural language processing (NLP) and computer vision. Although these models allow for integrated vi…

Caption Generationknowledge editingMisinformation

LaPIG: Cross-Modal Generation of Paired Thermal and Visible Facial Images

2025-03-20 · Leyang Wang, Joice Lin

The success of modern machine learning, particularly in facial translation networks, is highly dependent on the availability of high-quality, paired, large-scale datasets. However, acquiring sufficient data is often chal…

Caption GenerationDiversityImage GenerationTranslation

Will Pre-Training Ever End? A First Step Toward Next-Generation Foundation MLLMs via Self-Improving Systematic Cognition

2025-03-16 · Xiaoying Zhang, Da Peng, YiPeng Zhang, Zonghao Guo 외

Recent progress in (multimodal) large language models ((M)LLMs) has shifted focus from pre-training to inference-time compute scaling and post-training optimization, driven by concerns over limited high-quality real-worl…

Caption GenerationImage CaptioningMultimodal ReasoningSelf-Learning
1–20 / 310 다음 →