Papers Caption Generation
“Caption Generation” 태그가 달린 논문 310편 · 필터 해제
GNN-ViTCap: GNN-Enhanced Multiple Instance Learning with Vision Transformers for Whole Slide Image Classification and Captioning
Microscopic assessment of histopathology images is vital for accurate cancer diagnosis and treatment. Whole Slide Image (WSI) classification and captioning have become crucial tasks in computer-aided pathology. However, …
Caption GenerationClusteringGraph Neural Networkimage-classification+3DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Multimodal Large Language Models (MLLMs) demonstrate a complex understanding of scenes, benefiting from large-scale and high-quality datasets. Most existing caption datasets lack the ground locations and relations for vi…
Caption GenerationObjectVisual GroundingSonicVerse: Multi-Task Learning for Music Feature-Informed Captioning
Detailed captions that accurately reflect the characteristics of a music piece can enrich music databases and drive forward research in music AI. This paper introduces a multi-task music captioning model, SonicVerse, tha…
Caption GenerationDescriptiveKey DetectionLarge Language Model+2EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
Text-guided image editing, fueled by recent advancements in generative AI, is becoming increasingly widespread. This trend highlights the need for a comprehensive framework to verify text-guided edits and assess their qu…
Artifact DetectionCaption GenerationCommon Sense Reasoningtext-guided-image-editingAttention-based transformer models for image captioning across languages: An in-depth survey and evaluation
Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly im…
Caption GenerationImage CaptioningScene UnderstandingSurveyFusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
High-quality, large-scale audio captioning is crucial for advancing audio understanding, yet current automated methods often generate captions that lack fine-grained detail and contextual accuracy, primarily due to their…
Audio captioningCaption GenerationInstruction FollowingLarge Language ModelVCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
Video captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Although large vision-language models (VLMs) …
Caption GenerationLanguage ModelingLanguage ModellingLarge Language Model+2NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-ID
Multi-modal object re-identification (ReID) aims to extract identity features across heterogeneous spectral modalities to enable accurate recognition and retrieval in complex real-world scenarios. However, most existing …
AttributeCaption GenerationDescriptiveMixture-of-Experts+1GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance
Knowledge-Based Visual Question Answering (KB-VQA) methods focus on tasks that demand reasoning with information extending beyond the explicit content depicted in the image. Early methods relied on explicit knowledge bas…
Caption GenerationQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)Temporal Object Captioning for Street Scene Videos from LiDAR Tracks
Video captioning models have seen notable advancements in recent years, especially with regard to their ability to capture temporal information. While many research efforts have focused on architectural advancements, suc…
Caption GenerationVideo CaptioningLoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
Long videos contain a vast amount of information, making video-text retrieval an essential and challenging task in multimodal learning. However, existing benchmarks suffer from limited video duration, low-quality caption…
Caption GenerationRetrievalText RetrievalVideo Retrieval+2Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
Vision-language modeling (VLM) aims to bridge the information gap between images and natural language. Under the new paradigm of first pre-training on massive image-text pairs and then fine-tuning on task-specific data, …
Caption GenerationContrastive LearningGeneral KnowledgeImage Generation+2VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
Video Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content. However, many existing methods rely on feedi…
Caption GenerationEgoSchemaMultimodal ReasoningQuestion Answering+4TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
Soccer is a globally popular sporting event, typically characterized by long matches and distinctive highlight moments. Recent advances in Multimodal Large Language Models (MLLMs) offer promising capabilities in temporal…
Caption GenerationDense Video CaptioningLanguage ModelingLanguage Modelling+6Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
In recent years, the field of vision-language model pre-training has experienced rapid advancements, driven primarily by the continuous enhancement of textual capabilities in large language models. However, existing trai…
Caption GenerationHallucinationLanguage ModelingLanguage Modelling3D CoCa: Contrastive Learners are 3D Captioners
3D captioning, which aims to describe the content of 3D scenes in natural language, remains highly challenging due to the inherent sparsity of point clouds and weak cross-modal alignment in existing methods. To address t…
3D dense captioningCaption Generationcross-modal alignmentDecoder+2Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention
Recent advances in image captioning have focused on enhancing accuracy by substantially increasing the dataset and model size. While conventional captioning models exhibit high performance on established metrics such as …
Caption GenerationContrastive LearningImage CaptioningIdentifying Multi-modal Knowledge Neurons in Pretrained Transformers via Two-stage Filtering
Recent advances in large language models (LLMs) have led to the development of multimodal LLMs (MLLMs) in the fields of natural language processing (NLP) and computer vision. Although these models allow for integrated vi…
Caption Generationknowledge editingMisinformationLaPIG: Cross-Modal Generation of Paired Thermal and Visible Facial Images
The success of modern machine learning, particularly in facial translation networks, is highly dependent on the availability of high-quality, paired, large-scale datasets. However, acquiring sufficient data is often chal…
Caption GenerationDiversityImage GenerationTranslationWill Pre-Training Ever End? A First Step Toward Next-Generation Foundation MLLMs via Self-Improving Systematic Cognition
Recent progress in (multimodal) large language models ((M)LLMs) has shifted focus from pre-training to inference-time compute scaling and post-training optimization, driven by concerns over limited high-quality real-worl…
Caption GenerationImage CaptioningMultimodal ReasoningSelf-Learning