Papers Image-to-Text Retrieval
“Image-to-Text Retrieval” 태그가 달린 논문 66편 · 필터 해제
OrganLens: Organ-Specific Representation Learning for CT Foundation Models
A CT examination captures multiple organs, but many biomedical questions concern abnormalities, prognosis, or longitudinal change in a specific organ. These questions require a separate representation for each organ with…
Representation LearningImage-to-Text RetrievalOne Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness
The hubness problem, in which hub embeddings are close to many unrelated examples, occurs often in high-dimensional embedding spaces and may pose a practical threat for purposes such as information retrieval and automati…
Image-to-Text RetrievalInformation RetrievalImage CaptioningNegative Entity Suppression for Zero-Shot Captioning with Synthetic Images
Text-only training provides an attractive approach to address data scarcity challenges in zero-shot image captioning (ZIC), avoiding the expense of collecting paired image-text annotations. However, although these approa…
Image-to-Text RetrievalDomain GeneralizationImage CaptioningEvaluating Perspectival Biases in Cross-Modal Retrieval
Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query. In practice, however, retrieval outcomes systematically reflect perspectival biases: dev…
Representation LearningImage-to-Text RetrievalCross-Modal RetrievalImage RetrievalDualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts
Recent lightweight retrieval-augmented image caption models often utilize retrieved data solely as text prompts, thereby creating a semantic gap by leaving the original visual features unenhanced, particularly for object…
Image-to-Text RetrievalImage CaptioningImage RetrievalMeta CLIP 2: A Worldwide Scaling Recipe
Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully tra…
Image-to-Text RetrievalSynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
Zero-shot Image Captioning (ZIC) increasingly utilizes synthetic datasets generated by text-to-image (T2I) models to mitigate the need for costly manual annotation. However, these T2I models often produce images that exh…
Image-to-Text RetrievalImage CaptioningImproving Medical Visual Representation Learning with Pathological-level Cross-Modal Alignment and Correlation Exploration
Learning medical visual representations from image-report pairs through joint learning has garnered increasing research attention due to its potential to alleviate the data scarcity problem in the medical domain. The pri…
cross-modal alignmentImage to textImage-to-Text Retrievalobject-detection+4Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models
Medical vision-language alignment through cross-modal contrastive learning shows promising performance in image-text matching tasks, such as retrieval and zero-shot classification. However, conventional cross-modal contr…
Contrastive LearningImage-text matchingImage to textImage-to-Text Retrieval+5Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution
China has a long and rich history, encompassing a vast cultural heritage that includes diverse multimodal information, such as silk patterns, Dunhuang murals, and their associated historical narratives. Cross-modal retri…
Cross-Modal RetrievalImage to textImage-to-Text RetrievalRetrieval+1SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs
Cross-modal retrieval (CMR) is a fundamental task in multimedia research, focused on retrieving semantically relevant targets across different modalities. While traditional CMR methods match text and image via embedding-…
Cross-Modal RetrievalImage RetrievalImage to textImage-to-Text Retrieval+2DART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report Generation
The automatic generation of radiology reports has emerged as a promising solution to reduce a time-consuming task and accurately capture critical disease-relevant findings in X-ray images. Previous approaches for radiolo…
Contrastive LearningImage to textImage-to-Text RetrievalRetrieval+1ABC: Achieving Better Control of Multimodal Embeddings using VLMs
Visual embedding models excel at zero-shot tasks like visual retrieval and classification. However, these models cannot be used for tasks that contain ambiguity or require user instruction. These tasks necessitate a mult…
Image to textImage-to-Text RetrievalRetrievalText Retrieval+1Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher Distillation
Contrastive language-image pretraining models such as CLIP have demonstrated remarkable performance in various text-image alignment tasks. However, the inherent 77-token input limitation and reliance on predominantly…
image-classificationImage ClassificationImage to textImage-to-Text Retrieval+3DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding
Image captioning models often suffer from performance degradation when applied to novel datasets, as they are typically trained on domain-specific data. To enhance generalization in out-of-domain scenarios, retrieval-aug…
Caption GenerationDomain GeneralizationImage CaptioningImage to text+3Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization
State recognition of the environment and objects, such as the open/closed state of doors and the on/off of lights, is indispensable for robots that perform daily life support and security tasks. Until now, state recognit…
Image to textImage-to-Text RetrievalLanguage ModelingLanguage Modelling+1Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization
In order for robots to autonomously navigate and operate in diverse environments, it is essential for them to recognize the state of their environment. On the other hand, the environmental state recognition has tradition…
Image to textImage-to-Text RetrievalNavigateQuestion Answering+2GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models
Vision-language models (VLMs) are intensively used in many downstream tasks, including those requiring assessments of individuals appearing in the images. While VLMs perform well in simple single-person scenarios, in rea…
Image to textImage-to-Text RetrievalSelection biasText RetrievalTowards a text-based quantitative and explainable histopathology image analysis
Recently, vision-language pre-trained models have emerged in computational pathology. Previous works generally focused on the alignment of image-text pairs via the contrastive pre-training paradigm. Such pre-trained mode…
image-classificationImage ClassificationImage to textImage-to-Text Retrieval+5BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval
Existing Vision-Language Compositionality (VLC) benchmarks like SugarCrepe are formulated as image-to-text retrieval problems, where, given an image, the models need to select between the correct textual description and …
Image RetrievalImage to textImage-to-Text RetrievalRetrieval+1