paper-with-me

Papers Image-to-Text Retrieval

“Image-to-Text Retrieval” 태그가 달린 논문 66편 · 필터 해제

OrganLens: Organ-Specific Representation Learning for CT Foundation Models

2026-07-28 · Zhixuan Ge, Anqi Li, Sadeer Al-Kindi, Hanwen Xu 외 arxiv

A CT examination captures multiple organs, but many biomedical questions concern abnormalities, prognosis, or longitudinal change in a specific organ. These questions require a separate representation for each organ with…

Representation LearningImage-to-Text Retrieval

One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via Hubness

2026-04-30 · Hiroyuki Deguchi, Katsuki Chousa, Yusuke Sakai arxiv

The hubness problem, in which hub embeddings are close to many unrelated examples, occurs often in high-dimensional embedding spaces and may pose a practical threat for purposes such as information retrieval and automati…

Image-to-Text RetrievalInformation RetrievalImage Captioning

Negative Entity Suppression for Zero-Shot Captioning with Synthetic Images

2025-11-12 · Zimao Lu, Hui Xu, Bing Liu, Ke Wang arxiv

Text-only training provides an attractive approach to address data scarcity challenges in zero-shot image captioning (ZIC), avoiding the expense of collecting paired image-text annotations. However, although these approa…

Image-to-Text RetrievalDomain GeneralizationImage Captioning

Evaluating Perspectival Biases in Cross-Modal Retrieval

2025-10-30 · Teerapol Saengsukhiran, Peerawat Chomphooyod, Narabodee Rodjananant, Chompakorn Chaksangchaichot 외 arxiv

Multimodal retrieval systems are expected to operate in a semantic space, agnostic to the language or cultural origin of the query. In practice, however, retrieval outcomes systematically reflect perspectival biases: dev…

Representation LearningImage-to-Text RetrievalCross-Modal RetrievalImage Retrieval

DualCap: Enhancing Lightweight Image Captioning via Dual Retrieval with Similar Scenes Visual Prompts

2025-10-28 · Binbin Li, Guimiao Yang, Zisen Qi, Haiping Wang 외 arxiv

Recent lightweight retrieval-augmented image caption models often utilize retrieved data solely as text prompts, thereby creating a semantic gap by leaving the original visual features unenhanced, particularly for object…

Image-to-Text RetrievalImage CaptioningImage Retrieval

Meta CLIP 2: A Worldwide Scaling Recipe

2025-07-29 · Yung-Sung Chuang, Yang Li, Dong Wang, Ching-Feng Yeh 외 arxiv

Contrastive Language-Image Pretraining (CLIP) is a popular foundation model, supporting from zero-shot classification, retrieval to encoders for multimodal large language models (MLLMs). Although CLIP is successfully tra…

Image-to-Text Retrieval

SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning

2025-07-24 · Si-Woo Kim, MinJu Jeon, Ye-Chan Kim, Soeun Lee 외 arxiv

Zero-shot Image Captioning (ZIC) increasingly utilizes synthetic datasets generated by text-to-image (T2I) models to mitigate the need for costly manual annotation. However, these T2I models often produce images that exh…

Image-to-Text RetrievalImage Captioning

Improving Medical Visual Representation Learning with Pathological-level Cross-Modal Alignment and Correlation Exploration

2025-06-12 · Jun Wang, Lixing Zhu, Xiaohan Yu, Abhir Bhalerao 외

Learning medical visual representations from image-report pairs through joint learning has garnered increasing research attention due to its potential to alleviate the data scarcity problem in the medical domain. The pri…

cross-modal alignmentImage to textImage-to-Text Retrievalobject-detection+4

Efficient Medical Vision-Language Alignment Through Adapting Masked Vision Models

2025-06-10 · Chenyu Lian, Hong-Yu Zhou, Dongyun Liang, Jing Qin 외

Medical vision-language alignment through cross-modal contrastive learning shows promising performance in image-text matching tasks, such as retrieval and zero-shot classification. However, conventional cross-modal contr…

Contrastive LearningImage-text matchingImage to textImage-to-Text Retrieval+5

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution

2025-05-16 · Junyi Yuan, Jian Zhang, Fangyu Wu, Dongming Lu 외

China has a long and rich history, encompassing a vast cultural heritage that includes diverse multimodal information, such as silk patterns, Dunhuang murals, and their associated historical narratives. Cross-modal retri…

Cross-Modal RetrievalImage to textImage-to-Text RetrievalRetrieval+1

SemCORE: A Semantic-Enhanced Generative Cross-Modal Retrieval Framework with MLLMs

2025-04-17 · Haoxuan Li, Yi Bin, Yunshan Ma, Guoqing Wang 외

Cross-modal retrieval (CMR) is a fundamental task in multimedia research, focused on retrieving semantically relevant targets across different modalities. While traditional CMR methods match text and image via embedding-…

Cross-Modal RetrievalImage RetrievalImage to textImage-to-Text Retrieval+2

DART: Disease-aware Image-Text Alignment and Self-correcting Re-alignment for Trustworthy Radiology Report Generation

2025-04-16 · CVPR 2025 1 · Sang-Jun Park, Keun-Soo Heo, Dong-Hee Shin, Young-Han Son 외

The automatic generation of radiology reports has emerged as a promising solution to reduce a time-consuming task and accurately capture critical disease-relevant findings in X-ray images. Previous approaches for radiolo…

Contrastive LearningImage to textImage-to-Text RetrievalRetrieval+1

ABC: Achieving Better Control of Multimodal Embeddings using VLMs

2025-03-01 · Benjamin Schneider, Florian Kerschbaum, Wenhu Chen

Visual embedding models excel at zero-shot tasks like visual retrieval and classification. However, these models cannot be used for tasks that contain ambiguity or require user instruction. These tasks necessitate a mult…

Image to textImage-to-Text RetrievalRetrievalText Retrieval+1

Retaining Knowledge and Enhancing Long-Text Representations in CLIP through Dual-Teacher Distillation

2025-01-01 · CVPR 2025 1 · Yuheng Feng, Changsong Wen, Zelin Peng, Li jiaye 외

Contrastive language-image pretraining models such as CLIP have demonstrated remarkable performance in various text-image alignment tasks. However, the inherent 77-token input limitation and reliance on predominantly…

image-classificationImage ClassificationImage to textImage-to-Text Retrieval+3

DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding

2024-12-02 · Hao Wu, Zhihang Zhong, Xiao Sun

Image captioning models often suffer from performance degradation when applied to novel datasets, as they are typically trained on domain-specific data. To enhance generalization in out-of-domain scenarios, retrieval-aug…

Caption GenerationDomain GeneralizationImage CaptioningImage to text+3

Robotic State Recognition with Image-to-Text Retrieval Task of Pre-Trained Vision-Language Model and Black-Box Optimization

2024-10-30 · Kento Kawaharazuka, Yoshiki Obinata, Naoaki Kanazawa, Kei Okada 외

State recognition of the environment and objects, such as the open/closed state of doors and the on/off of lights, is indispensable for robots that perform daily life support and security tasks. Until now, state recognit…

Image to textImage-to-Text RetrievalLanguage ModelingLanguage Modelling+1

Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization

2024-09-26 · Kento Kawaharazuka, Yoshiki Obinata, Naoaki Kanazawa, Kei Okada 외

In order for robots to autonomously navigate and operate in diverse environments, it is essential for them to recognize the state of their environment. On the other hand, the environmental state recognition has tradition…

Image to textImage-to-Text RetrievalNavigateQuestion Answering+2

GABInsight: Exploring Gender-Activity Binding Bias in Vision-Language Models

2024-07-30 · Ali Abdollahi, Mahdi Ghaznavi, Mohammad Reza Karimi Nejad, Arash Mari Oriyad 외

Vision-language models (VLMs) are intensively used in many downstream tasks, including those requiring assessments of individuals appearing in the images. While VLMs perform well in simple single-person scenarios, in rea…

Image to textImage-to-Text RetrievalSelection biasText Retrieval

Towards a text-based quantitative and explainable histopathology image analysis

2024-07-10 · Anh Tien Nguyen, Trinh Thi Le Vuong, Jin Tae Kwak

Recently, vision-language pre-trained models have emerged in computational pathology. Previous works generally focused on the alignment of image-text pairs via the contrastive pre-training paradigm. Such pre-trained mode…

image-classificationImage ClassificationImage to textImage-to-Text Retrieval+5

BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval

2024-06-14 · Imanol Miranda, Ander Salaberria, Eneko Agirre, Gorka Azkune

Existing Vision-Language Compositionality (VLC) benchmarks like SugarCrepe are formulated as image-to-text retrieval problems, where, given an image, the models need to select between the correct textual description and …

Image RetrievalImage to textImage-to-Text RetrievalRetrieval+1
1–20 / 66 다음 →