paper-with-me

홈 › Papers

A Survey on Bridging VLMs and Synthetic Data

2025-05-09 · OpenReview 2025 5 · Mohammad Ghiasvand Mohammadkhani, Saeedeh Momtazi, Hamid Beigy

Vision-language models (VLMs) have significantly advanced multimodal AI by learning joint representations of visual and textual data. However, their progress is hindered by challenges in acquiring high-quality, aligned datasets, including issues of cost, privacy, and scarcity. On the other hand, synthetic data, created through the use of generative AI—which can even include VLMs—offers a scalable and cost-effective solution to these challenges. This paper presents the first comprehensive survey on bridging VLMs and synthetic data, exploring both the role of synthetic data in VLMs and the role of VLMs in synthetic data. First, we provide a preliminary overview by briefly explaining the architecture of two basic VLMs and, after studying a large number of previous works, offer an extensive survey of the previously proposed methodologies and potential future directions in this area.

📄 PDF Abstract BibTeX

Code (1)

mghiasvand1/Awesome-VLM-Synthetic-Data

Tasks

Survey

Similar Papers 제목 키워드 기반

Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench

2025-10-30 · Fenfen Lin, Yesheng Liu, Haiyu Xu, Chen Yue 외 arxiv

Reading measurement instruments is effortless for humans and requires relatively little domain expertise, yet it remains surprisingly challenging for current vision-language models (VLMs) as we find in preliminary evalua…

Vision-Language Models for Deployable Social Robot Navigation: Bridging Semantic Reasoning and Low-Level Control

2026-06-27 · Runji Cai, Toshihiko Yamasaki, Ling Xiao arxiv

Social robot navigation (SRN) requires more than geometric path planning; it demands understanding human intentions, social norms, and contextual cues to generate socially compliant behaviors. Although classical navigati…

Collision AvoidanceRobot Navigation

Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting

2025-08-06 · Yuyang Liu, Qiuhe Hong, Linlan Huang, Alexandra Gomez-Villa 외 arxiv

Vision-language models (VLMs), spanning predictive architectures to generative Multimodal Large Language Models (MLLMs), have revolutionized artificial intelligence through powerful cross-modal alignment and zero-shot ge…

Compositional Zero-Shot LearningZero-shot GeneralizationContinual Learning

Synthetic Tabular Data Generation: A Comparative Survey for Modern Techniques

2025-07-15 · Raju Challagundla, Mohsen Dorodchi, Pu Wang, Minwoo Lee arxiv

As privacy regulations become more stringent and access to real-world data becomes increasingly constrained, synthetic data generation has emerged as a vital solution, especially for tabular datasets, which are central t…

Synthetic Data GenerationTabular Data Generation

Bridging Resolution: A Survey of the State of the Art

2020-12-01 · COLING 2020 8 · Hideo Kobayashi, Vincent Ng

Bridging reference resolution is an anaphora resolution task that is arguably more challenging and less studied than entity coreference resolution. Given that significant progress has been made on coreference resolution …

coreference-resolutionCoreference ResolutionSurvey