paper-with-me

Papers

Topological Perspectives on Optimal Multimodal Embedding Spaces

2024-05-29 · Abdul Aziz A. B, A. B Abdul Rahim

Recent strides in multimodal model development have ignited a paradigm shift in the realm of text-to-image generation. Among these advancements, CLIP stands out as a remarkable achievement which is a sophisticated autoencoder adept at encoding both textual and visual information within a unified latent space. This paper delves into a comparative analysis between CLIP and its recent counterpart, CLOOB. To unravel the intricate distinctions within the embedding spaces crafted by these models, we employ topological data analysis. Our approach encompasses a comprehensive examination of the modality gap drivers, the clustering structures existing across both high and low dimensions, and the pivotal role that dimension collapse plays in shaping their respective embedding spaces. Empirical experiments substantiate the implications of our analyses on downstream performance across various contextual scenarios. Through this investigation, we aim to shed light on the nuanced intricacies that underlie the comparative efficacy of CLIP and CLOOB, offering insights into their respective strengths and weaknesses, and providing a foundation for further refinement and advancement in multimodal model research.

📄 PDF Abstract BibTeX arXiv:2405.18867

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image GenerationTopological Data Analysis

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Topological Alignment of Shared Vision-Language Embedding Space

2025-10-13 · Junwon You, Dasol Kang, Jae-Hun Jung arxiv

Contrastive Vision-Language Models (VLMs) have demonstrated strong zero-shot capabilities. However, their cross-modal alignment remains biased toward English due to limited multilingual multimodal data. Recent multilingu…

Representation Learning

GraphMoRE: Mitigating Topological Heterogeneity via Mixture of Riemannian Experts

2024-12-15 · Zihao Guo, Qingyun Sun, Haonan Yuan, Xingcheng Fu 외

Real-world graphs have inherently complex and diverse topological patterns, known as topological heterogeneity. Most existing works learn graph representation in a single constant curvature space that is insufficient to …

From Topology to Retrieval: Decoding Embedding Spaces with Unified Signatures

2025-11-27 · Florian Rottach, William Rudman, Bastian Rieck, Harrisen Scells 외 arxiv

Studying how embeddings are organized in space not only enhances model interpretability but also uncovers factors that drive downstream task performance. In this paper, we present a comprehensive analysis of topological …

Dimensionality reduction and width of deep neural networks based on topological degree theory

2025-11-10 · Xiao-Song Yang arxiv

In this paper we present a mathematical framework on linking of embeddings of compact topological spaces into Euclidean spaces and separability of linked embeddings under a specific class of dimension reduction maps. As …

Dimensionality Reduction

Explainable Mapper: Charting LLM Embedding Spaces Using Perturbation-Based Explanation and Verification Agents

2025-07-24 · Xinyuan Yan, Rita Sevastjanova, Sinie van der Ben, Mennatallah El-Assady 외 arxiv

Large language models (LLMs) produce high-dimensional embeddings that capture rich semantic and syntactic relationships between words, sentences, and concepts. Investigating the topological structures of LLM embedding sp…