paper-with-me

Papers

Conjuring Semantic Similarity

2024-10-21 · Tian Yu Liu, Stefano Soatto

The semantic similarity between sample expressions measures the distance between their latent 'meaning'. Such meanings are themselves typically represented by textual expressions, often insufficient to differentiate concepts at fine granularity. We propose a novel approach whereby the semantic similarity among textual expressions is based not on other expressions they can be rephrased as, but rather based on the imagery they evoke. While this is not possible with humans, generative models allow us to easily visualize and compare generated images, or their distribution, evoked by a textual prompt. Therefore, we characterize the semantic similarity between two textual expressions simply as the distance between image distributions they induce, or 'conjure.' We show that by choosing the Jensen-Shannon divergence between the reverse-time diffusion stochastic differential equations (SDEs) induced by each textual expression, this can be directly computed via Monte-Carlo sampling. Our method contributes a novel perspective on semantic similarity that not only aligns with human-annotated scores, but also opens up new avenues for the evaluation of text-conditioned generative models while offering better interpretability of their learnt representations.

📄 PDF Abstract BibTeX arXiv:2410.16431

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic SimilaritySemantic Textual Similarity

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Conjuring Positive Pairs for Efficient Unification of Representation Learning and Image Synthesis

2025-03-19 · Imanol G. Estepa, Jesús M. Rodríguez-de-Vera, Ignacio Sarasúa, Bhalaji Nagarajan 외

While representation learning and generative modeling seek to understand visual data, unifying both domains remains unexplored. Recent Unified Self-Supervised Learning (SSL) methods have started to bridge the gap between…

Few-Shot LearningImage GenerationRepresentation LearningSelf-Supervised Learning+2

Latent semantics of action verbs reflect phonetic parameters of intensity and emotional content

2014-05-06 · Michael Kai Petersen

Conjuring up our thoughts, language reflects statistical patterns of word co-occurrences which in turn come to describe how we perceive the world. Whether counting how frequently nouns and verbs combine in Google search …

Clustering

Social Conjuring: Multi-User Runtime Collaboration with AI in Building Virtual 3D Worlds

2024-09-30 · Amina Kobenova, Cyan DeVeaux, Samyak Parajuli, Andrzej Banburski-Fahey 외

Generative artificial intelligence has shown promise in prompting virtual worlds into existence, yet little attention has been given to understanding how this process unfolds as social interaction. We present Social Conj…

Spatial Reasoning

LivePhoto: Real Image Animation with Text-guided Motion Control

2023-12-05 · Xi Chen, Zhiheng Liu, Mengting Chen, Yutong Feng 외

Despite the recent progress in text-to-video generation, existing studies usually overlook the issue that only spatial contents but not temporal motions in synthesized videos are under the control of text. Towards such a…

Image AnimationText-to-Video GenerationVideo Generation

Semantic similarity prediction is better than other semantic similarity measures

2023-09-22 · Steffen Herbold

Semantic similarity between natural language texts is typically measured either by looking at the overlap between subsequences (e.g., BLEU) or by using embeddings (e.g., BERTScore, S-BERT). Within this paper, we argue th…

Semantic SimilaritySemantic Textual SimilaritySTSSTS-B