paper-with-me

홈 › Papers

On Geometrical Properties of Text Token Embeddings for Strong Semantic Binding in Text-to-Image Generation

2025-03-29 · Hoigi Seo, Junseo Bang, Haechang Lee, Joohoon Lee, Byung Hyun Lee, Se Young Chun

Text-to-Image (T2I) models often suffer from text-image misalignment in complex scenes involving multiple objects and attributes. Semantic binding aims to mitigate this issue by accurately associating the generated attributes and objects with their corresponding noun phrases (NPs). Existing methods rely on text or latent optimizations, yet the factors influencing semantic binding remain underexplored. Here we investigate the geometrical properties of text token embeddings and their cross-attention (CA) maps. We empirically and theoretically analyze that the geometrical properties of token embeddings, specifically both angular distances and norms, play a crucial role in CA map differentiation. Then, we propose \textbf{TeeMo}, a training-free text embedding-aware T2I framework with strong semantic binding. TeeMo consists of Causality-Aware Projection-Out (CAPO) for distinct inter-NP CA maps and Adaptive Token Mixing (ATM) with our loss to enhance inter-NP separation while maintaining intra-NP cohesion in CA maps. Extensive experiments confirm TeeMo consistently outperforms prior arts across diverse baselines and datasets.

📄 PDF Abstract BibTeX arXiv:2503.23011

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Improving OCR-Based Image Captioning by Incorporating Geometrical Relationship

2021-06-19 · CVPR 2021 1 · Jing Wang, Jinhui Tang, Mingkun Yang, Xiang Bai 외

OCR-based image captioning aims to automatically describe images based on all the visual entities (both visual objects and scene text) in images. Compared with conventional image captioning, the reasoning of scene te…

Image CaptioningOptical Character Recognition (OCR)Relation

Convergent Evolution: How Different Language Models Learn Similar Number Representations

2026-04-22 · Deqing Fu, Tianyi Zhou, Mikhail Belkin, Vatsal Sharan 외 arxiv

Language models trained on natural text learn to represent numbers using periodic features with dominant periods at $T=2, 5, 10$. In this paper, we identify a two-tiered hierarchy of these features: while Transformers, L…

LIME: Making LLM Data More Efficient with Linguistic Metadata Embeddings

2025-12-08 · Sebastian Sztwiertnia, Felix Friedrich, Kristian Kersting, Patrick Schramowski 외 arxiv

Pre-training decoder-only language models relies on vast amounts of high-quality data, yet the availability of such data is increasingly reaching its limits. While metadata is commonly used to create and curate these dat…

Monitoring geometrical properties of word embeddings for detecting the emergence of new topics

2021-11-05 · Clément Christophe, Julien Velcin, Jairo Cugliari, Manel Boumghar 외

Slow emerging topic detection is a task between event detection, where we aggregate behaviors of different words on short period of time, and language evolution, where we monitor their long term evolution. In this work, …

ArticlesEvent DetectionWord Embeddings

Monitoring geometrical properties of word embeddings for detecting the emergence of new topics.

2021-11-01 · EMNLP 2021 11 · Clément Christophe, Julien Velcin, Jairo Cugliari, Manel Boumghar 외

Slow emerging topic detection is a task between event detection, where we aggregate behaviors of different words on short period of time, and language evolution, where we monitor their long term evolution. In this work, …

ArticlesEvent DetectionWord Embeddings