paper-with-me

홈 › Papers

ANCHOR: LLM-driven News Subject Conditioning for Text-to-Image Synthesis

2024-04-15 · Aashish Anantha Ramakrishnan, Sharon X. Huang, Dongwon Lee

Text-to-Image (T2I) Synthesis has made tremendous strides in enhancing synthesized image quality, but current datasets evaluate model performance only on descriptive, instruction-based prompts. Real-world news image captions take a more pragmatic approach, providing high-level situational and Named-Entity (NE) information and limited physical object descriptions, making them abstractive. To evaluate the ability of T2I models to capture intended subjects from news captions, we introduce the Abstractive News Captions with High-level cOntext Representation (ANCHOR) dataset, containing 70K+ samples sourced from 5 different news media organizations. With Large Language Models (LLM) achieving success in language and commonsense reasoning tasks, we explore the ability of different LLMs to identify and understand key subjects from abstractive captions. Our proposed method Subject-Aware Finetuning (SAFE), selects and enhances the representation of key subjects in synthesized images by leveraging LLM-generated subject weights. It also adapts to the domain distribution of news images and captions through custom Domain Fine-tuning, outperforming current T2I baselines on ANCHOR. By launching the ANCHOR dataset, we hope to motivate research in furthering the Natural Language Understanding (NLU) capabilities of T2I models.

📄 PDF Abstract BibTeX arXiv:2404.10141

Code (1)

aashish2000/anchor 공식 구현

Tasks

DescriptiveImage CaptioningImage GenerationNatural Language Understanding

Similar Papers 제목 키워드 기반

FastComposer: Tuning-Free Multi-Subject Image Generation with Localized Attention

2023-05-17 · Guangxuan Xiao, Tianwei Yin, William T. Freeman, Frédo Durand 외

Diffusion models excel at text-to-image generation, especially in subject-driven generation for personalized images. However, existing methods are inefficient due to the subject-specific fine-tuning, which is computation…

DenoisingDiffusion PersonalizationDiffusion Personalization Tuning FreeImage Generation+3

Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation

2026-05-25 · Shuhong Zheng, Aashish Kumar Misraa, Yu-Teng Li, Yu-Jhe Li 외 arxiv

Subject-driven image generation aims to synthesize new images that preserve the identity of the given subject while following textual instructions. Existing approaches often encode text and reference images separately. T…

Instruction FollowingImage Generation

AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories

2026-02-16 · Zun Wang, Han Lin, Jaehong Yoon, Jaemin Cho 외 arxiv

Maintaining spatial world consistency over long horizons remains a central challenge for camera-controllable video generation. Existing memory-based approaches often condition generation on globally reconstructed 3D scen…

Depth EstimationVideo Generation

3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory

2025-12-22 · Xinyang Song, Libin Wang, Weining Wang, Zhiwei Li 외 arxiv

Recent image generation approaches often address subject, style, and structure-driven conditioning in isolation, leading to feature entanglement and limited task transferability. In this paper, we introduce 3SGen, a task…

Image Generation

MagicComp: Training-free Dual-Phase Refinement for Compositional Video Generation

2025-03-18 · Hongyu Zhang, Yufan Deng, Shenghai Yuan, Peng Jin 외

Text-to-video (T2V) generation has made significant strides with diffusion models. However, existing methods still struggle with accurately binding attributes, determining spatial relationships, and capturing complex act…

DenoisingVideo Generation