paper-with-me

홈 › Papers

TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models

2026-01-21 · Carolin Holtermann, Nina Krebs, Anne Lauscher arxiv

Time alters the visual appearance of entities in our world, like objects, places, and animals. Thus, for accurately generating contextually-relevant images, knowledge and reasoning about time can be crucial (e.g., for generating a landscape in spring vs. in winter). Yet, although substantial work exists on understanding and improving temporal knowledge in natural language processing, research on how temporal phenomena appear and are handled in text-to-image (T2I) models remains scarce. We address this gap with TempViz, the first data set to holistically evaluate temporal knowledge in image generation, consisting of 7.9k prompts and more than 600 reference images. Using TempViz, we study the capabilities of five T2I models across five temporal knowledge categories. Human evaluation shows that temporal competence is generally weak, with no model exceeding 75% accuracy across categories. Towards larger-scale studies, we also examine automated evaluation methods, comparing several established approaches against human judgments. However, none of these approaches provides a reliable assessment of temporal cues - further indicating the pressing need for future research on temporal knowledge in T2I.

📄 PDF Abstract BibTeX arXiv:2601.14951

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights

2025-12-01 · Juanxi Tian, Siyuan Li, Conghui He, Lijun Wu 외 arxiv

Current multimodal models aim to transcend the limitations of single-modality representations by unifying understanding and generation, often using text-to-image (T2I) tasks to calibrate semantic consistency. However, th…

Image Generation

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

2025-03-10 · Yuwei Niu, Munan Ning, Mengren Zheng, Bin Lin 외

Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content. However, existing research and evaluation standards predominantly focus on image realism and shallow text-image ali…

Common Sense ReasoningImage GenerationText to Image GenerationText-to-Image Generation+1

Beyond Known Facts: Generating Unseen Temporal Knowledge to Address Data Contamination in LLM Evaluation

2026-01-20 · Arthur Amalvy, Hen-Hsen Huang arxiv

The automatic extraction of information is important for populating large web knowledge bases such as Wikidata. The temporal version of that task, temporal knowledge graph extraction (TKGE), involves extracting temporall…

Text Generation

Temporal Knowledge-Aware Image Captioning

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Contextualized image captioning is a task that extends beyond generating a purely visual description of the image content and aims to produce a caption that is influenced by the context and informed by the real world kno…

Caption GenerationImage CaptioningWorld Knowledge

Revisiting Temporal Modeling for CLIP-based Image-to-Video Knowledge Transferring

2023-01-26 · CVPR 2023 1 · Ruyang Liu, Jingjia Huang, Ge Li, Jiashi Feng 외

Image-text pretrained models, e.g., CLIP, have shown impressive general multi-modal knowledge learned from large-scale image-text data pairs, thus attracting increasing attention for their potential to improve visual rep…

Representation LearningRetrievalText RetrievalVideo Recognition+2