paper-with-me

Papers

Context-aware Multimodal AI Reveals Hidden Pathways in Five Centuries of Art Evolution

2025-03-15 · Jin Kim, Byunghwee Lee, Taekho You, Jinhyuk Yun

The rise of multimodal generative AI is transforming the intersection of technology and art, offering deeper insights into large-scale artwork. Although its creative capabilities have been widely explored, its potential to represent artwork in latent spaces remains underexamined. We use cutting-edge generative AI, specifically Stable Diffusion, to analyze 500 years of Western paintings by extracting two types of latent information with the model: formal aspects (e.g., colors) and contextual aspects (e.g., subject). Our findings reveal that contextual information differentiates between artistic periods, styles, and individual artists more successfully than formal elements. Additionally, using contextual keywords extracted from paintings, we show how artistic expression evolves alongside societal changes. Our generative experiment, infusing prospective contexts into historical artworks, successfully reproduces the evolutionary trajectory of artworks, highlighting the significance of mutual interaction between society and art. This study demonstrates how multimodal AI expands traditional formal analysis by integrating temporal, cultural, and historical contexts.

📄 PDF Abstract BibTeX arXiv:2503.13531

Code (1)

aljinny/art-history 공식 구현 pytorch

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

InEdit-Bench: Benchmarking Intermediate Logical Pathways for Intelligent Image Editing Models

2026-03-04 · Zhiqiang Sheng, Xumeng Han, Zhiwei Zhang, Zenghui Xiong 외 arxiv

Multimodal generative models have made significant strides in image editing, demonstrating impressive performance on a variety of static tasks. However, their proficiency typically does not extend to complex scenarios re…

Image Editing

Lance: Unified Multimodal Modeling by Multi-Task Synergy

2026-05-18 · Fengyi Fu, Mengqi Huang, Shaojin Wu, Yunsheng Jiang 외 arxiv

We present Lance, a lightweight native unified model supporting multimodal understanding, generation, and editing for both images and videos. Rather than relying on model capacity scaling or text-image-dominant designs, …

Video Generation

Segregate, Refine, Integrate: Decomposing Multimodal Fusion for Sentiment Analysis

2026-07-14 · Alexios Filippakopoulos, Elias Kallioras, Nikolaos Xiros, Efthymios Georgiou 외 arxiv

Multimodal fusion must simultaneously refine modality-specific signals and model cross-modal interactions; two competing objectives typically entangled within the same operation. We propose \textbf{SeRIn} (\textbf{Se}gre…

Sentiment Analysis

PATHWAYS: Evaluating Investigation and Context Discovery in AI Web Agents

2026-02-05 · Shifat E. Arman, Syed Nazmus Sakib, Tapodhir Karmakar Taton, Nafiul Haque 외 arxiv

We introduce PATHWAYS, a benchmark of 250 multi-step decision tasks that test whether web-based agents can discover and correctly use hidden contextual information. Across both closed and open models, agents typically na…

Saying the Unsaid: Revealing the Hidden Language of Multimodal Systems Through Telephone Games

2025-11-12 · Juntu Zhao, Jialing Zhang, Chongxuan Li, Dequan Wang arxiv

Recent closed-source multimodal systems have made great advances, but their hidden language for understanding the world remains opaque because of their black-box architectures. In this paper, we use the systems' preferen…