paper-with-me

Papers

CogCanvas: A Benchmark for Evaluating Multi-Subject Reference-Based Image Generation

2026-06-14 · Long-Bao Nguyen, Quang-Khai Tran, Tam V. Nguyen, Minh-Triet Tran, Trung-Nghia Le arxiv

Multi-subject reference-based image generation requires jointly preserving multiple human identities, binding per-person objects and fashion items, and respecting a specified background scene, a regime where current diffusion models remain brittle. Existing benchmarks evaluate only one axis at a time and none jointly captures multi-identity composition with human-object interaction, background grounding, and spatial plausibility. We introduce CogCanvas, a benchmark of 1,952 curated reference images spanning 100 celebrity identities, 115 distinctive objects and fashion items, and 29 real-world background scenes including landmarks, from which we construct 1,361 compositional prompts covering 2-5 person group sizes. The curation pipeline combines DINOv2-based deduplication, two-stage aesthetic filtering, and automated derivation of structured interaction and position graphs that serve as ground-truth supervision. CogCanvas supports three tasks, reference-based multi-human-object generation (primary), text-to-image compositional generation, and reference retrieval, under a unified six-axis evaluation protocol. We introduce two metrics tailored to the multi-reference setting: BG-Sim, which scores background fidelity on SAM 3-masked regions via DINOv3 feature similarity, and Attr-VQA, which uses a multimodal LLM to verify per-subject attribute binding and inter-person interactions against the structured graphs. Benchmarking five SOTA methods reveals that every model degrades substantially as group size grows from 2 to 5, with near-complete failure on object/fashion binding beyond three subjects.

📄 PDF Abstract BibTeX arXiv:2606.15867

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

F-Eval: Assessing Fundamental Abilities with Refined Evaluation Methods

2024-01-26 · Yu Sun, Keyu Chen, Shujie Wang, Peiji Li 외

Large language models (LLMs) garner significant attention for their unprecedented performance, leading to an increasing number of researches evaluating LLMs. However, these evaluation benchmarks are limited to assessing …

Instruction Following

DEAR: Dataset for Evaluating the Aesthetics of Rendering

2025-12-04 · Vsevolod Plohotnuk, Artyom Panshin, Nikola Banić, Simone Bianco 외 arxiv

Traditional Image Quality Assessment~(IQA) focuses on quantifying technical degradations such as noise, blur, or compression artifacts, using both full-reference and no-reference objective metrics. However, evaluation of…

Image Quality Assessment

Multi-subject Open-set Personalization in Video Generation

2025-01-10 · CVPR 2025 1 · Tsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Yuwei Fang 외

Video personalization methods allow us to synthesize videos with specific concepts such as people, pets, and places. However, existing methods often focus on limited domains, require time-consuming optimization per subje…

Video Generation

Benchmarking Music Generation Models and Metrics via Human Preference Studies

2025-06-23 · Audio Imagination: NeurIPS 2024 Workshop 2024 10 · Florian Grötschla, Ahmet Solak, Luca A. Lanzendörfer, Roger Wattenhofer

Recent advancements have brought generated music closer to human-created compositions, yet evaluating these models remains challenging. While human preference is the gold standard for assessing quality, translating these…

BenchmarkingMusic Generation

Full Reference Objective Quality Assessment for Reconstructed Background Images

2018-03-12 · Aditee Shrotre, Lina Karam

With an increased interest in applications that require a clean background image, such as video surveillance, object tracking, street view imaging and location-based services on web-based maps, multiple algorithms have b…

Image Quality AssessmentObject Tracking