paper-with-me

Papers

BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion

2024-04-06 · Gwanghyun Kim, Hayeon Kim, Hoigi Seo, Dong Un Kang, Se Young Chun

Generating higher-resolution human-centric scenes with details and controls remains a challenge for existing text-to-image diffusion models. This challenge stems from limited training image size, text encoder capacity (limited tokens), and the inherent difficulty of generating complex scenes involving multiple humans. While current methods attempted to address training size limit only, they often yielded human-centric scenes with severe artifacts. We propose BeyondScene, a novel framework that overcomes prior limitations, generating exquisite higher-resolution (over 8K) human-centric scenes with exceptional text-image correspondence and naturalness using existing pretrained diffusion models. BeyondScene employs a staged and hierarchical approach to initially generate a detailed base image focusing on crucial elements in instance creation for multiple humans and detailed descriptions beyond token limit of diffusion model, and then to seamlessly convert the base image to a higher-resolution output, exceeding training image size and incorporating details aware of text and instances via our novel instance-aware hierarchical enlargement process that consists of our proposed high-frequency injected forward diffusion and adaptive joint diffusion. BeyondScene surpasses existing methods in terms of correspondence with detailed text descriptions and naturalness, paving the way for advanced applications in higher-resolution human-centric scene creation beyond the capacity of pretrained diffusion models without costly retraining. Project page: https://janeyeon.github.io/beyond-scene.

📄 PDF Abstract BibTeX arXiv:2404.04544

Code (0)

등록된 구현이 없습니다.

Tasks

8kScene Generation

Methods 이 논문이 사용한 방법론

BASE 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
AWARE We propose to theoretically and empirically examine the effect of incorporating weighting schemes into walk-aggregating GNNs. To this end, we propose a simple, interpretable, and…

Similar Papers 제목 키워드 기반

Scene-Centric Unsupervised Panoptic Segmentation

2025-04-02 · CVPR 2025 1 · Oliver Hahn, Christoph Reich, Nikita Araslanov, Daniel Cremers 외

Unsupervised panoptic segmentation aims to partition an image into semantically meaningful regions and distinct object instances without training on manually annotated data. In contrast to prior work on unsupervised pano…

Instance SegmentationPanoptic SegmentationPseudo LabelScene Understanding+5

Towards Flexible 3D Perception: Object-Centric Occupancy Completion Augments 3D Object Detection

2024-12-06 · Chaoda Zheng, Feng Wang, Naiyan Wang, Shuguang Cui 외

While 3D object bounding box (bbox) representation has been widely used in autonomous driving perception, it lacks the ability to capture the precise details of an object's intrinsic geometry. Recently, occupancy has eme…

3D Object DetectionAutonomous DrivingObjectobject-detection+1

PANDA: A Gigapixel-level Human-centric Video Dataset

2020-03-10 · CVPR 2020 6 · Xueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo 외

We present PANDA, the first gigaPixel-level humAN-centric viDeo dAtaset, for large-scale, long-term, and multi-object visual analysis. The videos in PANDA were captured by a gigapixel camera and cover real-world scenes w…

4kAttributeHuman Detection

Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models

2025-09-30 · Yuansen Liu, Haiming Tang, Jinlong Peng, Jiangning Zhang 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely been explored, primarily due to the abs…

Scene Understanding

Exploring object-centric and scene-centric CNN features and their complementarity for human rights violations recognition in images

2018-05-12 · Grigorios Kalliatakis, Shoaib Ehsan, Ales Leonardis, Klaus McDonald-Maier

Identifying potential abuses of human rights through imagery is a novel and challenging task in the field of computer vision, that will enable to expose human rights violations over large-scale data that may otherwise be…

Representation LearningTransfer Learning