paper-with-me

홈 › Papers

Toward Intelligent Scene Augmentation for Context-Aware Object Placement and Sponsor-Logo Integration

2025-12-25 · Unnati Saraswat, Tarun Rao, Namah Gupta, Shweta Swami, Shikhar Sharma, Prateek Narang, Dhruv Kumar arxiv

Intelligent image editing increasingly relies on advances in computer vision, multimodal reasoning, and generative modeling. While vision-language models (VLMs) and diffusion models enable guided visual manipulation, existing work rarely ensures that inserted objects are \emph{contextually appropriate}. We introduce two new tasks for advertising and digital media: (1) \emph{context-aware object insertion}, which requires predicting suitable object categories, generating them, and placing them plausibly within the scene; and (2) \emph{sponsor-product logo augmentation}, which involves detecting products and inserting correct brand logos, even when items are unbranded or incorrectly branded. To support these tasks, we build two new datasets with category annotations, placement regions, and sponsor-product labels.

📄 PDF Abstract BibTeX arXiv:2512.21560

Code (0)

등록된 구현이 없습니다.

Tasks

Multimodal ReasoningImage Editing

Similar Papers 제목 키워드 기반

Object-aware Contrastive Learning for Debiased Scene Representation

2021-07-30 · NeurIPS 2021 12 · Sangwoo Mo, Hyunwoo Kang, Kihyuk Sohn, Chun-Liang Li 외

Contrastive self-supervised learning has shown impressive results in learning visual representations from unlabeled images by enforcing invariance against different data augmentations. However, the learned representation…

Contrastive LearningObjectRepresentation LearningSelf-Supervised Learning

Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

2025-02-20 · Pengxiang Ding, Jianfei Ma, Xinyang Tong, Binghong Zou 외

This paper addresses the limitations of current humanoid robot control frameworks, which primarily rely on reactive mechanisms and lack autonomous interaction capabilities due to data scarcity. We propose Humanoid-VLA, a…

Data AugmentationHumanoid ControlMotion Generation

Point Cloud Recombination: Systematic Real Data Augmentation Using Robotic Targets for LiDAR Perception Validation

2025-05-05 · Hubert Padusinski, Christian Steinhauser, Christian Scherl, Julian Gaal 외

The validation of LiDAR-based perception of intelligent mobile systems operating in open-world applications remains a challenge due to the variability of real environmental conditions. Virtual simulations allow the gener…

Data Augmentation

Dynamic Context-Aware Scene Reasoning Using Vision-Language Alignment in Zero-Shot Real-World Scenarios

2025-10-30 · Manjunath Prasad Holenarasipura Rajiv, B. M. Vidyavathi arxiv

In real-world environments, AI systems often face unfamiliar scenarios without labeled data, creating a major challenge for conventional scene understanding models. The inability to generalize across unseen contexts limi…

Zero-shot GeneralizationScene Understanding

Intelligent Spatial Perception by Building Hierarchical 3D Scene Graphs for Indoor Scenarios with the Help of LLMs

2025-03-19 · Yao Cheng, Zhe Han, Fengyang Jiang, Huaizhen Wang 외

This paper addresses the high demand in advanced intelligent robot navigation for a more holistic understanding of spatial environments, by introducing a novel system that harnesses the capabilities of Large Language Mod…

ObjectRobot NavigationTask Planning