paper-with-me

Papers

Generating Object Stamps

2020-01-01 · Youssef Alami Mejjati, Zejiang Shen, Michael Snower, Aaron Gokaslan, Oliver Wang, James Tompkin, Kwang In Kim

We present an algorithm to generate diverse foreground objects and composite them into background images using a GAN architecture. Given an object class, a user-provided bounding box, and a background image, we first use a mask generator to create an object shape, and then use a texture generator to fill the mask such that the texture integrates with the background. By separating the problem of object insertion into these two stages, we show that our model allows us to improve the realism of diverse object generation that also agrees with the provided background image. Our results on the challenging COCO dataset show improved overall quality and diversity compared to state-of-the-art object insertion approaches.

📄 PDF Abstract BibTeX arXiv:2001.02595

Code (1)

AlamiMejjati/GeneratingObjectStamps 공식 구현 tf

Tasks

DiversityObject

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Dogecoin Customer Service Number +1-833-534-1729 설명 없음

Similar Papers 제목 키워드 기반

GenVDM: Generating Vector Displacement Maps From a Single Image

2025-03-01 · CVPR 2025 1 · Yuezhi Yang, Qimin Chen, Vladimir G. Kim, Siddhartha Chaudhuri 외

We introduce the first method for generating Vector Displacement Maps (VDMs): parameterized, detailed geometric stamps commonly used in 3D modeling. Given a single input image, our method first generates multi-view norma…

Encode Once, Decode Never: Reusing Audio LM Internals for Efficient Temporal Localization

2026-02-10 · Joesph An, Phillip Keung, Jiaqi Wang, Orevaoghene Ahia 외 arxiv

Audio language models process input audio into rich frame-level representations, but the standard approach to temporal localization generates timestamps as sequences of text tokens, which discards the frame-level represe…

Speaker DiarizationWord Alignment

Towards a Logic-Based Unifying Framework for Computing

2013-01-29 · Robert Kowalski, Fariba Sadri

In this paper we propose a logic-based, framework inspired by artificial intelligence, but scaled down for practical database and programming applications. Computation in the framework is viewed as the task of generating…

Predicting Long-horizon Futures by Conditioning on Geometry and Time

2024-04-17 · Tarasha Khurana, Deva Ramanan

Our work explores the task of generating future sensor observations conditioned on the past. We are motivated by `predictive coding' concepts from neuroscience as well as robotic applications such as self-driving vehicle…

Video Prediction

VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding

2025-04-10 · Henghao Zhao, Ge-Peng Ji, Rui Yan, Huan Xiong 외

The core challenge in video understanding lies in perceiving dynamic content changes over time. However, multimodal large language models struggle with temporal-sensitive video tasks, which requires generating timestamps…

Instruction FollowingVideo Understanding