paper-with-me

Papers

MeronymNet: A Hierarchical Approach for Unified and Controllable Multi-Category Object Generation

2021-10-17 · Rishabh Baghel, Abhishek Trivedi, Tejas Ravichandran, Ravi Kiran Sarvadevabhatla

We introduce MeronymNet, a novel hierarchical approach for controllable, part-based generation of multi-category objects using a single unified model. We adopt a guided coarse-to-fine strategy involving semantically conditioned generation of bounding box layouts, pixel-level part layouts and ultimately, the object depictions themselves. We use Graph Convolutional Networks, Deep Recurrent Networks along with custom-designed Conditional Variational Autoencoders to enable flexible, diverse and category-aware generation of 2-D objects in a controlled manner. The performance scores for generated objects reflect MeronymNet's superior performance compared to multiple strong baselines and ablative variants. We also showcase MeronymNet's suitability for controllable object generation and interactive object editing at various levels of structural and semantic granularity.

📄 PDF Abstract BibTeX arXiv:2110.08818

Code (1)

atmacvit/meronymnet 공식 구현 pytorch

Tasks

Object

Similar Papers 제목 키워드 기반

UniDGF: A Unified Detection-to-Generation Framework for Hierarchical Object Visual Recognition

2025-11-20 · Xinyu Nan, Lingtao Mao, Huangyu Dai, Zexin Zheng 외 arxiv

Achieving visual semantic understanding requires a unified framework that simultaneously handles object detection, category prediction, and attribute recognition. However, current advanced approaches rely on global simil…

Object Detection

MMGen: Unified Multi-modal Image Generation and Understanding in One Go

2025-03-26 · Jiepeng Wang, Zhaoqing Wang, Hao Pan, YuAn Liu 외

A unified diffusion framework for multi-modal generation and understanding has the transformative potential to achieve seamless and controllable image diffusion and other cross-modal tasks. In this paper, we introduce MM…

Image Generation

Omni-Effects: Unified and Spatially-Controllable Visual Effects Generation

2025-08-11 · Fangyuan Mao, Aiming Hao, Jintao Chen, Dongxia Liu 외 arxiv

Visual effects (VFX) are essential visual enhancements fundamental to modern cinematic production. Although video generation models offer cost-efficient solutions for VFX production, current methods are constrained by pe…

Video GenerationImage Editing

Deep Hierarchical Classification for Category Prediction in E-commerce System

2020-05-14 · WS 2020 7 · Dehong Gao, Wenjing Yang, Huiling Zhou, Yi Wei 외

In e-commerce system, category prediction is to automatically predict categories of given texts. Different from traditional classification where there are no relations between classes, category prediction is reckoned as …

ClassificationGeneral ClassificationPrediction

A Unified Model with Structured Output for Fashion Images Classification

2018-06-25 · Beatriz Quintino Ferreira, Luís Baía, João Faria, Ricardo Gamelas Sousa

A picture is worth a thousand words. Albeit a clich\'e, for the fashion industry, an image of a clothing piece allows one to perceive its category (e.g., dress), sub-category (e.g., day dress) and properties (e.g., white…

ClassificationGeneral ClassificationSpecificity