paper-with-me

Papers

Fashion Matrix: Editing Photos by Just Talking

2023-07-25 · Zheng Chong, Xujie Zhang, Fuwei Zhao, Zhenyu Xie, Xiaodan Liang

The utilization of Large Language Models (LLMs) for the construction of AI systems has garnered significant attention across diverse fields. The extension of LLMs to the domain of fashion holds substantial commercial potential but also inherent challenges due to the intricate semantic interactions in fashion-related generation. To address this issue, we developed a hierarchical AI system called Fashion Matrix dedicated to editing photos by just talking. This system facilitates diverse prompt-driven tasks, encompassing garment or accessory replacement, recoloring, addition, and removal. Specifically, Fashion Matrix employs LLM as its foundational support and engages in iterative interactions with users. It employs a range of Semantic Segmentation Models (e.g., Grounded-SAM, MattingAnything, etc.) to delineate the specific editing masks based on user instructions. Subsequently, Visual Foundation Models (e.g., Stable Diffusion, ControlNet, etc.) are leveraged to generate edited images from text prompts and masks, thereby facilitating the automation of fashion editing processes. Experiments demonstrate the outstanding ability of Fashion Matrix to explores the collaborative potential of functionally diverse pre-trained models in the domain of fashion editing.

📄 PDF Abstract BibTeX arXiv:2307.13240

Code (1)

zheng-chong/fashionmatric 공식 구현 pytorch

Tasks

Semantic Segmentation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head Generation

2024-05-12 · Changpeng Cai, Guinan Guo, Jiao Li, Junhao Su 외

Most earlier researches on talking face generation have focused on the synchronization of lip motion and speech content. However, head pose and facial emotions are equally important characteristics of natural faces. Whil…

DisentanglementFace GenerationTalking Face GenerationTalking Head Generation

Text-based Talking Video Editing with Cascaded Conditional Diffusion

2024-07-20 · Bo Han, Heqing Zou, Haoyang Li, Guangcong Wang 외

Text-based talking-head video editing aims to efficiently insert, delete, and substitute segments of talking videos through a user-friendly text editing approach. It is challenging because of \textbf{1)} generalizable ta…

Video Editing

FacEDiT: Unified Talking Face Editing and Generation via Facial Motion Infilling

2025-12-16 · Kim Sung-Bin, Joohyun Chang, David Harwath, Tae-Hyun Oh arxiv

Talking face editing and face generation have often been studied as distinct problems. In this work, we propose viewing both not as separate tasks but as subtasks of a unifying formulation, speech-conditional facial moti…

Talking Face Generation

How To Extract Fashion Trends From Social Media? A Robust Object Detector With Support For Unsupervised Learning

2018-06-28 · Vijay Gabale, Anand Prabhu Subramanian

With the proliferation of social media, fashion inspired from celebrities, reputed designers as well as fashion influencers has shortened the cycle of fashion design and manufacturing. However, with the explosion of fash…

Objectobject-detectionObject Detection

Instruct-NeuralTalker: Editing Audio-Driven Talking Radiance Fields with Instructions

2023-06-19 · Yuqi Sun, Ruian He, Weimin Tan, Bo Yan

Recent neural talking radiance field methods have shown great success in photorealistic audio-driven talking face synthesis. In this paper, we propose a novel interactive framework that utilizes human instructions to edi…

Face GenerationTalking Face Generation