paper-with-me

홈 › Papers

Navigating User Behavior toward Personalized Multimodal Generation

2026-06-23 · Hengji Zhou, Yufeng Liu, Ye Liu, Yong Xu, Lianghao Xia, Liqiang Nie arxiv

Modern AIGC pipelines deliver high-fidelity images and videos but presuppose a well-formed creation instruction, while end users rarely articulate visual details, leaving generators misaligned with user demand. We study personalized content generation, which turns a user's interaction history into an executable instruction for downstream synthesis, and identify two obstacles: behavior must be encoded in a form legible to language reasoning, and the model must acquire instruction-writing skill absent from both pretraining and behavior data. We propose NaviGen, which represents each item with a dual identifier coupling a collaborative code and a textual code as a behavioral substrate and a semantic bridge in one token stream. On this representation, a two-stage SFT+RL pipeline first distills preference reasoning and instruction writing from evolutionarily searched supervision, then aligns generation with user intent through hierarchical and self-consistent rewards. Experiments across product, game, and short-video domains show that NaviGen improves personalized image and video generation, strengthens next-item prediction, and yields more specific, relevant, and visually generatable instructions. Our code is released at: https://github.com/iLearn-Lab/NaviGen.

📄 PDF Abstract BibTeX arXiv:2606.24196

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal generationVideo Generation

Similar Papers 제목 키워드 기반

PMG : Personalized Multimodal Generation with Large Language Models

2024-04-07 · Xiaoteng Shen, Rui Zhang, Xiaoyan Zhao, Jieming Zhu 외

The emergence of large language models (LLMs) has revolutionized the capabilities of text comprehension and generation. Multi-modal generation attracts great attention from both the industry and academia, but there is li…

multimodal generationReading ComprehensionRecommendation Systems

Personalized Image Generation with Large Multimodal Models

2024-10-18 · Yiyan Xu, Wenjie Wang, Yang Zhang, Biao Tang 외

Personalized content filtering, such as recommender systems, has become a critical infrastructure to alleviate information overload. However, these systems merely filter existing content and are constrained by its limite…

Image GenerationPersonalized Image GenerationRecommendation SystemsText Generation

TailorMind: Towards Preference-Aligned Multimodal Content Generation

2026-06-22 · Hengji Zhou, Ye Liu, Yufeng Liu, Si Wu 외 arxiv

Personalized content systems depend on available UGC and struggle when suitable content is absent, delayed, or costly to create. Although multimodal generators can synthesize content on demand, how to translate behaviora…

Collaborative Filteringmultimodal generation

Design Your Ad: Personalized Advertising Image and Text Generation with Unified Autoregressive Models

2026-05-12 · Yexing Xu, Wei Feng, Shen Zhang, Haohan Wang 외 arxiv

Generating realistic and user-preferred advertisements is a key challenge in e-commerce. Existing approaches utilize multiple independent models driven by click-through-rate (CTR) to controllably create attractive image …

Text Generation

Unified Personalized Understanding, Generating and Editing

2026-01-11 · Yu Zhong, Tianwei Lin, Ruike Zhu, Yuqian Yuan 외 arxiv

Unified large multimodal models (LMMs) have achieved remarkable progress in general-purpose multimodal understanding and generation. However, they still operate under a ``one-size-fits-all'' paradigm and struggle to mode…

Image Editing