paper-with-me

홈 › Papers

UMAIR-FPS: User-aware Multi-modal Animation Illustration Recommendation Fusion with Painting Style

2024-02-16 · Yan Kang, Hao Lin, Mingjian Yang, Shin-Jye Lee

The rapid advancement of high-quality image generation models based on AI has generated a deluge of anime illustrations. Recommending illustrations to users within massive data has become a challenging and popular task. However, existing anime recommendation systems have focused on text features but still need to integrate image features. In addition, most multi-modal recommendation research is constrained by tightly coupled datasets, limiting its applicability to anime illustrations. We propose the User-aware Multi-modal Animation Illustration Recommendation Fusion with Painting Style (UMAIR-FPS) to tackle these gaps. In the feature extract phase, for image features, we are the first to combine image painting style features with semantic features to construct a dual-output image encoder for enhancing representation. For text features, we obtain text embeddings based on fine-tuning Sentence-Transformers by incorporating domain knowledge that composes a variety of domain text pairs from multilingual mappings, entity relationships, and term explanation perspectives, respectively. In the multi-modal fusion phase, we novelly propose a user-aware multi-modal contribution measurement mechanism to weight multi-modal features dynamically according to user features at the interaction level and employ the DCN-V2 module to model bounded-degree multi-modal crosses effectively. UMAIR-FPS surpasses the stat-of-the-art baselines on large real-world datasets, demonstrating substantial performance enhancements.

📄 PDF Abstract BibTeX arXiv:2402.10381

Code (1)

oysterqaq/umair-fps 공식 구현 tf

Tasks

Image GenerationMulti-modal RecommendationRecommendation SystemsSentence

Methods 이 논문이 사용한 방법론

DCN-V2 DCN-V2 is an architecture for learning-to-rank that improves upon the original DCN model. It first learns explicit feature interactions…

Similar Papers 제목 키워드 기반

EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation

2025-07-05 · Rang Meng, Yan Wang, Weipeng Wu, Ruobing Zheng 외 arxiv

Recent work on human animation usually incorporates large-scale video models, thereby achieving more vivid performance. However, the practical use of such methods is hindered by the slow inference speed and high computat…

Edge Federated Learning Via Unit-Modulus Over-The-Air Computation

2021-01-28 · Shuai Wang, Yuncong Hong, Rui Wang, Qi Hao 외

Edge federated learning (FL) is an emerging paradigm that trains a global parametric model from distributed datasets based on wireless communications. This paper proposes a unit-modulus over-the-air computation (UMAirCom…

Autonomous DrivingFederated Learning

Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal Space

2025-04-14 · Kangwei Liu, Junwu Liu, Xiaowei Yi, Jinlin Guo 외

Audio-driven emotional 3D facial animation encounters two significant challenges: (1) reliance on single-modal control signals (videos, text, or emotion labels) without leveraging their complementary strengths for compre…

Contrastive LearningDiversity

Allo-AVA: A Large-Scale Multimodal Conversational AI Dataset for Allocentric Avatar Gesture Animation

2024-10-21 · Saif Punjwani, Larry Heck

The scarcity of high-quality, multimodal training data severely hinders the creation of lifelike avatar animations for conversational AI in virtual environments. Existing datasets often lack the intricate synchronization…

PMMTalk: Speech-Driven 3D Facial Animation from Complementary Pseudo Multi-modal Features

2023-12-05 · Tianshun Han, Shengnan Gui, Yiqing Huang, Baihui Li 외

Speech-driven 3D facial animation has improved a lot recently while most related works only utilize acoustic modality and neglect the influence of visual and textual cues, leading to unsatisfactory results in terms of pr…

cross-modal alignmentDecoderspeech-recognitionSpeech Recognition+1