Deciphering Compatibility Relationships with Textual Descriptions via Extraction and Explanation
Understanding and accurately explaining compatibility relationships between fashion items is a challenging problem in the burgeoning domain of AI-driven outfit recommendations. Present models, while making strides in this area, still occasionally fall short, offering explanations that can be elementary and repetitive. This work aims to address these shortcomings by introducing the Pair Fashion Explanation (PFE) dataset, a unique resource that has been curated to illuminate these compatibility relationships. Furthermore, we propose an innovative two-stage pipeline model that leverages this dataset. This fine-tuning allows the model to generate explanations that convey the compatibility relationships between items. Our experiments showcase the model's potential in crafting descriptions that are knowledgeable, aligned with ground-truth matching correlations, and that produce understandable and informative descriptions, as assessed by both automatic metrics and human evaluation. Our code and data are released at https://github.com/wangyu-ustc/PairFashionExplanation
Code (1)
Similar Papers 제목 키워드 기반
Inductive learning for product assortment graph completion
Global retailers have assortments that contain hundreds of thousands of products that can be linked by several types of relationships like style compatibility, "bought together", "watched together", etc. Graphs are a nat…
Inductive LearningFast Contextual Scene Graph Generation With Unbiased Context Augmentation
Scene graph generation (SGG) methods have historically suffered from long-tail bias and slow inference speed. In this paper, we notice that humans can analyze relationships between objects relying solely on context d…
Graph GenerationScene Graph GenerationExpert Knowledge & Machine Understanding: Bridging Reactome's Ontology with LLM Semantic Embeddings
Biological knowledgebases like Reactome provide high-quality pathways that include biological elements' relationships and textual descriptions (metadata). The quality of such pathways is granted by manual curation, that …
Probing Language Models' Gesture Understanding for Enhanced Human-AI Interaction
The rise of Large Language Models (LLMs) has affected various disciplines that got beyond mere text generation. Going beyond their textual nature, this project proposal aims to investigate the interaction between LLMs an…
Probing Language ModelsText GenerationComplete 3d relationships extraction modality alignment network for 3d dense captioning
3D dense captioning aims to semantically describe each object detected in a 3D scene, which plays a significant role in 3D scene understanding. Previous works lack a complete definition of 3D spatial relationships and th…
3D dense captioning3D Object DetectionDense Captioningobject-detection+2