paper-with-me

홈 › Papers

MMIS: Multimodal Dataset for Interior Scene Visual Generation and Recognition

2024-07-08 · Hozaifa Kassab, Ahmed Mahmoud, Mohamed Bahaa, Ammar Mohamed, Ali Hamdi

We introduce MMIS, a novel dataset designed to advance MultiModal Interior Scene generation and recognition. MMIS consists of nearly 160,000 images. Each image within the dataset is accompanied by its corresponding textual description and an audio recording of that description, providing rich and diverse sources of information for scene generation and recognition. MMIS encompasses a wide range of interior spaces, capturing various styles, layouts, and furnishings. To construct this dataset, we employed careful processes involving the collection of images, the generation of textual descriptions, and corresponding speech annotations. The presented dataset contributes to research in multi-modal representation learning tasks such as image generation, retrieval, captioning, and classification.

📄 PDF Abstract BibTeX arXiv:2407.05980

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationRepresentation LearningRetrievalScene Generation

Similar Papers 제목 키워드 기반

Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation

2024-08-19 · Yunxin Li, Haoyuan Shi, Baotian Hu, Longyue Wang 외

Traditional animation generation methods depend on training generative models with human-labelled data, entailing a sophisticated multi-stage pipeline that demands substantial human effort and incurs high training costs.…

Image GenerationVideo Generation

Towards Multimodal Multitask Scene Understanding Models for Indoor Mobile Agents

2022-09-27 · Yao-Hung Hubert Tsai, Hanlin Goh, Ali Farhadi, Jian Zhang

The perception system in personalized mobile agents requires developing indoor scene understanding models, which can understand 3D geometries, capture objectiveness, analyze human behaviors, etc. Nonetheless, this direct…

3D Object DetectionAutonomous DrivingComputational EfficiencyDepth Completion+8

VIDES: Virtual Interior Design via Natural Language and Visual Guidance

2023-08-26 · Minh-Hien Le, Chi-Bien Chu, Khanh-Duy Le, Tam V. Nguyen 외

Interior design is crucial in creating aesthetically pleasing and functional indoor spaces. However, developing and editing interior design concepts requires significant time and expertise. We propose Virtual Interior DE…

Incorporating Eye-Tracking Signals Into Multimodal Deep Visual Models For Predicting User Aesthetic Experience In Residential Interiors

2026-01-23 · Chen-Ying Chien, Po-Chih Kuo arxiv

Understanding how people perceive and evaluate interior spaces is essential for designing environments that promote well-being. However, predicting aesthetic experiences remains difficult due to the subjective nature of …

What Looks Good with my Sofa: Multimodal Search Engine for Interior Design

2017-07-21 · Ivona Tautkute, Aleksandra Możejko, Wojciech Stokowiec, Tomasz Trzciński 외

In this paper, we propose a multi-modal search engine for interior design that combines visual and textual queries. The goal of our engine is to retrieve interior objects, e.g. furniture or wall clocks, that share visual…