paper-with-me

Papers

InstaIndoor and Multi-modal Deep Learning for Indoor Scene Recognition

2021-12-23 · Andreea Glavan, Estefania Talavera

Indoor scene recognition is a growing field with great potential for behaviour understanding, robot localization, and elderly monitoring, among others. In this study, we approach the task of scene recognition from a novel standpoint, using multi-modal learning and video data gathered from social media. The accessibility and variety of social media videos can provide realistic data for modern scene recognition techniques and applications. We propose a model based on fusion of transcribed speech to text and visual features, which is used for classification on a novel dataset of social media videos of indoor scenes named InstaIndoor. Our model achieves up to 70% accuracy and 0.7 F1-Score. Furthermore, we highlight the potential of our approach by benchmarking on a YouTube-8M subset of indoor scenes as well, where it achieves 74% accuracy and 0.74 F1-Score. We hope the contributions of this work pave the way to novel research in the challenging field of indoor scene recognition.

📄 PDF Abstract BibTeX arXiv:2112.12409

Code (1)

andreea-glavan/multimodal-audiovisual-scene-recognition 공식 구현 tf

Tasks

BenchmarkingDeep LearningScene RecognitionSpeech-to-Text

Similar Papers 제목 키워드 기반

Dynamic Graph Neural Network with Adaptive Features Selection for RGB-D Based Indoor Scene Recognition

2026-04-01 · Qiong Liu, Ruofei Xiong, Xingzhen Chen, Muyao Peng 외 arxiv

Multi-modality of color and depth, i.e., RGB-D, is of great importance in recent research of indoor scene recognition. In this kind of data representation, depth map is able to describe the 3D structure of scenes and geo…

Graph Neural NetworkScene Recognition

Indoor scene recognition from images under visual corruptions

2024-08-23 · Willams de Lima Costa, Raul Ismayilov, Nicola Strisciuglio, Estefania Talavera Martinez

The classification of indoor scenes is a critical component in various applications, such as intelligent robotics for assistive living. While deep learning has significantly advanced this field, models often suffer from …

Scene Recognition

Towards Multimodal Multitask Scene Understanding Models for Indoor Mobile Agents

2022-09-27 · Yao-Hung Hubert Tsai, Hanlin Goh, Ali Farhadi, Jian Zhang

The perception system in personalized mobile agents requires developing indoor scene understanding models, which can understand 3D geometries, capture objectiveness, analyze human behaviors, etc. Nonetheless, this direct…

3D Object DetectionAutonomous DrivingComputational EfficiencyDepth Completion+8

RoomPilot: Controllable Indoor Scene Synthesis via Multimodal Semantic Parsing

2025-12-12 · Wentang Chen, Shougao Zhang, Yiman Zhang, Tianhao Zhou 외 arxiv

Generating controllable indoor scenes is fundamental to applications in game development, architectural visualization, and embodied AI. However, existing approaches either support a limited input modalities or rely on im…

Indoor Scene SynthesisScene GenerationSemantic Parsing

A Discriminative Representation of Convolutional Features for Indoor Scene Recognition

2015-06-17 · Salman H. Khan, Munawar Hayat, Mohammed Bennamoun, Roberto Togneri 외

Indoor scene recognition is a multi-faceted and challenging problem due to the diverse intra-class variations and the confusing inter-class similarities. This paper presents a novel approach which exploits rich mid-level…

ObjectObject RecognitionScene ClassificationScene Recognition