paper-with-me

Papers

Multi-View Task-Driven Recognition in Visual Sensor Networks

2017-05-30 · Ali Taalimi, Alireza Rahimpour, Liu Liu, Hairong Qi

Nowadays, distributed smart cameras are deployed for a wide set of tasks in several application scenarios, ranging from object recognition, image retrieval, and forensic applications. Due to limited bandwidth in distributed systems, efficient coding of local visual features has in fact been an active topic of research. In this paper, we propose a novel approach to obtain a compact representation of high-dimensional visual data using sensor fusion techniques. We convert the problem of visual analysis in resource-limited scenarios to a multi-view representation learning, and we show that the key to finding properly compressed representation is to exploit the position of cameras with respect to each other as a norm-based regularization in the particular signal representation of sparse coding. Learning the representation of each camera is viewed as an individual task and a multi-task learning with joint sparsity for all nodes is employed. The proposed representation learning scheme is referred to as the multi-view task-driven learning for visual sensor network (MT-VSN). We demonstrate that MT-VSN outperforms state-of-the-art in various surveillance recognition tasks.

📄 PDF Abstract BibTeX arXiv:1705.10715

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalMulti-Task LearningObject RecognitionRepresentation LearningRetrievalSensor Fusion

Similar Papers 제목 키워드 기반

Context-driven Visual Object Recognition based on Knowledge Graphs

2022-10-20 · Sebastian Monka, Lavdim Halilaj, Achim Rettinger

Current deep learning methods for object recognition are purely data-driven and require a large number of training samples to achieve good results. Due to their sole dependence on image data, these methods tend to fail w…

Knowledge GraphsObjectObject RecognitionTransfer Learning

Visual Recognition-Driven Image Restoration for Multiple Degradation With Intrinsic Semantics Recovery

2023-01-01 · CVPR 2023 1 · Zizheng Yang, Jie Huang, Jiahao Chang, Man Zhou 외

Deep image recognition models suffer a significant performance drop when applied to low-quality images since they are trained on high-quality images. Although many studies have investigated to solve the issue through…

Domain AdaptationImage EnhancementImage RestorationPerson Re-Identification

A large scale multi-view RGBD visual affordance learning dataset

2022-03-26 · Zeyad Khalifa, Syed Afaq Ali Shah

The physical and textural attributes of objects have been widely studied for recognition, detection and segmentation tasks in computer vision.~A number of datasets, such as large scale ImageNet, have been proposed for fe…

Affordance RecognitionSegmentation

Combining Multiple Views for Visual Speech Recognition

2017-10-19 · Marina Zimmermann, Mostafa Mehdipour Ghazi, Hazim Kemal Ekenel, Jean-Philippe Thiran

Visual speech recognition is a challenging research problem with a particular practical application of aiding audio speech recognition in noisy scenarios. Multiple camera setups can be beneficial for the visual speech re…

Sentencespeech-recognitionSpeech RecognitionVisual Speech Recognition

OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset

2023-01-16 · Jeongkyun Park, Jung-Wook Hwang, Kwanghee Choi, Seung-Hyun Lee 외

Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, induce dependencies with various prediction models d…

Audio-Visual Speech RecognitionLip ReadingSpeaker Recognitionspeech-recognition+2