Multi-Glimpse LSTM with Color-Depth Feature Fusion for Human Detection
With the development of depth cameras such as Kinect and Intel Realsense, RGB-D based human detection receives continuous research attention due to its usage in a variety of applications. In this paper, we propose a new Multi-Glimpse LSTM (MG-LSTM) network, in which multi-scale contextual information is sequentially integrated to promote the human detection performance. Furthermore, we propose a feature fusion strategy based on our MG-LSTM network to better incorporate the RGB and depth information. To the best of our knowledge, this is the first attempt to utilize LSTM structure for RGB-D based human detection. Our method achieves superior performance on two publicly available datasets.
Code (0)
등록된 구현이 없습니다.
Tasks
Human DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Learning of Colors from Color Names: Distribution and Point Estimation
Color names are often made up of multiple words. As a task in natural language understanding we investigate in depth the capacity of neural networks based on sums of word embeddings (SOWE), recurrence (LSTM and GRU based…
Natural Language UnderstandingWord EmbeddingsGlimpse Clouds: Human Activity Recognition from Unstructured Feature Points
We propose a method for human activity recognition from RGB data that does not rely on any pose information during test time and does not explicitly calculate pose information internally. Instead, a visual attention modu…
Action RecognitionActivity PredictionActivity RecognitionHuman Activity Recognition+2GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
Existing video benchmarks often resemble image-based benchmarks, with question types like "What actions does the person perform throughout the video?" or "What color is the woman's dress in the video?" For these, models …
Representation and Correlation Enhanced Encoder-Decoder Framework for Scene Text Recognition
Attention-based encoder-decoder framework is widely used in the scene text recognition task. However, for the current state-of-the-art(SOTA) methods, there is room for improvement in terms of the efficient usage of local…
DecoderScene Text RecognitionAction Anticipation at a Glimpse: To What Extent Can Multimodal Cues Replace Video?
Anticipating actions before they occur is a core challenge in action understanding research. While conventional methods rely on extracting and aggregating temporal information from videos, as humans we can often predict …
Action UnderstandingAction AnticipationSpatial Reasoning