paper-with-me

Papers

Learning a Convolutional Bilinear Sparse Code for Natural Videos

2019-09-11 · NeurIPS Workshop Neuro_AI 2019 12 · Dimitrios C. Gklezakos, Rajesh P. N. Rao

In contrast to the monolithic deep architectures used in deep learning today for computer vision, the visual cortex processes retinal images via two functionally distinct but interconnected networks: the ventral pathway for processing object-related information and the dorsal pathway for processing motion and transformations. Inspired by this cortical division of labor and properties of the magno- and parvocellular systems, we explore an unsupervised approach to feature learning that jointly learns object features and their transformations from natural videos. We propose a new convolutional bilinear sparse coding model that (1) allows independent feature transformations and (2) is capable of processing large images. Our learning procedure leverages smooth motion in natural videos. Our results show that our model can learn groups of features and their transformations directly from natural videos in a completely unsupervised manner. The learned "dynamic filters" exhibit certain equivariance properties, resemble cortical spatiotemporal filters, and capture the statistics of transitions between video frames. Our model can be viewed as one of the first approaches to demonstrate unsupervised learning of primary "capsules" (proposed by Hinton and colleagues for supervised learning) and has strong connections to the Lie group approach to visual perception.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Multiregion Bilinear Convolutional Neural Networks for Person Re-Identification

2015-12-16 · Evgeniya Ustinova, Yaroslav Ganin, Victor Lempitsky

In this work we propose a new architecture for person re-identification. As the task of re-identification is inherently associated with embedding learning and non-rigid appearance description, our architecture is based o…

Person Re-Identification

Submanifold Sparse Convolutional Networks

2017-06-05 · Benjamin Graham, Laurens van der Maaten

Convolutional network are the de-facto standard for analysing spatio-temporal data such as images, videos, 3D shapes, etc. Whilst some of this data is naturally dense (for instance, photos), many other data sources are i…

3D Part Segmentation

BiSparse-AAS: Bilinear Sparse Attention and Adaptive Spans Framework for Scalable and Efficient Text Summarization

2025-10-31 · Desta Haileselassie Hagos, Legand L. Burge, Anietie Andy, Anis Yazidi 외 arxiv

Transformer-based architectures have advanced text summarization, yet their quadratic complexity limits scalability on long documents. This paper introduces BiSparse-AAS (Bilinear Sparse Attention with Adaptive Spans), a…

Text Summarization

Local Temporal Bilinear Pooling for Fine-grained Action Parsing

2018-12-05 · CVPR 2019 6 · Yan Zhang, Siyu Tang, Krikamol Muandet, Christian Jarvers 외

Fine-grained temporal action parsing is important in many applications, such as daily activity understanding, human motion analysis, surgical robotics and others requiring subtle and precise operations in a long-term per…

Action ParsingDecoder

Optimization Methods for Convolutional Sparse Coding

2014-06-10 · Hilton Bristow, Simon Lucey

Sparse and convolutional constraints form a natural prior for many optimization problems that arise from physical processes. Detecting motifs in speech and musical passages, super-resolving images, compressing videos, an…