paper-with-me

홈 › Papers

Poolformer: Recurrent Networks with Pooling for Long-Sequence Modeling

2025-10-02 · Daniel Gallo Fernández arxiv

Sequence-to-sequence models have become central in Artificial Intelligence, particularly following the introduction of the transformer architecture. While initially developed for Natural Language Processing, these models have demonstrated utility across domains, including Computer Vision. Such models require mechanisms to exchange information along the time dimension, typically using recurrent or self-attention layers. However, self-attention scales quadratically with sequence length, limiting its practicality for very long sequences. We introduce Poolformer, a sequence-to-sequence model that replaces self-attention with recurrent layers and incorporates pooling operations to reduce sequence length. Poolformer is defined recursively using SkipBlocks, which contain residual blocks, a down-pooling layer, a nested SkipBlock, an up-pooling layer, and additional residual blocks. We conduct extensive experiments to support our architectural choices. Our results show that pooling greatly accelerates training, improves perceptual metrics (FID and IS), and prevents overfitting. Our experiments also suggest that long-range dependencies are handled by deep layers, while shallow layers take care of short-term features. Evaluated on raw audio, which naturally features long sequence lengths, Poolformer outperforms state-of-the-art models such as SaShiMi and Mamba. Future directions include applications to text and vision, as well as multi-modal scenarios, where a Poolformer-based LLM could effectively process dense representations of images and videos.

📄 PDF Abstract BibTeX arXiv:2510.02206

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Quasi-Recurrent Neural Networks

2016-11-05 · James Bradbury, Stephen Merity, Caiming Xiong, Richard Socher

Recurrent neural networks are a powerful tool for modeling sequential data, but the dependence of each timestep's computation on the previous timestep's output limits parallelism and makes RNNs unwieldy for very long seq…

Language ModelingLanguage ModellingMachine TranslationSentiment Analysis+3

MetaFormer Is Actually What You Need for Vision

2021-11-22 · CVPR 2022 1 · Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si 외

Transformers have shown great potential in computer vision tasks. A common belief is their attention-based token mixer module contributes most to their competence. However, recent works show the attention-based module in…

Image ClassificationObject DetectionRecommendation SystemsSemantic Segmentation

Spatiotemporal Pooling on Appropriate Topological Maps Represented as Two-Dimensional Images for EEG Classification

2024-03-07 · Takuto Fukushima, Ryusuke Miyamoto

Motor imagery classification based on electroencephalography (EEG) signals is one of the most important brain-computer interface applications, although it needs further improvement. Several methods have attempted to obta…

Brain Computer InterfaceClassificationEEGMotor Imagery

Poolingformer: Long Document Modeling with Pooling Attention

2021-05-10 · Hang Zhang, Yeyun Gong, Yelong Shen, Weisheng Li 외

In this paper, we introduce a two-level attention schema, Poolingformer, for long document modeling. Its first level uses a smaller sliding window pattern to aggregate information from neighbors. Its second level employs…

Advanced LSTM: A Study about Better Time Dependency Modeling in Emotion Recognition

2017-10-27 · Fei Tao, Gang Liu

Long short-term memory (LSTM) is normally used in recurrent neural network (RNN) as basic recurrent unit. However,conventional LSTM assumes that the state at current time step depends on previous time step. This assumpti…

Emotion ClassificationEmotion RecognitionGeneral Classification