paper-with-me

Papers

Encoding Longer-term Contextual Multi-modal Information in a Predictive Coding Model

2018-04-17 · Junpei Zhong, Tetsuya OGATA, Angelo Cangelosi

Studies suggest that within the hierarchical architecture, the topological higher level possibly represents a conscious category of the current sensory events with slower changing activities. They attempt to predict the activities on the lower level by relaying the predicted information. On the other hand, the incoming sensory information corrects such prediction of the events on the higher level by the novel or surprising signal. We propose a predictive hierarchical artificial neural network model that examines this hypothesis on neurorobotic platforms, based on the AFA-PredNet model. In this neural network model, there are different temporal scales of predictions exist on different levels of the hierarchical predictive coding, which are defined in the temporal parameters in the neurons. Also, both the fast and the slow-changing neural activities are modulated by the active motor activities. A neurorobotic experiment based on the architecture was also conducted based on the data collected from the VRep simulator.

📄 PDF Abstract BibTeX arXiv:1804.06774

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Phase Diagram of Vision Large Language Models Inference: A Perspective from Interaction across Image and Instruction

2024-11-01 · Houjing Wei, Yuting Shi, Naoya Inoue

Vision Large Language Models (VLLMs) usually take input as a concatenation of image token embeddings and text token embeddings and conduct causal modeling. However, their internal behaviors remain underexplored, raising …

multimodal interaction

TULIP: Token-length Upgraded CLIP

2024-10-13 · Ivona Najdenkoska, Mohammad Mahdi Derakhshani, Yuki M. Asano, Nanne van Noord 외

We address the challenge of representing long captions in vision-language models, such as CLIP. By design these models are limited by fixed, absolute positional encodings, restricting inputs to a maximum of 77 tokens and…

Image GenerationPositionText to Image GenerationText-to-Image Generation

CMDR: Contextual Multimodal Document Retrieval

2026-07-07 · Ryota Tanaka, Taku Hasegawa, Kyosuke Nishida arxiv

Multimodal document retrieval aims to retrieve relevant pages while preserving both textual and visual content from the original document. However, existing benchmarks primarily evaluate simple lexical or semantic matchi…

Contrastive Learning

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations

2025-05-29 · Yiming Lei, Zhizheng Yang, Zeming Liu, Haitao Leng 외

Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal models suffer from the weak capability of mu…

Contextually Structured Token Dependency Encoding for Large Language Models

2025-01-30 · James Blades, Frederick Somerfield, William Langley, Susan Everingham 외

Token representation strategies within large-scale neural architectures often rely on contextually refined embeddings, yet conventional approaches seldom encode structured relationships explicitly within token interactio…

Computational EfficiencyText Generation