paper-with-me

홈 › Papers

ACORT: A Compact Object Relation Transformer for Parameter Efficient Image Captioning

2022-02-11 · Jia Huei Tan, Ying Hua Tan, Chee Seng Chan, Joon Huang Chuah

Recent research that applies Transformer-based architectures to image captioning has resulted in state-of-the-art image captioning performance, capitalising on the success of Transformers on natural language tasks. Unfortunately, though these models work well, one major flaw is their large model sizes. To this end, we present three parameter reduction methods for image captioning Transformers: Radix Encoding, cross-layer parameter sharing, and attention parameter sharing. By combining these methods, our proposed ACORT models have 3.7x to 21.6x fewer parameters than the baseline model without compromising test performance. Results on the MS-COCO dataset demonstrate that our ACORT models are competitive against baselines and SOTA approaches, with CIDEr score >=126. Finally, we present qualitative results and ablation studies to demonstrate the efficacy of the proposed changes further. Code and pre-trained models are publicly available at https://github.com/jiahuei/sparse-image-captioning.

📄 PDF Abstract BibTeX arXiv:2202.05451

Code (1)

jiahuei/sparse-image-captioning 공식 구현 pytorch

Tasks

Image CaptioningRelation

Similar Papers 제목 키워드 기반

3D-aCortex: An Ultra-Compact Energy-Efficient Neurocomputing Platform Based on Commercial 3D-NAND Flash Memories

2019-08-07 · Mohammad Bavandpour, Shubham Sahay, Mohammad Reza Mahmoodi, Dmitri B. Strukov

The first contribution of this paper is the development of extremely dense, energy-efficient mixed-signal vector-by-matrix-multiplication (VMM) circuits based on the existing 3D-NAND flash memory blocks, without any need…

Decoding the decoder: Contextual sequence-to-sequence modeling for intracortical speech decoding

2026-03-10 · Michal Olak, Tommaso Boccato, Matteo Ferrante arxiv

Speech brain--computer interfaces require decoders that translate intracortical activity into linguistic output while remaining robust to limited data and day-to-day variability. While prior high-performing systems have …

Adaptively Pruned Spiking Neural Networks for Energy-Efficient Intracortical Neural Decoding

2025-04-15 · Francesca Rivelli, Martin Popov, Charalampos S. Kouzinopoulos, Guangzhi Tang

Intracortical brain-machine interfaces demand low-latency, energy-efficient solutions for neural decoding. Spiking Neural Networks (SNNs) deployed on neuromorphic hardware have demonstrated remarkable efficiency in neura…

SMART SLM: Structured Memory and Reasoning Transformer, A Small Language Model for Accurate Document Assistance

2025-12-24 · Divij Dudeja, Mayukha Pal arxiv

The user of Engineering Manuals (EM) finds it difficult to read EM s because they are long, have a dense format which includes written documents, step by step procedures, and standard parameter lists for engineering equi…

Neural Decoding of Overt Speech from ECoG Using Vision Transformers and Contrastive Representation Learning

2025-12-04 · Mohamed Baha Ben Ticha, Xingchen Ran, Guillaume Saldanha, Gaël Le Godais 외 arxiv

Speech Brain Computer Interfaces (BCIs) offer promising solutions to people with severe paralysis unable to communicate. A number of recent studies have demonstrated convincing reconstruction of intelligible speech from …

Representation LearningContrastive Learning