paper-with-me

홈 › Papers

YotoR-You Only Transform One Representation

2024-05-30 · José Ignacio Díaz Villa, Patricio Loncomilla, Javier Ruiz-del-Solar

This paper introduces YotoR (You Only Transform One Representation), a novel deep learning model for object detection that combines Swin Transformers and YoloR architectures. Transformers, a revolutionary technology in natural language processing, have also significantly impacted computer vision, offering the potential to enhance accuracy and computational efficiency. YotoR combines the robust Swin Transformer backbone with the YoloR neck and head. In our experiments, YotoR models TP5 and BP4 consistently outperform YoloR P6 and Swin Transformers in various evaluations, delivering improved object detection performance and faster inference speeds than Swin Transformer models. These results highlight the potential for further model combinations and improvements in real-time object detection with Transformers. The paper concludes by emphasizing the broader implications of YotoR, including its potential to enhance transformer-based models for image-related tasks.

📄 PDF Abstract BibTeX arXiv:2405.19629

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyObjectobject-detectionObject DetectionReal-Time Object Detection

Methods 이 논문이 사용한 방법론

Attention 설명 없음
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Adam 설명 없음
Position-Wise Feed-Forward Layer 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…

Similar Papers 제목 키워드 기반

Algebras of actions in an agent's representations of the world

2023-10-02 · Alexander Dean, Eduardo Alonso, Esther Mondragon

In this paper, we propose a framework to extract the algebra of the transformations of worlds from the perspective of an agent. As a starting point, we use our framework to reproduce the symmetry-based representations fr…

Representation Learning

Robustness to Transformations Across Categories: Is Robustness To Transformations Driven by Invariant Neural Representations?

2020-06-30 · Hojin Jang, Syed Suleman Abbas Zaidi, Xavier Boix, Neeraj Prasad 외

Deep Convolutional Neural Networks (DCNNs) have demonstrated impressive robustness to recognize objects under transformations (eg. blur or noise) when these transformations are included in the training set. A hypothesis …

No Other Representation Component Is Needed: Diffusion Transformers Can Provide Representation Guidance by Themselves

2025-05-05 · Dengyang Jiang, Mengmeng Wang, Liuzhuozheng Li, Lei Zhang 외

Recent studies have demonstrated that learning a meaningful internal representation can both accelerate generative training and enhance the generation quality of diffusion transformers. However, existing approaches neces…

Image GenerationRepresentation Learning

Gauge Freedom and Metric Dependence in Neural Representation Spaces

2026-03-06 · Jericho Cain arxiv

Neural network representations are often analyzed as vectors in a fixed Euclidean space. However, their coordinates are not uniquely defined. If a hidden representation is transformed by an invertible linear map, the net…

Probing Spectrum-Like Organization of States of Mind in Transformer Representation Spaces

2025-12-23 · Sophie Zhao arxiv

We investigate whether graded states of mind form spectrum-like structure in transformer representation spaces. To do so, we construct a dataset of 636 short natural-language sentences annotated with both a continuous sc…