paper-with-me

홈 › Papers

Integrating Human Vision Perception in Vision Transformers for Classifying Waste Items

2023-12-19 · Akshat Kishore Shrivastava, Tapan Kumar Gandhi

In this paper, we propose an novel methodology aimed at simulating the learning phenomenon of nystagmus through the application of differential blurring on datasets. Nystagmus is a biological phenomenon that influences human vision throughout life, notably by diminishing head shake from infancy to adulthood. Leveraging this concept, we address the issue of waste classification, a pressing global concern. The proposed framework comprises two modules, with the second module closely resembling the original Vision Transformer, a state-of-the-art model model in classification tasks. The primary motivation behind our approach is to enhance the model's precision and adaptability, mirroring the real-world conditions that the human visual system undergoes. This novel methodology surpasses the standard Vision Transformer model in waste classification tasks, exhibiting an improvement with a margin of 2%. This improvement underscores the potential of our methodology in improving model precision by drawing inspiration from human vision perception. Further research in the proposed methodology could yield greater performance results, and can be extrapolated to other global issues.

📄 PDF Abstract BibTeX arXiv:2312.12143

Code (0)

등록된 구현이 없습니다.

Tasks

Classification

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Position-Wise Feed-Forward Layer 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Evaluating Graphical Perception Capabilities of Vision Transformers

2026-02-20 · Poonam Poonam, Pere-Pau Vázquez, Timo Ropinski arxiv

Vision Transformers, ViTs, have emerged as a powerful alternative to convolutional neural networks, CNNs, in a variety of image-based tasks. While CNNs have previously been evaluated for their ability to perform graphica…

Do Vision Transformers See Like Humans? Evaluating their Perceptual Alignment

2025-08-13 · Pablo Hernández-Cámara, Jose Manuel Jaén-Lorites, Jorge Vila-Tomás, Valero Laparra 외 arxiv

Vision Transformers (ViTs) achieve remarkable performance in image recognition tasks, yet their alignment with human perception remains largely unexplored. This study systematically analyzes how model size, dataset size,…

Data Augmentation

Texture Representations in Deep Vision Models: Comparing CNNs, Vision Transformers, and Human Perception

2026-07-09 · Ludovica de Paolis, Marco Baroni, Alessandro Laio, Eugenio Piasini arxiv

In computational vision science, Convolutional Neural Networks (CNNs) have emerged as a popular model of biological vision because of the alignment they can exhibit with neural and behavioral data in humans and animals. …

Object Recognition

Divergences in Color Perception between Deep Neural Networks and Humans

2023-09-11 · Ethan O. Nadler, Elise Darragh-Ford, Bhargav Srinivasa Desikan, Christian Conaway 외

Deep neural networks (DNNs) are increasingly proposed as models of human vision, bolstered by their impressive performance on image classification and object recognition tasks. Yet, the extent to which DNNs capture funda…

image-classificationImage ClassificationImage SegmentationObject Recognition+2

Dynamic Granularity Matters: Rethinking Vision Transformers Beyond Fixed Patch Splitting

2025-11-24 · Qiyang Yu, Yu Fang, Tianrui Li, Xuemei Cao 외 arxiv

Vision Transformers (ViTs) have demonstrated strong capabilities in capturing global dependencies but often struggle to efficiently represent fine-grained local details. Existing multi-scale approaches alleviate this iss…

Computational Efficiency