paper-with-me

Papers

HawkDrive: A Transformer-driven Visual Perception System for Autonomous Driving in Night Scene

2024-04-06 · Ziang Guo, Stepan Perminov, Mikhail Konenkov, Dzmitry Tsetserukou

Many established vision perception systems for autonomous driving scenarios ignore the influence of light conditions, one of the key elements for driving safety. To address this problem, we present HawkDrive, a novel perception system with hardware and software solutions. Hardware that utilizes stereo vision perception, which has been demonstrated to be a more reliable way of estimating depth information than monocular vision, is partnered with the edge computing device Nvidia Jetson Xavier AGX. Our software for low light enhancement, depth estimation, and semantic segmentation tasks, is a transformer-based neural network. Our software stack, which enables fast inference and noise reduction, is packaged into system modules in Robot Operating System 2 (ROS2). Our experimental results have shown that the proposed end-to-end system is effective in improving the depth estimation and semantic segmentation performance. Our dataset and codes will be released at https://github.com/ZionGo6/HawkDrive.

📄 PDF Abstract BibTeX arXiv:2404.04653

Code (1)

ziongo6/hawkdrive 공식 구현 pytorch

Tasks

Autonomous DrivingDepth EstimationEdge-computingSegmentationSemantic Segmentation

Similar Papers 제목 키워드 기반

GPTR: Gestalt-Perception Transformer for Diagram Object Detection

2022-12-29 · Xin Hu, Lingling Zhang, Jun Liu, Jinfu Fan 외

Diagram object detection is the key basis of practical applications such as textbook question answering. Because the diagram mainly consists of simple lines and color blocks, its visual features are sparser than those of…

DecoderObjectobject-detectionObject Detection+1

Evaluating Graphical Perception Capabilities of Vision Transformers

2026-02-20 · Poonam Poonam, Pere-Pau Vázquez, Timo Ropinski arxiv

Vision Transformers, ViTs, have emerged as a powerful alternative to convolutional neural networks, CNNs, in a variety of image-based tasks. While CNNs have previously been evaluated for their ability to perform graphica…

Perceptual Reality Transformer: Neural Architectures for Simulating Neurological Perception Conditions

2025-08-13 · Baihan Lin arxiv

Neurological conditions affecting visual perception create profound experiential divides between affected individuals and their caregivers, families, and medical professionals. We present the Perceptual Reality Transform…

VL2Spike: Spike-driven Distillation from VLMs for Low-Power Visual Perception in Embodied AI

2026-06-14 · Zinan Liu, Eric Zheng, Soumyaratna Debnath, Hao Shi 외 arxiv

Spiking neural networks (SNNs) are brain-inspired, event-driven models that compute with sparse spikes, which enables highly efficient visual perception in resource-constrained embodied AI models. The emergence of Spikin…

Visual Place RecognitionKnowledge Distillation

Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT

2025-08-12 · Muhammad A. Muttaqien, Tomohiro Motoda, Ryo Hanai, Yukiyasu Domae arxiv

Robotic pick-and-place tasks in convenience stores pose challenges due to dense object arrangements, occlusions, and variations in object properties such as color, shape, size, and texture. These factors complicate traje…

Trajectory Planning