paper-with-me

Papers

BlindSpotNet: Seeing Where We Cannot See

2022-07-08 · Taichi Fukuda, Kotaro Hasegawa, Shinya Ishizaki, Shohei Nobuhara, Ko Nishino

We introduce 2D blind spot estimation as a critical visual task for road scene understanding. By automatically detecting road regions that are occluded from the vehicle's vantage point, we can proactively alert a manual driver or a self-driving system to potential causes of accidents (e.g., draw attention to a road region from which a child may spring out). Detecting blind spots in full 3D would be challenging, as 3D reasoning on the fly even if the car is equipped with LiDAR would be prohibitively expensive and error prone. We instead propose to learn to estimate blind spots in 2D, just from a monocular camera. We achieve this in two steps. We first introduce an automatic method for generating ``ground-truth'' blind spot training data for arbitrary driving videos by leveraging monocular depth estimation, semantic segmentation, and SLAM. The key idea is to reason in 3D but from 2D images by defining blind spots as those road regions that are currently invisible but become visible in the near future. We construct a large-scale dataset with this automatic offline blind spot estimation, which we refer to as Road Blind Spot (RBS) dataset. Next, we introduce BlindSpotNet (BSN), a simple network that fully leverages this dataset for fully automatic estimation of frame-wise blind spot probability maps for arbitrary driving videos. Extensive experimental results demonstrate the validity of our RBS Dataset and the effectiveness of our BSN.

📄 PDF Abstract BibTeX arXiv:2207.03870

Code (0)

등록된 구현이 없습니다.

Tasks

Depth EstimationMonocular Depth Estimationroad scene understandingScene UnderstandingSemantic Segmentation

Similar Papers 제목 키워드 기반

Responsibility in Extensive Form Games

2023-12-12 · Qi Shi

Two different forms of responsibility, counterfactual and seeing-to-it, have been extensively discussed in the philosophy and AI in the context of a single agent or multiple agents acting simultaneously. Although the gen…

counterfactualFormPhilosophy

Vision-Language Asymmetry in Bistable Image Captioning

2026-06-06 · Arohan Agate arxiv

Wittgenstein's duck-rabbit poses a question for vision-language models: when a model captions an ambiguous image, where in the model is the commitment to one aspect made? We address this with a 3,320-generation behaviora…

Image Captioning

Vision: looking and seeing through our brain's information bottleneck

2025-03-24 · Li Zhaoping

Our brain recognizes only a tiny fraction of sensory input, due to an information processing bottleneck. This blinds us to most visual inputs. Since we are blind to this blindness, only a recent framework highlights this…

Seeing the Goal, Missing the Truth: Human Accountability for AI Bias

2026-02-10 · Sean Cao, Wei Jiang, Hui Xu arxiv

This research explores how human-defined goals influence the behavior of Large Language Models (LLMs) through purpose-conditioned cognition. Using financial prediction tasks, we show that revealing the downstream use (e.…

Seeing What Is Not There: Learning Context to Determine Where Objects Are Missing

2017-02-26 · CVPR 2017 7 · Jin Sun, David W. Jacobs

Most of computer vision focuses on what is in an image. We propose to train a standalone object-centric context representation to perform the opposite task: seeing what is not there. Given an image, our context model can…

Objectobject-detectionObject Detection