Semantic bottleneck for computer vision tasks
This paper introduces a novel method for the representation of images that is semantic by nature, addressing the question of computation intelligibility in computer vision tasks. More specifically, our proposition is to introduce what we call a semantic bottleneck in the processing pipeline, which is a crossing point in which the representation of the image is entirely expressed with natural language , while retaining the efficiency of numerical representations. We show that our approach is able to generate semantic representations that give state-of-the-art results on semantic content-based image retrieval and also perform very well on image classification tasks. Intelligibility is evaluated through user centered experiments for failure detection.
Code (0)
등록된 구현이 없습니다.
Tasks
Content-Based Image RetrievalGeneral Classificationimage-classificationImage ClassificationImage RetrievalRetrievalSimilar Papers 제목 키워드 기반
Hypergraph Vision Transformers: Images are More than Nodes, More than Edges
Recent advancements in computer vision have highlighted the scalability of Vision Transformers (ViTs) across various tasks, yet challenges remain in balancing adaptability, computational efficiency, and the ability to mo…
ClusteringComputational EfficiencyDiversityimage-classification+3Assistive Image Annotation Systems with Deep Learning and Natural Language Capabilities: A Review
While supervised learning has achieved significant success in computer vision tasks, acquiring high-quality annotated data remains a bottleneck. This paper explores both scholarly and non-scholarly works in AI-assistive …
Active LearningImage Captioningimage-classificationImage Classification+7Computer Vision for Road Imaging and Pothole Detection: A State-of-the-Art Review of Systems and Algorithms
Computer vision algorithms have been prevalently utilized for 3-D road imaging and pothole detection for over two decades. Nonetheless, there is a lack of systematic survey articles on state-of-the-art (SoTA) computer vi…
ArticlesSegmentationSemantic SegmentationSurveyOutline Objects using Deep Reinforcement Learning
Image segmentation needs both local boundary position information and global object context information. The performance of the recent state-of-the-art method, fully convolutional networks, reaches a bottleneck due to th…
Deep Reinforcement LearningImage SegmentationObjectPosition+5Exploring Deep Spiking Neural Networks for Automated Driving Applications
Neural networks have become the standard model for various computer vision tasks in automated driving including semantic segmentation, moving object detection, depth estimation, visual odometry, etc. The main flavors of …
Depth EstimationMoving Object Detectionobject-detectionObject Detection+2