Papers Zero-Shot Object Detection
“Zero-Shot Object Detection” 태그가 달린 논문 63편 · 필터 해제
Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection
Vision foundation models (VFMs) offer the promise of zero-shot object detection without task-specific training data, yet their performance in complex agricultural scenes remains highly sensitive to text prompt constructi…
Zero-Shot Object DetectionPrompt EngineeringPET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
Open-Set Object Detection (OSOD) enables recognition of novel categories beyond fixed classes but faces challenges in aligning text representations with complex visual concepts and the scarcity of image-text pairs for ra…
Zero-Shot Object DetectionTinyVLM: Zero-Shot Object Detection on Microcontrollers via Vision-Language Distillation with Matryoshka Embeddings
Zero-shot object detection enables recognising novel objects without task-specific training, but current approaches rely on large vision language models (VLMs) like CLIP that require hundreds of megabytes of memory - far…
Zero-Shot Object DetectionRobust Object Detection with Pseudo Labels from VLMs using Per-Object Co-teaching
Foundation models, especially vision-language models (VLMs), offer compelling zero-shot object detection for applications like autonomous driving, a domain where manual labelling is prohibitively expensive. However, thei…
Zero-Shot Object DetectionRobust Object DetectionAutonomous DrivingA Computer Vision Pipeline for Individual-Level Behavior Analysis: Benchmarking on the Edinburgh Pig Dataset
Animal behavior analysis plays a crucial role in understanding animal welfare, health status, and productivity in agricultural settings. However, traditional manual observation methods are time-consuming, subjective, and…
Zero-Shot Object DetectionFine-Grained Zero-Shot Object Detection
Zero-shot object detection (ZSD) aims to leverage semantic descriptions to localize and recognize objects of both seen and unseen classes. Existing ZSD works are mainly coarse-grained object detection, where the classes …
Zero-Shot Object DetectionVisionReasoner: Unified Visual Perception and Reasoning via Reinforcement Learning
Large vision-language models exhibit inherent capabilities to handle diverse visual perception tasks. In this paper, we introduce VisionReasoner, a unified framework capable of reasoning and solving multiple visual perce…
2D Object DetectionObject CountingReasoning SegmentationReferring Expression Segmentation+4Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety
Detecting anomalous hazards in visual data, particularly in video streams, is a critical challenge in autonomous driving. Existing models often struggle with unpredictable, out-of-label hazards due to their reliance on p…
Anomaly DetectionAutonomous DrivingDenoisingLanguage Modeling+7Finding the Reflection Point: Unpadding Images to Remove Data Augmentation Artifacts in Large Open Source Image Datasets for Machine Learning
In this paper, we address a novel image restoration problem relevant to machine learning dataset curation: the detection and removal of noisy mirrored padding artifacts. While data augmentation techniques like padding ar…
Data AugmentationHuman DetectionImage Restorationobject-detection+2The Power of One: A Single Example is All it Takes for Segmentation in VLMs
Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities by implicitly learning associations between textual descriptions and imag…
Allobject-detectionObject DetectionPrompt Engineering+2LangGas: Introducing Language in Selective Zero-Shot Background Subtraction for Semi-Transparent Gas Leak Detection with a New Dataset
Gas leakage poses a significant hazard that requires prevention. Traditionally, human inspection has been used for detection, a slow and labour-intensive process. Recent research has applied machine learning techniques t…
Classificationobject-detectionObject DetectionSegmentation+1UniFa: A unified feature hallucination framework for any-shot object detection
Any-shot object detection seeks to simultaneously detect base (many-shot), few-shot and zero-shot categories. The primary challenge lies in insufficient visual data for rare (few-shot and zero-shot) categories, hindering…
Generalized Zero-Shot Object DetectionHallucinationobject-detectionObject Detection+1CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection
Recent research on universal object detection aims to introduce language in a SoTA closed-set detector and then generalize the open-set concepts by constructing large-scale (text-region) datasets for training. However, t…
object-detectionObject DetectionZero-Shot Object DetectionNo Annotations for Object Detection in Art through Stable Diffusion
Object detection in art is a valuable tool for the digital humanities, as it allows for faster identification of objects in artistic and historical images compared to humans. However, annotating such images poses signifi…
Objectobject-detectionObject DetectionZero-Shot Object DetectionGaussian Splatting Under Attack: Investigating Adversarial Noise in 3D Objects
3D Gaussian Splatting has advanced radiance field reconstruction, enabling high-quality view synthesis and fast rendering in 3D modeling. While adversarial attacks on object detection models are well-studied for 2D image…
Autonomous Drivingobject-detectionObject DetectionZero-Shot Object DetectionDINO-X: A Unified Vision Model for Open-World Object Detection and Understanding
In this paper, we introduce DINO-X, which is a unified object-centric vision model developed by IDEA Research with the best open-world object detection performance to date. DINO-X employs the same Transformer-based encod…
Long-tailed Object DetectionObjectobject-detectionObject Detection+3OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion
Open-vocabulary detection is a challenging task due to the requirement of detecting objects based on class names, including those not encountered during training. Existing methods have shown strong zero-shot detection ca…
Object DetectionZero-Shot Object DetectionSegment Anything Model for automated image data annotation: empirical studies using text prompts from Grounding DINO
Grounding DINO and the Segment Anything Model (SAM) have achieved impressive performance in zero-shot object detection and image segmentation, respectively. Together, they have a great potential to revolutionize applicat…
Image SegmentationMedical Image Segmentationobject-detectionObject Detection+6Eating Smart: Advancing Health Informatics with the Grounding DINO based Dietary Assistant App
The Smart Dietary Assistant utilizes Machine Learning to provide personalized dietary advice, focusing on users with conditions like diabetes. This app leverages the Grounding DINO model, which combines a text encoder an…
ManagementNutritionobject-detectionObject Detection+1OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision
Open-vocabulary detection aims to detect objects from novel categories beyond the base categories on which the detector is trained. However, existing open-vocabulary detectors trained on base category data tend to assign…
Contrastive LearningDenoisingobject-detectionObject Detection+2