paper-with-me

Papers Zero-Shot Object Detection

“Zero-Shot Object Detection” 태그가 달린 논문 63편 · 필터 해제

Does Your VFM Speak Plant? The Botanical Grammar of Vision Foundation Models for Object Detection

2026-04-10 · Lars Lundqvist, Earl Ranario, Hamid Kamangir, Heesup Yun 외 arxiv

Vision foundation models (VFMs) offer the promise of zero-shot object detection without task-specific training data, yet their performance in complex agricultural scenes remains highly sensitive to text prompt constructi…

Zero-Shot Object DetectionPrompt Engineering

PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training

2026-04-01 · Weifu Fu, Jinyang Li, Bin-Bin Gao, Jialin Li 외 arxiv

Open-Set Object Detection (OSOD) enables recognition of novel categories beyond fixed classes but faces challenges in aligning text representations with complex visual concepts and the scarcity of image-text pairs for ra…

Zero-Shot Object Detection

TinyVLM: Zero-Shot Object Detection on Microcontrollers via Vision-Language Distillation with Matryoshka Embeddings

2026-02-24 · Bibin Wilson arxiv

Zero-shot object detection enables recognising novel objects without task-specific training, but current approaches rely on large vision language models (VLMs) like CLIP that require hundreds of megabytes of memory - far…

Zero-Shot Object Detection

Robust Object Detection with Pseudo Labels from VLMs using Per-Object Co-teaching

2025-11-13 · Uday Bhaskar, Rishabh Bhattacharya, Avinash Patel, Sarthak Khoche 외 arxiv

Foundation models, especially vision-language models (VLMs), offer compelling zero-shot object detection for applications like autonomous driving, a domain where manual labelling is prohibitively expensive. However, thei…

Zero-Shot Object DetectionRobust Object DetectionAutonomous Driving

A Computer Vision Pipeline for Individual-Level Behavior Analysis: Benchmarking on the Edinburgh Pig Dataset

2025-09-15 · Haiyu Yang, Enhong Liu, Jennifer Sun, Sumit Sharma 외 arxiv

Animal behavior analysis plays a crucial role in understanding animal welfare, health status, and productivity in agricultural settings. However, traditional manual observation methods are time-consuming, subjective, and…

Zero-Shot Object Detection

Fine-Grained Zero-Shot Object Detection

2025-07-14 · Hongxu Ma, Chenbo Zhang, Lu Zhang, Jiaogen Zhou 외 arxiv

Zero-shot object detection (ZSD) aims to leverage semantic descriptions to localize and recognize objects of both seen and unseen classes. Existing ZSD works are mainly coarse-grained object detection, where the classes …

Zero-Shot Object Detection

VisionReasoner: Unified Visual Perception and Reasoning via Reinforcement Learning

2025-05-17 · Yuqi Liu, Tianyuan Qu, Zhisheng Zhong, Bohao Peng 외

Large vision-language models exhibit inherent capabilities to handle diverse visual perception tasks. In this paper, we introduce VisionReasoner, a unified framework capable of reasoning and solving multiple visual perce…

2D Object DetectionObject CountingReasoning SegmentationReferring Expression Segmentation+4

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety

2025-04-18 · Shashank Shriram, Srinivasa Perisetla, Aryan Keskar, Harsha Krishnaswamy 외

Detecting anomalous hazards in visual data, particularly in video streams, is a critical challenge in autonomous driving. Existing models often struggle with unpredictable, out-of-label hazards due to their reliance on p…

Anomaly DetectionAutonomous DrivingDenoisingLanguage Modeling+7

Finding the Reflection Point: Unpadding Images to Remove Data Augmentation Artifacts in Large Open Source Image Datasets for Machine Learning

2025-04-04 · Lucas Choi, Ross Greer

In this paper, we address a novel image restoration problem relevant to machine learning dataset curation: the detection and removal of noisy mirrored padding artifacts. While data augmentation techniques like padding ar…

Data AugmentationHuman DetectionImage Restorationobject-detection+2

The Power of One: A Single Example is All it Takes for Segmentation in VLMs

2025-03-13 · Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal, James J. Little

Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities by implicitly learning associations between textual descriptions and imag…

Allobject-detectionObject DetectionPrompt Engineering+2

LangGas: Introducing Language in Selective Zero-Shot Background Subtraction for Semi-Transparent Gas Leak Detection with a New Dataset

2025-03-04 · Wenqi Guo, Yiyang Du, Shan Du

Gas leakage poses a significant hazard that requires prevention. Traditionally, human inspection has been used for detection, a slow and labour-intensive process. Recent research has applied machine learning techniques t…

Classificationobject-detectionObject DetectionSegmentation+1

UniFa: A unified feature hallucination framework for any-shot object detection

2025-03-01 · journal 2025 3 · Hui Nie, Ruiping Wang, Xilin Chen

Any-shot object detection seeks to simultaneously detect base (many-shot), few-shot and zero-shot categories. The primary challenge lies in insufficient visual data for rare (few-shot and zero-shot) categories, hindering…

Generalized Zero-Shot Object DetectionHallucinationobject-detectionObject Detection+1

CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object Detection

2024-12-13 · Qibo Chen, Weizhong Jin, Jianyue Ge, Mengdi Liu 외

Recent research on universal object detection aims to introduce language in a SoTA closed-set detector and then generalize the open-set concepts by constructing large-scale (text-region) datasets for training. However, t…

object-detectionObject DetectionZero-Shot Object Detection

No Annotations for Object Detection in Art through Stable Diffusion

2024-12-09 · Patrick Ramos, Nicolas Gonthier, Selina Khan, Yuta Nakashima 외

Object detection in art is a valuable tool for the digital humanities, as it allows for faster identification of objects in artistic and historical images compared to humans. However, annotating such images poses signifi…

Objectobject-detectionObject DetectionZero-Shot Object Detection

Gaussian Splatting Under Attack: Investigating Adversarial Noise in 3D Objects

2024-12-03 · Abdurrahman Zeybey, Mehmet Ergezer, Tommy Nguyen

3D Gaussian Splatting has advanced radiance field reconstruction, enabling high-quality view synthesis and fast rendering in 3D modeling. While adversarial attacks on object detection models are well-studied for 2D image…

Autonomous Drivingobject-detectionObject DetectionZero-Shot Object Detection

DINO-X: A Unified Vision Model for Open-World Object Detection and Understanding

2024-11-21 · Tianhe Ren, Yihao Chen, Qing Jiang, Zhaoyang Zeng 외

In this paper, we introduce DINO-X, which is a unified object-centric vision model developed by IDEA Research with the best open-world object detection performance to date. DINO-X employs the same Transformer-based encod…

Long-tailed Object DetectionObjectobject-detectionObject Detection+3

OV-DINO: Unified Open-Vocabulary Detection with Language-Aware Selective Fusion

2024-07-10 · Hao Wang, Pengzhen Ren, Zequn Jie, Xiao Dong 외

Open-vocabulary detection is a challenging task due to the requirement of detecting objects based on class names, including those not encountered during training. Existing methods have shown strong zero-shot detection ca…

Object DetectionZero-Shot Object Detection

Segment Anything Model for automated image data annotation: empirical studies using text prompts from Grounding DINO

2024-06-27 · Fuseini Mumuni, Alhassan Mumuni

Grounding DINO and the Segment Anything Model (SAM) have achieved impressive performance in zero-shot object detection and image segmentation, respectively. Together, they have a great potential to revolutionize applicat…

Image SegmentationMedical Image Segmentationobject-detectionObject Detection+6

Eating Smart: Advancing Health Informatics with the Grounding DINO based Dietary Assistant App

2024-06-02 · Abdelilah Nossair, Hamza El Housni

The Smart Dietary Assistant utilizes Machine Learning to provide personalized dietary advice, focusing on users with conditions like diabetes. This app leverages the Grounding DINO model, which combines a text encoder an…

ManagementNutritionobject-detectionObject Detection+1

OV-DQUO: Open-Vocabulary DETR with Denoising Text Query Training and Open-World Unknown Objects Supervision

2024-05-28 · Junjie Wang, Bin Chen, Bin Kang, Yulin Li 외

Open-vocabulary detection aims to detect objects from novel categories beyond the base categories on which the detector is trained. However, existing open-vocabulary detectors trained on base category data tend to assign…

Contrastive LearningDenoisingobject-detectionObject Detection+2
1–20 / 63 다음 →