paper-with-me

홈 › Papers

When Does Supervised Training Pay Off? The Hidden Economics of Object Detection in the Era of Vision-Language Models

2025-10-13 · Samer Al-Hamadani arxiv

Object detection traditionally relies on costly manual annotation. We present the first comprehensive cost-effectiveness analysis comparing supervised YOLO and zero-shot vision-language models (Gemini Flash 2.5 and GPT-4). Evaluated on 5,000 stratified COCO images and 500 diverse product images, combined with Total Cost of Ownership modeling, we derive break-even thresholds for architecture selection. Results show supervised YOLO attains 91.2% accuracy versus 68.5% for Gemini and 71.3% for GPT-4 on standard categories; the annotation expense for a 100-category system is $10,800, and the accuracy advantage only pays off beyond 55 million inferences (151,000 images/day for one year). On diverse product categories Gemini achieves 52.3% and GPT-4 55.1%, while supervised YOLO cannot detect untrained classes. Cost-per-correct-detection favors Gemini ($0.00050) and GPT-4 ($0.00067) over YOLO ($0.143) at 100,000 inferences. We provide decision frameworks showing that optimal architecture choice depends on inference volume, category stability, budget, and accuracy requirements.

📄 PDF Abstract BibTeX arXiv:2510.11302

Code (0)

등록된 구현이 없습니다.

Tasks

Object Detection

Similar Papers 제목 키워드 기반

How does unlabeled data improve generalization in self-training? A one-hidden-layer theoretical analysis

2022-01-21 · Shuai Zhang, Meng Wang, Sijia Liu, Pin-Yu Chen 외

Self-training, a semi-supervised learning algorithm, leverages a large amount of unlabeled data to improve learning when the labeled data are limited. Despite empirical successes, its theoretical characterization remains…

Generalization of Auto-Regressive Hidden Markov Models to Non-Linear Dynamics and Unit Quaternion Observation Space

2023-02-23 · Michele Ginesi, Paolo Fiorini

Latent variable models are widely used to perform unsupervised segmentation of time series in different context such as robotics, speech recognition, and economics. One of the most widely used latent variable model is th…

speech-recognitionSpeech RecognitionTime SeriesTime Series Analysis

An Introduction to Animal Movement Modeling with Hidden Markov Models using Stan for Bayesian Inference

2018-06-27

Hidden Markov models (HMMs) are popular time series model in many fields including ecology, economics and genetics. HMMs can be defined over discrete or continuous time, though here we only cover the former. In the field…

Bayesian InferenceTime SeriesTime Series Analysis

Diversified Hidden Markov Models for Sequential Labeling

2019-04-05 · Maoying Qiao, Wei Bian, Richard Yida Xu, DaCheng Tao

Labeling of sequential data is a prevalent meta-problem for a wide range of real world applications. While the first-order Hidden Markov Models (HMM) provides a fundamental approach for unsupervised sequential labeling, …

DiversityOptical Character RecognitionOptical Character Recognition (OCR)Part-Of-Speech Tagging+2

An Empirical Study of Training Self-Supervised Vision Transformers

2021-04-05 · ICCV 2021 10 · Xinlei Chen, Saining Xie, Kaiming He

This paper does not describe a novel method. Instead, it studies a straightforward, incremental, yet must-know baseline given the recent progress in computer vision: self-supervised learning for Vision Transformers (ViT)…

Out-of-Distribution GeneralizationSelf-Supervised Image ClassificationSelf-Supervised Learning