paper-with-me

홈 › Papers

Mind the Gap: Evaluating Vision Systems in Small Data Applications

2025-04-08 · Samuel Stevens, S M Rayeed, Jenna Kline

The practical application of AI tools for specific computer vision tasks relies on the "small-data regime" of hundreds to thousands of labeled samples. This small-data regime is vital for applications requiring expensive expert annotations, such as ecological monitoring, medical diagnostics or industrial quality control. We find, however, that computer vision research has ignored the small data regime as evaluations increasingly focus on zero- and few-shot learning. We use the Natural World Tasks (NeWT) benchmark to compare multi-modal large language models (MLLMs) and vision-only methods across varying training set sizes. MLLMs exhibit early performance plateaus, while vision-only methods improve throughout the small-data regime, with performance gaps widening beyond 10 training examples. We provide the first comprehensive comparison between these approaches in small-data contexts and advocate for explicit small-data evaluations in AI research to better bridge theoretical advances with practical deployments.

📄 PDF Abstract BibTeX arXiv:2504.06486

Code (1)

samuelstevens/mindthegap 공식 구현 jax

Tasks

Few-Shot Learning

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception

2025-10-09 · Nikos Theodoridis, Tim Brophy, Reenu Mohandas, Ganesh Sistu 외 arxiv

Vision-Language Models (VLMs) are becoming increasingly powerful, demonstrating strong performance on a variety of tasks that require both visual and textual understanding. Their strong generalisation abilities make them…

Visual Question Answering

Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices

2025-09-25 · Yilong Li, Shuai Zhang, Yijing Zeng, Hao Zhang 외 arxiv

Large Multimodal Models (LMMs) are inherently modular, comprising vision and audio encoders, a projector, and a language backbone. Yet existing systems execute them monolithically, underutilizing the heterogeneous accele…

Bounded Minds, Generative Machines: Envisioning Conversational AI that Works with Human Heuristics and Reduces Bias Risk

2026-01-19 · Jiqun Liu arxiv

Conversational AI is rapidly becoming a primary interface for information seeking and decision making, yet most systems still assume idealized users. In practice, human reasoning is bounded by limited attention, uneven k…

Decision Making

Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language

2025-11-19 · Seungbeen Lee, Jinhong Jeong, Donghyun Kim, Yejin Son 외 arxiv

Our ability to interpret others' mental states through nonverbal cues (NVCs) is fundamental to our survival and social cohesion. While existing Theory of Mind (ToM) benchmarks have primarily focused on false-belief tasks…

MIND: Monge Inception Distance for Generative Models Evaluation

2026-05-07 · Quentin Berthet, Yu-Han Wu, Clement Crepy, Romuald Elie 외 arxiv

We propose the Monge Inception Distance (MIND), a metric for evaluating generative models that addresses key limitations of the widely adopted Fréchet Inception Distance (FID). The MIND metric leverages the sliced Wasser…