paper-with-me

Papers

Benchmarking Attribute Discrimination in Infant-Scale Vision-Language Models

2025-12-22 · Patrick Batsell, Satoshi Tsutsui, Bihan Wen arxiv

Infants learn not only object categories but also fine-grained visual attributes such as color, size, and texture from limited experience. Prior infant-scale vision--language models have mainly been evaluated on object recognition, leaving open whether they support within-class attribute discrimination. We introduce a controlled benchmark that varies color, size, and texture across 67 everyday object classes using synthetic rendering to decouple attribute values from object identity. We evaluate infant-trained models (CVCL and an infant-trained DINO baseline) against web-scale and ImageNet models (CLIP, SigLIP, ResNeXt) under two complementary settings: an image-only prototype test and a text--vision test with attribute--object prompts. We find a dissociation between visual and linguistic attribute information: infant-trained models form strong visual representations for size and discriminate texture comparably to other models, but perform poorly on visual color discrimination, and in the text--vision setting they struggle to ground color and show only modest size grounding. In contrast, web-trained vision--language models strongly ground color from text while exhibiting weaker visual size discrimination.

📄 PDF Abstract BibTeX arXiv:2512.18951

Code (0)

등록된 구현이 없습니다.

Tasks

Object Recognition

Similar Papers 제목 키워드 기반

BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models

2025-12-11 · Shengao Wang, Wenqi Wang, Zecheng Wang, Max Whitton 외 arxiv

Early children's developmental trajectories set up a natural goal for sample-efficient pretraining of vision foundation models. We introduce BabyVLM-V2, a developmentally grounded framework for infant-inspired vision-lan…

Spatial Reasoning

Evaluating Visual Number Discrimination in Deep Neural Networks

2023-03-13 · Ivana Kajić, Aida Nematzadeh

The ability to discriminate between large and small quantities is a core aspect of basic numerical competence in both humans and animals. In this work, we examine the extent to which the state-of-the-art neural networks …

AggPose: Deep Aggregation Vision Transformer for Infant Pose Estimation

2022-05-11 · Xu Cao, Xiaoye Li, Liya Ma, Yi Huang 외

Movement and pose assessment of newborns lets experienced pediatricians predict neurodevelopmental disorders, allowing early intervention for related diseases. However, most of the newest AI approaches for human pose est…

Keypoint DetectionPose Estimation

Evaluating computational models of infant phonetic learning across languages

2020-08-06 · Yevgen Matusevych, Thomas Schatz, Herman Kamper, Naomi H. Feldman 외

In the first year of life, infants' speech perception becomes attuned to the sounds of their native language. Many accounts of this early phonetic learning exist, but computational models predicting the attunement patter…

Multi-stream 3D FCN with Multi-scale Deep Supervision for Multi-modality Isointense Infant Brain MR Image Segmentation

2017-11-28 · Guodong Zeng, Guoyan Zheng

We present a method to address the challenging problem of segmentation of multi-modality isointense infant brain MR images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF). Our method is based on c…

Image SegmentationInfant Brain Mri SegmentationMRI segmentationSegmentation+1