paper-with-me

홈 › Papers

BioCLIP 2: Emergent Properties from Scaling Hierarchical Contrastive Learning

2025-05-29 · Jianyang Gu, Samuel Stevens, Elizabeth G Campolongo, Matthew J Thompson, Net Zhang, Jiaman Wu, Andrei Kopanev, Zheda Mai, Alexander E. White, James Balhoff, Wasila Dahdul, Daniel Rubenstein, Hilmar Lapp, Tanya Berger-Wolf, Wei-Lun Chao, Yu Su

Foundation models trained at scale exhibit remarkable emergent behaviors, learning new capabilities beyond their initial training objectives. We find such emergent behaviors in biological vision models via large-scale contrastive vision-language training. To achieve this, we first curate TreeOfLife-200M, comprising 214 million images of living organisms, the largest and most diverse biological organism image dataset to date. We then train BioCLIP 2 on TreeOfLife-200M to distinguish different species. Despite the narrow training objective, BioCLIP 2 yields extraordinary accuracy when applied to various biological visual tasks such as habitat classification and trait prediction. We identify emergent properties in the learned embedding space of BioCLIP 2. At the inter-species level, the embedding distribution of different species aligns closely with functional and ecological meanings (e.g., beak sizes and habitats). At the intra-species level, instead of being diminished, the intra-species variations (e.g., life stages and sexes) are preserved and better separated in subspaces orthogonal to inter-species distinctions. We provide formal proof and analyses to explain why hierarchical supervision and contrastive objectives encourage these emergent properties. Crucially, our results reveal that these properties become increasingly significant with larger-scale training data, leading to a biologically meaningful embedding space.

📄 PDF Abstract BibTeX arXiv:2505.23883

Code (1)

imageomics/treeoflife-toolbox 공식 구현 jax

Tasks

Contrastive Learning

Similar Papers 제목 키워드 기반

BioCLIP: A Vision Foundation Model for the Tree of Life

2023-11-30 · CVPR 2024 1 · Samuel Stevens, Jiaman Wu, Matthew J Thompson, Elizabeth G Campolongo 외

Images of the natural world, collected by a variety of cameras, from drones to individual phones, are increasingly abundant sources of biological information. There is an explosion of computational methods and tools, par…

Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification

2026-06-20 · Zhiyuan Tao, Srikumar Sastry, Matthew J Thompson, Elizabeth G Campolongo 외 arxiv

Multimodal contrastive learning has enabled zero-shot visual classification by aligning images with textual categories. However, in hierarchically structured label spaces, existing methods often produce predictions that …

Contrastive Learning

EB-DEVS: A Formal Framework for Modeling and Simulation of Emergent Behavior in Dynamic Complex Systems

2020-10-10 · Daniel J. Foguelman, Philipp Henning, Adelinde Uhrmacher, Rodrigo Castro

Emergent behavior is a key feature defining a system under study as a complex system. Simulation has been recognized as the only way to deal with the study of the emergency of properties (at a macroscopic level) among gr…

Predicting Emergent Abilities with Infinite Resolution Evaluation

2023-10-05 · Shengding Hu, Xin Liu, Xu Han, Xinrong Zhang 외

The scientific scale-up of large language models (LLMs) necessitates a comprehensive understanding of their scaling properties. However, the existing literature on the scaling properties only yields an incomplete answer:…

Code Generation

Audio-to-Image Bird Species Retrieval without Audio-Image Pairs via Text Distillation

2026-01-31 · Ilyass Moummad, Marius Miron, Lukas Rauch, David Robinson 외 arxiv

Audio-to-image retrieval offers an interpretable alternative to audio-only classification for bioacoustic species recognition, but learning aligned audio-image representations is challenging due to the scarcity of paired…

Image Retrieval