paper-with-me

Papers

C-RADIOv4 (Tech Report)

2026-01-24 · Mike Ranzinger, Greg Heinrich, Collin McCarthy, Jan Kautz, Andrew Tao, Bryan Catanzaro, Pavlo Molchanov arxiv

By leveraging multi-teacher distillation, agglomerative vision backbones provide a unified student model that retains and improves the distinct capabilities of multiple teachers. In this tech report, we describe the most recent release of the C-RADIO family of models, C-RADIOv4, which builds upon AM-RADIO/RADIOv2.5 in design, offering strong improvements on key downstream tasks at the same computational complexity. We release -SO400M (412M params), and -H (631M) model variants, both trained with an updated set of teachers: SigLIP2, DINOv3, and SAM3. In addition to improvements on core metrics and new capabilities from imitating SAM3, the C-RADIOv4 model family further improves any-resolution support, brings back the ViTDet option for drastically enhanced efficiency at high-resolution, and comes with a permissive license.

📄 PDF Abstract BibTeX arXiv:2601.17237

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RadioVIL: Anomaly-Aware Diffusion Models for Radio Map Inpainting and Zero-Shot Vehicle Localization

2026-08-17 · Ruixin Zhao, Xiucheng Wang, Qiming Zhang, Nan Cheng 외 arxiv

High-precision radio map construction is essential for emerging 6G Integrated Sensing and Communication (ISAC) applications, including digital twins and intelligent transportation. However, existing deep learning methods…

RADIOv2.5: Improved Baselines for Agglomerative Vision Foundation Models

2025-01-01 · CVPR 2025 1 · Greg Heinrich, Mike Ranzinger, Hongxu Yin, Yao Lu 외

Agglomerative models have recently emerged as a powerful approach to training vision foundation models, leveraging multi-teacher distillation from existing models such as CLIP, DINO, and SAM. This strategy enables th…

Cross-Domain Generalization Limits of Vision Foundation Models in Facial Deepfake Detection

2026-05-24 · Ibrahim Delibasoglu arxiv

The rapid evolution of generative models has enabled the creation of hyper-realistic facial deepfakes, exposing a critical vulnerability in modern digital forensics: the inability of detectors to generalize to unseen man…

Domain GeneralizationDeepFake Detection

Image Recognition with Vision and Language Embeddings of VLMs

2025-09-11 · Illia Volkov, Nikita Kisel, Klara Janouskova, Jiri Matas arxiv

Vision-language models (VLMs) have enabled strong zero-shot classification through image-text alignment. Yet, their purely visual inference capabilities remain under-explored. In this work, we conduct a comprehensive eva…

Image Classification

Duplicate Bug Report Detection With a Combination of Information Retrieval and Topic Modeling

2013-04-08 · 27th IEEE/ACM International Conference on Automated Software Engineering 2013 4 · Anh Tuan Nguyen, Tung Thanh Nguyen, Tien N. Nguyen, David Lo 외

Detecting duplicate bug reports helps reduce triaging efforts and save time for developers in fixing the same issues. Among several automated detection approaches, text-based information retrieval (IR) approaches have be…

DescriptiveInformation RetrievalRetrieval