paper-with-me

홈 › Papers

Do Foundation Models Know Geometry? Probing Frozen Features for Continuous Physical Measurement

2026-03-06 · Yakov Pyotr Shkolnikov arxiv

Vision-language models encode continuous geometry that their text pathway fails to express: a 6,000-parameter linear probe extracts hand joint angles at 6.1 degrees MAE from frozen features, while the best text output achieves only 20.0 degrees -- a 3.3x bottleneck. LoRA fine-tuning (r=16, 2,000 images) narrows this gap to 6.5 degrees, providing evidence for a pathway-training deficit rather than a representational one. Training objective determines accuracy more than architecture: five encoders spanning self-supervised, contrastive, and hybrid paradigms converge to statistically equivalent accuracy (R^2 approximately 0.55, TOST-equivalent at delta=0.03) despite sharing as little as CKA=0.41 representational similarity -- functional convergence without representational convergence. Autoregressive generation damages geometric fidelity, but the damage originates in the generation process, not in language alignment: Qwen2.5-VL's LLM layers actually improve probe accuracy over its raw vision encoder. Layer-wise analysis reveals a universal mid-network accuracy peak across all architectures, with attention heads in layers 18-22 carrying disproportionate geometric signal. These findings enable a single frozen backbone to function as a multi-task geometric sensor through lightweight probes, without fine-tuning or text generation.

📄 PDF Abstract BibTeX arXiv:2603.06459

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

2026-05-27 · Haozhan Shen, Tiancheng Zhao, Kangjia Zhao, Jianwei Yin arxiv

Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schemes are now widely used as foundation bac…

3D Geometry PredictionVideo Generation

Discovery of a Hematopoietic Manifold in scGPT Yields a Method for Extracting Performant Algorithms from Biological Foundation Model Internals

2026-03-10 · Ihor Kendiukhov arxiv

We report the discovery and extraction of a compact hematopoietic algorithm from the single-cell foundation model scGPT, to our knowledge the first biologically useful, competitive algorithm extracted from a foundation m…

DeeperBrain: A Neuro-Grounded EEG Foundation Model Towards Universal BCI

2026-01-05 · Jiquan Wang, Sha Zhao, Yangxuan Zhou, Yiming Kang 외 arxiv

Electroencephalography (EEG) foundation models hold significant promise for universal Brain-Computer Interfaces (BCIs). However, existing approaches often rely on end-to-end fine-tuning and exhibit limited efficacy under…

Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis

2026-06-08 · Samuele Punzo, Niccolò Caselli, Ippokratis Pantelidis, Francesco Massafra 외 arxiv

We study whether pretrained video foundation models encode intuitive-physics information in their frozen representations, and how this information varies across model families, layers, and probe types. Using frozen-featu…

Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models

2025-12-01 · Benjamin Ramtoula, Pierre-Yves Lajoie, Paul Newman, Daniele De Martini arxiv

Foundation models (FMs) trained with different objectives and data learn diverse representations, making some more effective than others for specific downstream tasks. Existing adaptation strategies, such as parameter-ef…

parameter-efficient fine-tuning