paper-with-me

홈 › Papers

OceanCBM: A Concept Bottleneck Model for Mechanistic Interpretability in Ocean Forecasting

2026-05-12 · Sanah Suri, Kieran Ringel, Maike Sonnewald arxiv

Extreme ocean phenomena are challenging not only to predict but to diagnose, as accurate forecasts alone do not reveal the underlying physical drivers. While recent machine learning approaches achieve strong predictive skill, they remain largely opaque and provide limited guarantees of fidelity to ground-truth physics. We introduce OceanCBM, the first concept bottleneck model (CBM) for spatiotemporal prediction and mechanistic interrogation of ocean dynamics. OceanCBM uses mixed supervision to predict mixed layer heat content, a key precursor of marine heatwaves, while routing information through an intermediate layer of prescribed concepts derived from geophysical fluid dynamics and a 'free' concept. This design imposes soft physical structure without over-constraining the model, and the free concept both regularizes concept predictions and captures residual physical processes. Across ensemble initializations, we show that mixed supervision yields consistent mechanistic representations, whereas prediction-only and prescription-only baselines learn highly variable latent structures despite similar predictive performance. OceanCBM achieves interpretable, physically grounded representations without sacrificing skill, explicitly characterizing the interpretability-performance trade-off.

📄 PDF Abstract BibTeX arXiv:2605.12639

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Interpretability without actionability: mechanistic methods cannot correct language model errors despite near-perfect internal representations

2026-03-18 · Sanjay Basu, Sadiq Y. Patel, Parth Sheth, Bhairavi Muralidharan 외 arxiv

Language models encode task-relevant knowledge in internal representations that far exceeds their output performance, but whether mechanistic interpretability methods can bridge this knowledge-action gap has not been sys…

Interpretable and Steerable Concept Bottleneck Sparse Autoencoders

2025-12-11 · Akshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy, Shusen Liu 외 arxiv

Sparse autoencoders (SAEs) promise a unified approach for mechanistic interpretability, concept discovery, and model steering in LLMs and LVLMs. However, realizing this potential requires learned features to be both inte…

Image Generation

Learning Concept Bottleneck Models from Mechanistic Explanations

2026-03-07 · Antonio De Santis, Schrasing Tong, Marco Brambilla, Lalana Kagal arxiv

Concept Bottleneck Models (CBMs) aim for ante-hoc interpretability by learning a bottleneck layer that predicts interpretable concepts before the decision. State-of-the-art approaches typically select which concepts to l…

Knowledge Graphs

Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions

2026-06-27 · David Courtis, Ting Hu arxiv

Large Language Models (LLMs) have demonstrated the ability to simulate human-like OCEAN personality traits in generated text. Previous efforts have focused on prompt engineering or fine-tuning to shape LLM personality. I…

Prompt Engineering

Mechanistic Interpretability for AI Safety -- A Review

2024-04-22 · Leonard Bereska, Efstratios Gavves

Understanding AI systems' inner workings is critical for ensuring value alignment and safety. This review explores mechanistic interpretability: reverse engineering the computational mechanisms and representations learne…