paper-with-me

홈 › Papers

HessFormer: Hessians at Foundation Scale

2025-05-16 · Diego Granziol

Whilst there have been major advancements in the field of first order optimisation of deep learning models, where state of the art open source mixture of expert models go into the hundreds of billions of parameters, methods that rely on Hessian vector products, are still limited to run on a single GPU and thus cannot even work for models in the billion parameter range. We release a software package \textbf{HessFormer}, which integrates nicely with the well known Transformers package and allows for distributed hessian vector computation across a single node with multiple GPUs. Underpinning our implementation is a distributed stochastic lanczos quadrature algorithm, which we release for public consumption. Using this package we investigate the Hessian spectral density of the recent Deepseek $70$bn parameter model.

📄 PDF Abstract BibTeX arXiv:2505.11564

Code (0)

등록된 구현이 없습니다.

Tasks

GPU

Similar Papers 제목 키워드 기반

Chessformer: A Unified Architecture for Chess Modeling

2026-05-18 · Daniel Monroe, George Eilender, Philip Chalmers, Zhenwei Tang 외 arxiv

Chess has long served as a canonical testbed for artificial intelligence, but modeling approaches for its central tasks have diverged. Maximizing playing strength, predicting human play, and enabling interpretability are…

Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws

2025-10-19 · Egor Petrov, Nikita Kiselev, Vladislav Meshkov, Andrey Grabovoy arxiv

The lack of theoretical results for Layer Normalization and feedforward Hessians has left a gap in the study of Transformer optimization landscapes. We address this by deriving explicit second-order expressions for these…

Towards Fast, Specialized Machine Learning Force Fields: Distilling Foundation Models via Energy Hessians

2025-01-15 · Ishan Amin, Sanjeev Raja, Aditi Krishnapriyan

The foundation model (FM) paradigm is transforming Machine Learning Force Fields (MLFFs), leveraging general-purpose representations and scalable training to perform a variety of computational chemistry tasks. Although M…

Computational chemistryKnowledge Distillation

HIP: Hessian Interatomic Potentials without derivatives

2025-09-25 · Andreas Burger, Luca Thiede, Nikolaj Rønne, Varinia Bernales 외 arxiv

Molecular Hessians, the second derivatives of the potential energy, are fundamental to many workflows in computational chemistry. Usually, accurate Hessians are computationally expensive to calculate and scale poorly wit…

Local properties of neural networks through the lens of layer-wise Hessians

2025-10-20 · Maxim Bolshim, Alexander Kugaevskikh arxiv

We introduce a methodology for analyzing neural networks through the lens of layer-wise Hessian matrices. The local Hessian of each functional block (layer) is defined as the matrix of second derivatives of a scalar func…