ICML 2024: From Neurons to Neutrons

Main conference paper | Vienna, Austria

International Conference on Machine Learning 2024.

By the standards of theoretical-physics meetings, ICML 2024 was immense: it brought 9,095 in-person and virtual attendees to Vienna. In that cross-disciplinary setting, we presented our work, From Neurons to Neutrons: A Case Study in Interpretability, with Ouail Kitouni and Niklas Nolte.

Mechanistic Interpretability (MI) promises a path toward fully understanding how neural networks make their predictions. Prior work demonstrates that even when trained to perform simple arithmetic, models can implement a variety of algorithms (sometimes concurrently) depending on initialization and hyperparameters. Does this mean neuron-level interpretability techniques have limited applicability? Here, we argue that high-dimensional neural networks can learn useful low-dimensional representations of the data they were trained on, going beyond simply making good predictions. Such representations can be understood with the MI lens and provide insights that are surprisingly faithful to human-derived domain knowledge. This indicates that such approaches to interpretability can be useful for deriving a new understanding of a problem from models trained to solve it. As a case study, we extract nuclear physics concepts by studying models trained to reproduce nuclear data.

With Ouail Kitouni and Niklas Nolte at our ICML 2024 poster in Vienna.
With Ouail Kitouni and Niklas Nolte at our ICML 2024 poster in Vienna.

The conference page has further information about ICML 2024.

Sokratis Trifinopoulos
Sokratis Trifinopoulos
Research Associate

Fundamental Physics at the intersection of Quantum Field Theory and Artificial Intelligence.