A machine-learned interatomic potential (MLIP) is a regression of the Born–Oppenheimer potential-energy surface: a flexible function fit to a reference dataset of DFT energies, forces, and stresses, evaluated at near-classical-MD cost while approaching DFT accuracy inside the training manifold. The two design axes that matter are how local atomic environments are represented, and how the model is trained and grown.
Locality, invariance, equivariance
Almost all MLIPs assume a finite cutoff : the total energy is a sum of atomic contributions , each depending only on neighbours within . The representation must respect the symmetries of the PES — invariance of under translation, rotation , and permutation of like atoms — while forces and stresses (and any internal geometric features) transform equivariantly. This invariance/equivariance distinction is the main fault line between families.
Descriptor-based (invariant features → regression). Hand-built rotationally invariant descriptors of each environment feed a regressor:
- GAP/SOAP — the Smooth Overlap of Atomic Positions power spectrum as descriptor, Gaussian-process regression as the model;
- ACE (the Atomic Cluster Expansion) — a complete, systematically improvable basis of body-ordered invariants (-body terms via products of atomic-base functions), with a linear or shallow model; the unifying formalism that GAP/SNAP/MTP are special cases of.
Equivariant message passing (learned features). Graph neural networks pass messages between atoms; features carry an angular-momentum label and transform under irreps, so internal representations are equivariant rather than invariant:
- NequIP — -equivariant convolutions on the atomic graph, tensor products of spherical harmonics with learned radial channels;
- MACE — folds the higher body order of ACE into the message-passing layers, so a small number of layers reaches high body order; the current workhorse for accuracy-per-cost and the basis of several foundation models.
Higher body order and equivariance buy data efficiency and smoothness; they do not buy physics that is absent from the data.
Training on energies, forces, and stress
Force training is essential, not optional. The forces are , so fitting them constrains the gradient of the surface — exactly the quantity dynamics, phonons, and relaxations depend on — and supplies constraints per configuration against the single energy constraint. The loss is the weighted multi-target form
with stress targets required for any application at variable cell (equations of state, pressure, thermal expansion). An energy-only fit can have an excellent energy RMSE and badly wrong forces — the relevant failure for everything downstream.
Active learning
Datasets are grown, not assembled once. Active learning runs the potential, flags configurations where it is uncertain, recomputes those with DFT, and refits — iterating until the model is confident over the sampled region. Uncertainty estimators differ by family: GP posterior variance for GAP; query-by-committee (ensemble spread) for neural potentials; the MTP/ACE extrapolation grade (D-optimality / leverage score) that measures how far a feature vector lies outside the training set’s linear span. The point is to sample the relevant part of configuration space (e.g. the high-temperature anharmonic region for phonons) rather than to blanket it.
Extrapolation failure modes
The defining weakness: an MLIP is smooth but wrong outside its training manifold. A neural potential does not blow up off-distribution — it interpolates its features into a confident, plausible, incorrect surface, which is far more dangerous than an obvious error. Specific traps:
- No physics beyond the cutoff. Long-range electrostatics, dispersion, and polarization are absent unless added explicitly; a purely local model cannot represent them and cannot know it is missing them.
- Reactive / rare regions undersampled. Bond breaking, defects, and transition states are sparse in routine MD-sampled data; the surface there is whatever the smoothness prior invented.
- Silent extrapolation. Without an uncertainty signal, the model gives no warning when it has left the data — hence active learning’s uncertainty gate is the safety mechanism, not a nicety.
The DC-DFT caveat at the source
The reference labels are themselves DFT, so the potential inherits the reference functional’s errors — including its density-driven error. For a density-sensitive (“abnormal”) system, the self-consistent density is wrong, the labelled energies and (worse) forces are wrong in a correlated way, and the MLIP faithfully learns a corrupted PES. This is invisible to energy-only validation and is the cleanest example of garbage-in: a good fit to bad labels. The fix is at the data level — use a density-corrected (e.g. HF-DFT) reference where density sensitivity is diagnosed, per density-driven error (DC-DFT) — and it propagates directly into learned force constants, the subject of MLIP phonons.
Prerequisites
The Born–Oppenheimer potential-energy surface · Hellmann–Feynman forces · DFT as the reference labelling method · density-driven error (DC-DFT)
Builds toward: force constants & anharmonic sampling from MLIPs
Key references
- Foundational potentials — J. Behler & M. Parrinello, PRL 98, 146401 (2007) (HD-NNP); A. P. Bartók, M. C. Payne, R. Kondor & G. Csányi, PRL 104, 136403 (2010) (GAP).
- Reviews — J. Behler, J. Chem. Phys. 145, 170901 (2016); V. L. Deringer, M. A. Caro & G. Csányi, Adv. Mater. 31, 1902765 (2019); O. T. Unke et al., Chem. Rev. 121, 10142 (2021).
- Descriptors (SOAP, ACE, MTP, SNAP) — A. P. Bartók, R. Kondor & G. Csányi, PRB 87, 184115 (2013); R. Drautz, PRB 99, 014104 (2019) (ACE); A. V. Shapeev, Multiscale Model. Simul. 14, 1153 (2016) (MTP); A. P. Thompson et al., J. Comput. Phys. 285, 316 (2015) (SNAP).
- Equivariant message passing — K. T. Schütt et al., J. Chem. Phys. 148, 241722 (2018) (SchNet); S. Batzner et al., Nat. Commun. 13, 2453 (2022) (NequIP); I. Batatia, D. P. Kovács, G. Simm, C. Ortner & G. Csányi, Adv. Neural Inf. Process. Syst. 35, 11423 (2022) (MACE); A. Musaelian et al., Nat. Commun. 14, 579 (2023) (Allegro).
- Training & active learning — E. V. Podryabinkin & A. V. Shapeev, Comput. Mater. Sci. 140, 171 (2017) (extrapolation grade); R. Jinnouchi, J. Lahnsteiner, F. Karsai, G. Kresse & M. Bokdam, PRL 122, 225701 (2019) (on-the-fly Bayesian).
- Long-range physics & locality limits — A. Grisafi & M. Ceriotti, J. Chem. Phys. 151, 204105 (2019) (LODE); T. W. Ko, J. A. Finkler, S. Goedecker & J. Behler, Nat. Commun. 12, 398 (2021) (4G-HDNNP).
- Foundation models — C. Chen & S. P. Ong, Nat. Comput. Sci. 2, 718 (2022) (M3GNet); B. Deng et al., Nat. Mach. Intell. 5, 1031 (2023) (CHGNet); I. Batatia et al., arXiv:2401.00096 (2024) (MACE-MP-0).
- Density-driven error in the labels — M.-C. Kim, E. Sim & K. Burke, PRL 111, 073003 (2013).