A cheap, local correction that bends an approximate exchange–correlation functional back toward the right answer on a chosen localized subspace — usually the open or shell of a transition-metal or rare-earth ion. The whole construction is a targeted patch for one disease, delocalization error, applied only where that disease bites hardest and a full hybrid or many-body treatment would be unaffordable. It is the workhorse of correlated-oxide DFT, and also the method whose single tunable number — — is most routinely abused.

Why localized electrons need it

Semilocal functionals are convex in fractional electron number: they spuriously lower the energy of a non-integer occupation, so an electron prefers to delocalize and smear its fractional charge rather than sit as an integer on one site (this is the delocalization-error half of SIE). For dispersive bands the error is mild and the self-interaction is partly screened by the kinetic-energy gain of delocalization. For the near-atomic manifold it is severe: those states are spatially compact, so the un-cancelled self-Hartree repulsion of an electron with itself is large, and the curvature of on the subspace is strongly convex. The visible symptoms are the textbook failures of plain LDA/GGA on correlated insulators — gaps far too small or absent, magnetic moments quenched, and the charge-transfer / Mott insulators NiO, MnO, FeO, and CoO coming out metallic or as tiny-gap semiconductors. The fix has to raise the energy of the partially-occupied, delocalized solution relative to the integer-occupied, localized one — i.e. add back the missing on-site Coulomb discontinuity. That is exactly what a Hubbard-like penalty does.

The occupation matrix and its projector ambiguity

Everything is built on the on-site occupation matrix, the density matrix of the Kohn–Sham state projected onto a set of localized orbitals (angular index ) on atom :

with the occupations of the KS states . The eigenvalues of are sub-shell occupations in ; its trace is the total spin- count in the correlated shell. A fully occupied or empty shell has idempotent (, eigenvalues or ); fractional eigenvalues measure exactly the hybridization-induced departure from integer occupation that the correction targets.

The catch, and the reason no two codes agree on a number, is that is not an observable — it depends entirely on the choice of projector . Common choices:

  • Atomic / pseudo-atomic orbitals — the free-ion or pseudized radial function. Simple, but the result depends on the (somewhat arbitrary) radial cutoff and normalization.
  • PAW / LAPW projectors — the partial-wave projectors inside the augmentation sphere; tied to the sphere radius and the pseudopotential’s reference configuration.
  • Wannier functions — maximally-localized Wannier functions spanning the correlated manifold. The most physically defensible (they actually represent the in-solid orbital), and the natural choice when the same feeds DMFT, but more work to construct.

Because a more diffuse projector captures less of the in-solid charge, the same physical correction requires a different numerical for each projector. A from one code’s atomic projectors is not portable to another’s PAW projectors. The single most important hygiene rule in DFT+U is therefore to quote together with the projector that defines it, and to compute with the same projector it will be used with (below).

The Dudarev functional (rotationally invariant, single )

The form I reach for first, because it is rotationally invariant in the subspace and needs only one number. Dudarev et al. fold and into a single and write a correction that penalizes any departure from idempotency:

where are the eigenvalues of . The term vanishes for any integer (idempotent) occupation and is maximal at half-filling ; minimizing the total energy therefore drives the eigenvalues toward , i.e. toward the localized, integer-occupied solution. The corresponding potential is

a level shift that pushes occupied subspace states () down by and empty ones () up by — manufacturing precisely the on-site gap (the would-be Hubbard splitting) that semilocal DFT failed to open. This is the form that makes the SIE connection most transparent: is a local convexity penalty, the subspace analogue of flattening the convex curve, and it is exactly the term I cite in the SIE note as DFT+U’s contribution to piecewise-linearity restoration.

The full Liechtenstein form ( and separately)

The Dudarev functional is a spherical average; the more general Liechtenstein–Anisimov–Zaanen construction keeps the full anisotropy of the on-site interaction. One writes the mean-field energy of a screened multi-orbital Hubbard interaction on the subspace,

where the antisymmetric (exchange) second term acts only within a spin channel. The matrix elements are expanded in Slater integrals via Gaunt coefficients (for electrons ; for electrons also ), and the physically meaningful averages are

with the ratios (and ) fixed close to their atomic values, so that specifying determines the whole tensor. Keeping and separate matters when the multipole / Hund’s-rule structure is physics, not detail: orbital ordering, the relative stability of high- vs. low-spin states, and the anisotropy of the exchange splitting all live in and in the integrals that the Dudarev average discards. The price is the genuine symmetry-breaking freedom of a full (or ) occupation matrix, which is also where DFT+U’s notorious metastable orbital-occupation minima come from: the SCF can converge to the wrong orbital ordering and sit there. Occupation-matrix control (initializing and, if needed, constraining ) is the standard remedy. In the limit where the anisotropy is averaged away the Liechtenstein energy reduces to the Dudarev expression with — the two are not rival theories, but the anisotropic form and its spherical average.

Double counting: FLL vs. AMF

Both functionals add a Hubbard energy on top of an approximate that already contains some of that interaction in an averaged way. The added must therefore subtract a double-counting (DC) term for the mean interaction the base functional already holds:

and because is not a Hubbard model there is no exact — this is a real ambiguity, not bookkeeping, and it changes where the correction bites. Two limits are in use:

  • Fully-localized limit (FLL / “around-atomic-limit”). Assumes the subspace is close to integer occupation, subtracting the atomic-limit mean-field energy , with . This is the limit implicitly built into the Dudarev functional, and the right choice for localized, near-integer systems: Mott and charge-transfer insulators, rare-earth shells. It maximally favours integer occupation and can over-localize.
  • Around-mean-field (AMF). Subtracts the energy of a uniformly occupied shell (each orbital at the average filling ). Appropriate when the system is weakly correlated / itinerant and the occupations are genuinely close to uniform — early transition metals, metallic correlated systems — where FLL would spuriously open a gap.

FLL is the common default for insulating oxides; AMF (or interpolations between the two) is the safer choice when the shell is far from integer filling. The DC choice can swing computed gaps and even the sign of relative phase stabilities, so it is part of the method specification, not a hidden knob.

is not universal — and how to fix it ab initio

The most important conceptual point: there is no transferable . The effective depends on (i) the projector that defines , (ii) the underlying functional (LDA vs. PBE vs. SCAN screen differently), (iii) the structure and coordination (bond lengths, pressure, the chemical environment set the screening), and (iv) the oxidation / spin state (the same element needs a different as, e.g., a vs. cation). Taking a literature for “Fe in oxides” and reusing it across these axes is the single most common methodological error. There are two principled, system-specific routes to compute from first principles rather than fitting it to experiment:

  • Linear-response / constrained DFT (Cococcioni–de Gironcoli). Define as the curvature the exact functional should have but the approximate one lacks. Apply a small potential shift to the subspace on site and measure the occupation response. The bare (non-interacting, Kohn–Sham re-screening) and interacting responses are and , and the difference removing the trivial rehybridization so that what remains is the missing on-site electron–electron curvature. This is internally consistent — the is computed with the very projector and functional it will be used with — and it is the route I default to when the value matters to the physics. A subtlety: applied to the converged ground state it should be iterated to self-consistency in , since adding changes the screening that sets .
  • Constrained RPA (cRPA). Compute the screened on-site interaction directly as the partially-screened Coulomb interaction in which screening processes within the correlated subspace are explicitly excluded (so they are not double-counted by the subsequent Hubbard treatment): , with the polarizability of everything but the subspace. cRPA naturally yields the full frequency-dependent and the inter-orbital matrix, and is the standard provider of the Hubbard parameters that feed DFT+DMFT. Its weak point is the partitioning when the correlated and screening manifolds are strongly entangled (e.g. bands overlapping the ligand band), where the disentanglement choice noticeably moves .

Whichever route, the discipline is the same: a computed, projector-matched, state-specific — not a borrowed constant — and an honest statement of the double-counting choice.

What it is not

DFT+U is a static, on-site mean-field correction. It opens the right gaps and stabilizes the right moments and orbital orderings for the price of a near-zero-cost addition to the SCF, which is why it is the default for high-throughput correlated-oxide work and for supplying the starting point and the Hubbard parameters to DMFT. But it carries no frequency dependence and no dynamical correlation: it cannot produce a Mott transition driven by spectral-weight transfer, quasiparticle renormalization, or satellite structure — those need the explicit self-energy of DMFT or GW. And because the correction lives on the chosen subspace, it inherits all the projector and double-counting ambiguities above. It is the cheapest rung that breaks the semilocal SIE barrier, and it is best understood as exactly that.

Prerequisites

Builds toward: delocalization error · the beyond-DFT ladder · exchange tensors from DFT (the -dependence of extracted ) · density-driven error (DC-DFT)

Key references

  • LDA+U, original — V. I. Anisimov, J. Zaanen & O. K. Andersen, Phys. Rev. B 44, 943 (1991).
  • Rotationally-invariant (full ) form — A. I. Liechtenstein, V. I. Anisimov & J. Zaanen, Phys. Rev. B 52, R5467 (1995).
  • Simplified single- (Dudarev) form — S. L. Dudarev, G. A. Botton, S. Y. Savrasov, C. J. Humphreys & A. P. Sutton, Phys. Rev. B 57, 1505 (1998).
  • Linear-response — M. Cococcioni & S. de Gironcoli, Phys. Rev. B 71, 035105 (2005).
  • cRPA for the screened interaction — F. Aryasetiawan, M. Imada, A. Georges, G. Kotliar, S. Biermann & A. I. Lichtenstein, Phys. Rev. B 70, 195104 (2004).
  • Double counting — around-mean-field (AMF) — M. T. Czyżyk & G. A. Sawatzky, Phys. Rev. B 49, 14211 (1994) (introduces the AMF double-counting form). The fully-localized-limit (FLL) / atomic-limit form belongs to the original LDA+U line — Anisimov, Zaanen & Andersen (1991) and Liechtenstein, Anisimov & Zaanen (1995), above.