13 training structures, 4 held out (struct_006, struct_011, struct_012, struct_016)
baseline kinetic functional: TF + 0.333 vW, xc = PBE

  target dv: std 1.753 eV over the training set
  (this is the uncorrected Euler-Lagrange residual)

HELD OUT (eV, std of the Euler-Lagrange residual)
  model                                  before      after      skill
  -------------------------------------------------------------------
  no correction (baseline)                1.675      1.675      1.00x
  semi-local least squares (GGA)          1.675      0.883      1.90x
    epoch   50  train 0.0107   held-out rel 0.4700
    epoch  100  train 0.0061   held-out rel 0.4713
    epoch  150  train 0.0030   held-out rel 0.4686
    epoch  200  train 0.0001   held-out rel 0.4679
    epoch  250  train 0.0000   held-out rel 0.4680
    epoch  300  train 0.0000   held-out rel 0.4679
    epoch  350  train 0.0000   held-out rel 0.4679
    epoch  400  train 0.0000   held-out rel 0.4679

  FNO (w24 m10 L4)                        1.675      0.788      2.13x

    ... same model, TRAINING fit          1.748      0.002    920.92x

  9,228,025 parameters, 1195 s
