13 training structures, 4 held out (struct_006, struct_011, struct_012, struct_016)
baseline kinetic functional: TF + 0.200 vW, xc = LDA

  target dv: std 1.735 eV over the training set
  (this is the uncorrected Euler-Lagrange residual)

HELD OUT (eV, std of the Euler-Lagrange residual)
  model                                  before      after      skill
  -------------------------------------------------------------------
  no correction (baseline)                1.712      1.712      1.00x
  semi-local least squares (GGA)          1.712      0.559      3.06x
    epoch   50  train 0.0025   held-out rel 0.3119
    epoch  100  train 0.0059   held-out rel 0.3175
    epoch  150  train 0.0003   held-out rel 0.3079
    epoch  200  train 0.0002   held-out rel 0.3084
    epoch  250  train 0.0000   held-out rel 0.3082
    epoch  300  train 0.0000   held-out rel 0.3082
    epoch  350  train 0.0000   held-out rel 0.3081
    epoch  400  train 0.0000   held-out rel 0.3081

  FNO (w24 m10 L4)                        1.712      0.529      3.24x

    ... same model, TRAINING fit          1.735      0.002   1029.92x

  9,228,025 parameters, 1210 s
