aligned: 10 prompts x 64 teacher-forced decode steps, scored from step 1 (63 positions per prompt); every file decoded A:REF's tokens; prefill positions [8, 8, 8, 8, 8, 8, 8, 8, 8, 8]

NULL CONTROLS (must read exactly 0 at every position):
  A:REF2 vs A:REF   KL median 0.000e+00 max 0.000e+00  (n=630)
  B:REF  vs A:REF   KL median 0.000e+00 max 0.000e+00  (n=630)   <- cross-boot

KL vs A:REF at teacher-forced DECODE positions (the precision doc's columns):
  arm                median        p95        max   argmax flips
  STD             1.937e-04  4.968e-03  6.136e-02       6/630   
  INT8            2.557e-04  7.586e-03  2.699e-01       5/630   
  INT8G1          2.202e-04  7.873e-03  6.445e-02       5/630   
  k4v4 (the bar)  1.327e-03  2.527e-02  1.390e+01      14/630   

residual rms entering lm_head, relative, vs A:REF at the same positions:
  arm                median        p95        max
  STD             2.333e-02  1.529e-01  6.794e-01
  INT8            3.398e-02  1.672e-01  6.635e-01
  INT8G1          2.861e-02  1.830e-01  7.149e-01
  k4v4 (the bar)  6.944e-02  3.252e-01  1.141e+00

THE BAR -- arm / k4v4, paired by position, cluster bootstrap over 10 prompts x 10000:
  arm              paired ratio median [95% CI]        ratio of medians [95% CI]   verdict
  STD             0.191 [0.154, 0.214]              0.146 [0.084, 0.192]   INSIDE (KL)
                  0.343 [0.320, 0.369]              0.336 [0.315, 0.362]   INSIDE (residual rms)
  INT8            0.308 [0.254, 0.383]              0.192 [0.147, 0.311]   INSIDE (KL)
                  0.468 [0.411, 0.591]              0.489 [0.401, 0.682]   INSIDE (residual rms)
  INT8G1          0.253 [0.215, 0.312]              0.166 [0.115, 0.228]   INSIDE (KL)
                  0.407 [0.375, 0.450]              0.412 [0.364, 0.501]   INSIDE (residual rms)
