arms: ['REF', 'REF2', 'STD', 'INT8', 'INT8G', 'INT8G256', 'INT8GN', 'SIMROW', 'SIMG128', 'SIMG256', 'SIMG512']  prompts: 20  positions/prompt: [8, 8, 8] ...
  TV    SIMROW vs STD  mean  2.8497 pp  n=160
  TV   SIMG128 vs STD  mean  1.3488 pp  n=160
  TV   SIMG256 vs STD  mean  2.3982 pp  n=160
  TV   SIMG512 vs STD  mean  1.2376 pp  n=160
  TV      INT8 vs STD  mean  3.2851 pp  n=160
  TV     INT8G vs STD  mean  2.6943 pp  n=160
  TV  INT8G256 vs STD  mean  2.8819 pp  n=160
  TV    INT8GN vs STD  mean  2.7129 pp  n=160
  TV       REF vs STD  mean  1.4165 pp  n=160
  TV      REF2 vs REF  mean  0.0000 pp  n=160

PAIRED DIFFERENCES, mean TV pp, 95% cluster bootstrap over prompts.
A negative interval that excludes 0 is a real reduction.
  SIMG128-STD  minus  SIMROW-STD      -1.5009 pp [-3.7387, -0.2299]   what per-128-group activation scaling buys over the shipped per-token amax
  SIMG32-STD   minus  SIMG128-STD    SKIPPED (arm missing)
  SIMROW-STD   minus  INT8-STD        -0.4355 pp [-1.4403, +0.3650]   NULL: does the torch simulator reproduce the real kernel?
  INT8NOGDN-STD minus  INT8-STD      SKIPPED (arm missing)
  SIMSPLIT8-STD minus SIMROW-STD     SKIPPED (arm missing)
  SIMSPLIT8-STD minus SIMG128-STD    SKIPPED (arm missing)
  SIMSPLIT2-STD minus SIMG128-STD    SKIPPED (arm missing)
  SIMG128-STD  minus  REF-STD         -0.0677 pp [-0.4302, +0.2736]   is grouped int8 inside the leg cloud? <=0 means at or below it
  INT8-STD     minus  REF-STD         +1.8686 pp [+0.5788, +3.7197]   the shipped kernel against the same cloud
  INT8G-STD    minus  INT8-STD        -0.5908 pp [-1.2440, -0.0577]   THE HEADLINE: what per-group scaling buys ON THE REAL KERNEL
  INT8G-STD    minus  SIMG128-STD     +1.3455 pp [+0.1804, +3.0730]   NULL: did the torch PROJECTION survive implementation? 0 means yes
  INT8GN-STD   minus  INT8-STD        -0.5722 pp [-1.6901, +0.1231]   NULL: the grouped kernel with no grouping in its table must be INT8
  INT8GN-STD   minus  INT8G-STD       +0.0186 pp [-1.0093, +0.8676]   SEPARATION: >0 and clear of 0, or the treatment is its own control
  SIMG256-STD  minus  SIMG128-STD     +1.0494 pp [+0.0681, +2.5874]   what a 256-wide group gives up; the GEMM drains half as often
  SIMG512-STD  minus  SIMG128-STD     -0.1112 pp [-0.4059, +0.1946]   same at 512 -- a quarter of the drains
  SIMG512-STD  minus  SIMROW-STD      -1.6121 pp [-3.7966, -0.3904]   is a 512-wide group still worth having over the shipped per-token amax?
  INT8G256-STD minus  INT8G-STD       +0.1876 pp [-0.6386, +1.1307]   what the 256-wide group costs ON THE KERNEL; it halves the GEMM's drains
  INT8G256-STD minus  SIMG256-STD     +0.4837 pp [-0.3855, +1.5869]   NULL: the projection at 256, against its kernel
  INT8G-STD    minus  REF-STD         +1.2778 pp [+0.2587, +2.8027]   is the BUILT grouped kernel inside the leg cloud? <=0 means at or below
