[i8] basket: 6 prompts x 16384 tok, domains ['code', 'longctx', 'math', 'sci', 'wiki']
boot: this card's residency could NOT be split between this process and any other — neither NVML's per-process walk nor a pre-context mark was available — so driver.cuda_context is being booked from the DEVICE-WIDE reading (397 MiB). If anything else is resident on this device, its bytes are inside that number and the KV pool is sized that much smaller for the life of this boot.
compile-cache MISS: xgrammar — xgrammar compiled grammars (re)compiles this boot [/cache/xgrammar]
VRAM at profile: this card's residency could NOT be split between this process and any other — NVML's per-process walk did not identify us — so the whole 14.26 GiB resident on device 0 is being booked as driver.cuda_context. Any co-tenant's bytes are inside that number and KV is sized that much smaller for this boot.
[TKV autotune-provenance] fp=f4e763cb2f31eb74 regime=cache-hit cells=8 sweep_src=b5332b8dfc1c8389 tile_tokens=4 num_splits=32,64,128,256 min_blocks_per_sm=0 resolved=0 held=0 inherited=0
  return isinstance(obj, torch.Tensor)
[i8] trellis->path map: 401 linears
JIT compile AFTER serving-ready [cpp_ext]: exl3_i8_gemm_k4_cb2 cached .so load — a live request paid this compile's latency. This is a boot-warmup coverage gap: extend warmup to pre-compile this kernel/specialization. Counter jit_compile_serving (must-not-fire) at GET /v1/admin/flag_truth.
[i8] prefix cache: server default resolved=True, per-request cache_enabled=False; any cache hit below is refused
JIT compile AFTER serving-ready [cpp_ext]: arbi_serve_exl3_a_prep_v3 cached .so load — a live request paid this compile's latency. This is a boot-warmup coverage gap: extend warmup to pre-compile this kernel/specialization. Counter jit_compile_serving (must-not-fire) at GET /v1/admin/flag_truth.
activation admission RE-ARMED (JIT compile after serving-ready): the step budget FELL to 6381 MiB, -2 MiB on the 6383 MiB it was enforcing and 6383 MiB the boot reading armed. The widest slate admission will build is narrower from here — narrower prefill chunks and deferred rows instead of a forward that OOMs after the layout freeze.
[i8] warmup done
  [0] code           REF      16384tok  wall=  5.03s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [0] code           REF2     16384tok  wall=  5.03s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [0] code           INT8FO   16384tok  wall=  3.32s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [0] code           INT8GFO  16384tok  wall=  3.51s  finish=length  leg=0F/3200S acc=3200x16/0x32
driver serving-growth OVER RESERVE: the driver holds 2 MiB more than it did when the boot brackets closed, against the 0 MiB the serving floor's transient.serving_step.driver_growth term held for it. Those bytes are in no pool and no allocator counter -- the live card reaches them as driver.modules_serving and the boot ledger only as driver.residual -- and the KV layout is frozen, so the overage comes out of the free VRAM the floor is holding for one in-flight step (the verify tail, the DFlash context-assemble) and those reserves are the ones that fail first. The reading is on record for this configuration, so the next boot holds it; this process serves the rest of its life short by the overage.
activation admission RE-ARMED (driver residency over reserve): the step budget FELL to 6313 MiB, -68 MiB on the 6381 MiB it was enforcing and 6383 MiB the boot reading armed. The widest slate admission will build is narrower from here — narrower prefill chunks and deferred rows instead of a forward that OOMs after the layout freeze.
  [1] code           REF      16384tok  wall=  5.04s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [1] code           REF2     16384tok  wall=  5.05s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [1] code           INT8FO   16384tok  wall=  3.32s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [1] code           INT8GFO  16384tok  wall=  3.52s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [2] wiki           REF      16384tok  wall=  5.06s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [2] wiki           REF2     16384tok  wall=  5.07s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [2] wiki           INT8FO   16384tok  wall=  3.33s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [2] wiki           INT8GFO  16384tok  wall=  3.53s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [3] math           REF      16384tok  wall=  5.08s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [3] math           REF2     16384tok  wall=  5.09s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [3] math           INT8FO   16384tok  wall=  3.33s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [3] math           INT8GFO  16384tok  wall=  3.53s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [4] longctx        REF      16384tok  wall=  5.09s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [4] longctx        REF2     16384tok  wall=  5.10s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [4] longctx        INT8FO   16384tok  wall=  3.34s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [4] longctx        INT8GFO  16384tok  wall=  3.54s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [5] sci            REF      16384tok  wall=  5.10s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [5] sci            REF2     16384tok  wall=  5.10s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [5] sci            INT8FO   16384tok  wall=  3.34s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [5] sci            INT8GFO  16384tok  wall=  3.54s  finish=length  leg=0F/3200S acc=3200x16/0x32

[i8] kernel census: int8 calls=38400 rows=78643200 reconstructs SKIPPED=38400 fallback (non-4bpw) calls=0
[i8] shapes served by the kernel (K,N)->calls: {(5120, 1024): 3072, (5120, 6144): 4608, (5120, 10240): 4608, (5120, 12288): 1536, (5120, 17408): 12288, (6144, 5120): 6144, (17408, 5120): 6144}
[i8] excluded projections: none  excluded calls=0
...
  model.layers.9.linear_attn.in_proj_qkv                     98      0             -            -            -
  model.layers.9.linear_attn.in_proj_z                       98      0             -            -            -
  model.layers.9.linear_attn.out_proj                        98      0             -            -            -
  model.layers.9.mlp.down_proj                               98      0             -            -            -
  model.layers.9.mlp.gate_proj                               98      0             -            -            -
  model.layers.9.mlp.up_proj                                 98      0             -            -            -
[i8] per-arm census (int8 kernel calls): {'REF': 0, 'REF2': 0, 'INT8FO': 19200, 'INT8GFO': 19200}

ARM RECEIPTS -- what each arm actually executed, from the engine's
own per-call census.  fused/standalone is the reconstruct variant;
acc16/acc32 is the cuBLAS compute type leg B asked hgemm for.
  arm       legB calls          rows     fused  standalone     acc16     acc32   fp32pin
  REF            19200      39321600     19200           0     19200         0         0
  REF2           19200      39321600     19200           0     19200         0         0
  INT8FO         19200      39321600         0       19200     19200         0         0
  INT8GFO        19200      39321600         0       19200     19200         0         0
  all arms match their expected leg/accumulator signature
  prefill positions aligned: [8, 8, 8, 8, 8, 8] per prompt, identical across all 4 arms
[i8] null control: REF2 vs REF max KL is exactly 0

THE TAIL over PREFILL positions only -- one per prefill chunk, so
every arm scores the SAME context. Decode positions are reported
separately and are NOT a numerics measure: once a token flips, the
arms are reading different contexts. The control proves it -- pooled
over decode, STD (which contains no int8 at all) has the same max KL
as INT8, so a pooled tail cannot attribute anything to the kernel.

arm            n        mean         p50         p95         p99         MAX     argmax flip
REF           48   0.000e+00   0.000e+00   0.000e+00   0.000e+00   0.000e+00 0/48 =   0.00%
REF2          48   0.000e+00   0.000e+00   0.000e+00   0.000e+00   0.000e+00 0/48 =   0.00%
INT8FO        48   1.224e-01   8.966e-04   3.133e-01   4.259e+00   4.259e+00 2/48 =   4.17%
INT8GFO       48   1.023e-01   7.028e-04   2.343e-01   3.841e+00   3.841e+00 1/48 =   2.08%

Decode positions (divergence-contaminated, shown for completeness):
arm            n        mean         MAX    greedy agree
REF           48   0.000e+00   0.000e+00 48/48 =  100.0%
REF2          48   0.000e+00   0.000e+00 48/48 =  100.0%
INT8FO        48   4.465e+00   3.254e+01 40/48 =   83.3%
INT8GFO       48   1.783e-03   1.695e-02 48/48 =  100.0%

SAMPLER-RELEVANT TAIL @ canonical (T=1.0, top_p=0.95, top_k=20), prefill positions only.
TV is the headline: under a drafter that tracks the target, speculative
decoding's expected accept rate is sum_t min(p_ref, p_arm), so TV IS the
expected accept-rate loss.
arm            n  setOverlap    TV mean     TV p95     TV max  truncKL p95  bndryChurn  drawDiff
REF           48     100.00%  0.000e+00  0.000e+00  0.000e+00    0.000e+00 0/ 48     0.00%
REF2          48     100.00%  0.000e+00  0.000e+00  0.000e+00    0.000e+00 0/ 48     0.00%
INT8FO        48      98.29%  4.575e-02  3.620e-01  7.058e-01          inf 11/ 48     4.43%
INT8GFO       48      98.18%  4.945e-02  2.842e-01  8.840e-01          inf 10/ 48     7.81%

SAMPLER-RELEVANT TAIL @ owner's k~40 (T=1.0, top_p=1.0, top_k=40), prefill positions only.
TV is the headline: under a drafter that tracks the target, speculative
decoding's expected accept rate is sum_t min(p_ref, p_arm), so TV IS the
expected accept-rate loss.
arm            n  setOverlap    TV mean     TV p95     TV max  truncKL p95  bndryChurn  drawDiff
REF           48     100.00%  0.000e+00  0.000e+00  0.000e+00    0.000e+00 0/ 48     0.00%
REF2          48     100.00%  0.000e+00  0.000e+00  0.000e+00    0.000e+00 0/ 48     0.00%
INT8FO        48      93.53%  4.584e-02  3.459e-01  6.971e-01          inf 41/ 48     6.25%
INT8GFO       48      95.49%  4.856e-02  2.723e-01  8.649e-01          inf 33/ 48     9.64%

FULL-VOCAB numbers below, kept and labelled: they are the right
instrument for GREEDY or beam decode, where every rank matters.

arm         first-prefill-position KL       median          max
REF                         0.000e+00    0.000e+00    0.000e+00
REF2                        0.000e+00    0.000e+00    0.000e+00
INT8FO                      5.704e-04    3.044e-05    3.246e-03
INT8GFO                     2.510e-04    4.335e-05    1.298e-03

PAIRED TV @ canonical (T=1.0, top_p=0.95, top_k=20), prefill positions.
Each row is A scored against B -- B is that ROW's reference, not the
table's.  TV is the expected speculative accept-rate loss, so the
column is read directly in percentage points.  The interval is a
cluster bootstrap over PROMPTS; the IQR is the raw spread over
positions with no floor applied to it.

  A       vs B        n  mean TV pp           95% CI pp    p25 pp    p50 pp    p75 pp    p95 pp    max pp
  REF2    REF        48      0.0000 [  0.0000,  0.0000]    0.0000    0.0000    0.0000    0.0000    0.0000

lm_head CUT -- the divergence entering the head vs leaving it,
prefill positions, relative rms.  'in' is the residual stream after
24 layers; 'out' is the centred logits.  amp = out/in: at 1.0 the
head passes the body's divergence through unchanged, above it the
head adds its own.

  A       vs B        n   rel rms IN  rel rms OUT      amp
  REF2    REF        48   0.0000e+00   0.0000e+00      nan

crest report: no activation statistics captured, skipped

[i8fo] folded GEMMs=39200 of 38400 served  displaced (A fragments already clobbered)=0  unmatched (served the shipped way)=0
[i8f]  fused prep calls=39200  deferred-rotation flushes=0  unmatched=0
[i8fo] REFUSED: 39200 of 38400 int8 calls served to ['INT8FO', 'INT8GFO'] were folded; the arm's wall clock is a mixture -- or, at zero, is the unfolded one and equal to INT8F by construction.
EXIT=1
