[i8] basket: 6 prompts x 16384 tok, domains ['code', 'longctx', 'math', 'sci', 'wiki']
compile-cache MISS: xgrammar — xgrammar compiled grammars (re)compiles this boot [/cache/xgrammar]
VRAM at profile: this card's residency could NOT be split between this process and any other — NVML's per-process walk did not identify us — so the whole 14.26 GiB resident on device 0 is being booked as driver.cuda_context. Any co-tenant's bytes are inside that number and KV is sized that much smaller for this boot.
[TKV autotune-provenance] fp=f4e763cb2f31eb74 regime=cache-hit cells=8 sweep_src=b5332b8dfc1c8389 tile_tokens=4 num_splits=32,64,128,256 min_blocks_per_sm=0 resolved=0 held=0 inherited=0
  return isinstance(obj, torch.Tensor)
[i8] trellis->path map: 401 linears
JIT compile AFTER serving-ready [cpp_ext]: exl3_i8_gemm_k4_cb2 cached .so load — a live request paid this compile's latency. This is a boot-warmup coverage gap: extend warmup to pre-compile this kernel/specialization. Counter jit_compile_serving (must-not-fire) at GET /v1/admin/flag_truth.
[i8] prefix cache: server default resolved=True, per-request cache_enabled=False; any cache hit below is refused
JIT compile AFTER serving-ready [cpp_ext]: arbi_serve_exl3_a_prep_v3 cached .so load — a live request paid this compile's latency. This is a boot-warmup coverage gap: extend warmup to pre-compile this kernel/specialization. Counter jit_compile_serving (must-not-fire) at GET /v1/admin/flag_truth.
activation admission RE-ARMED (JIT compile after serving-ready): the step budget FELL to 6381 MiB, -2 MiB on the 6383 MiB it was enforcing and 6383 MiB the boot reading armed. The widest slate admission will build is narrower from here — narrower prefill chunks and deferred rows instead of a forward that OOMs after the layout freeze.
[i8] warmup done
  [0] code           REF      16384tok  wall=  5.03s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [0] code           REF2     16384tok  wall=  5.01s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [0] code           INT8FO   16384tok  wall=  3.47s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [0] code           INT8GFO  16384tok  wall=  3.66s  finish=length  leg=0F/3200S acc=3200x16/0x32
driver serving-growth OVER RESERVE: the driver holds 2 MiB more than it did when the boot brackets closed, against the 0 MiB the serving floor's transient.serving_step.driver_growth term held for it. Those bytes are in no pool and no allocator counter -- the live card reaches them as driver.modules_serving and the boot ledger only as driver.residual -- and the KV layout is frozen, so the overage comes out of the free VRAM the floor is holding for one in-flight step (the verify tail, the DFlash context-assemble) and those reserves are the ones that fail first. The reading is on record for this configuration, so the next boot holds it; this process serves the rest of its life short by the overage.
activation admission RE-ARMED (driver residency over reserve): the step budget FELL to 6313 MiB, -68 MiB on the 6381 MiB it was enforcing and 6383 MiB the boot reading armed. The widest slate admission will build is narrower from here — narrower prefill chunks and deferred rows instead of a forward that OOMs after the layout freeze.
  [1] code           REF      16384tok  wall=  5.06s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [1] code           REF2     16384tok  wall=  5.08s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [1] code           INT8FO   16384tok  wall=  3.51s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [1] code           INT8GFO  16384tok  wall=  3.71s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [2] wiki           REF      16384tok  wall=  5.07s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [2] wiki           REF2     16384tok  wall=  5.09s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [2] wiki           INT8FO   16384tok  wall=  3.50s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [2] wiki           INT8GFO  16384tok  wall=  3.71s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [3] math           REF      16384tok  wall=  5.09s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [3] math           REF2     16384tok  wall=  5.10s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [3] math           INT8FO   16384tok  wall=  3.51s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [3] math           INT8GFO  16384tok  wall=  3.70s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [4] longctx        REF      16384tok  wall=  5.11s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [4] longctx        REF2     16384tok  wall=  5.10s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [4] longctx        INT8FO   16384tok  wall=  3.52s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [4] longctx        INT8GFO  16384tok  wall=  3.72s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [5] sci            REF      16384tok  wall=  5.10s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [5] sci            REF2     16384tok  wall=  5.11s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [5] sci            INT8FO   16384tok  wall=  3.51s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [5] sci            INT8GFO  16384tok  wall=  3.76s  finish=length  leg=0F/3200S acc=3200x16/0x32

[i8] kernel census: int8 calls=38400 rows=78643200 reconstructs SKIPPED=38400 fallback (non-4bpw) calls=0
[i8] shapes served by the kernel (K,N)->calls: {(5120, 1024): 3072, (5120, 6144): 4608, (5120, 10240): 4608, (5120, 12288): 1536, (5120, 17408): 12288, (6144, 5120): 6144, (17408, 5120): 6144}
...
  model.layers.9.linear_attn.in_proj_z                       98      0             -            -            -
  model.layers.9.linear_attn.out_proj                        98      0             -            -            -
  model.layers.9.mlp.down_proj                               98      0             -            -            -
  model.layers.9.mlp.gate_proj                               98      0             -            -            -
  model.layers.9.mlp.up_proj                                 98      0             -            -            -
[i8] per-arm census (int8 kernel calls): {'REF': 0, 'REF2': 0, 'INT8FO': 19200, 'INT8GFO': 19200}

ARM RECEIPTS -- what each arm actually executed, from the engine's
own per-call census.  fused/standalone is the reconstruct variant;
acc16/acc32 is the cuBLAS compute type leg B asked hgemm for.
  arm       legB calls          rows     fused  standalone     acc16     acc32   fp32pin
  REF            19200      39321600     19200           0     19200         0         0
  REF2           19200      39321600     19200           0     19200         0         0
  INT8FO         19200      39321600         0       19200     19200         0         0
  INT8GFO        19200      39321600         0       19200     19200         0         0
  all arms match their expected leg/accumulator signature
  prefill positions aligned: [8, 8, 8, 8, 8, 8] per prompt, identical across all 4 arms
[i8] null control: REF2 vs REF max KL is exactly 0

THE TAIL over PREFILL positions only -- one per prefill chunk, so
every arm scores the SAME context. Decode positions are reported
separately and are NOT a numerics measure: once a token flips, the
arms are reading different contexts. The control proves it -- pooled
over decode, STD (which contains no int8 at all) has the same max KL
as INT8, so a pooled tail cannot attribute anything to the kernel.

arm            n        mean         p50         p95         p99         MAX     argmax flip
REF           48   0.000e+00   0.000e+00   0.000e+00   0.000e+00   0.000e+00 0/48 =   0.00%
REF2          48   0.000e+00   0.000e+00   0.000e+00   0.000e+00   0.000e+00 0/48 =   0.00%
INT8FO        48   1.224e-01   8.966e-04   3.133e-01   4.259e+00   4.259e+00 2/48 =   4.17%
INT8GFO       48   1.023e-01   7.028e-04   2.343e-01   3.841e+00   3.841e+00 1/48 =   2.08%

Decode positions (divergence-contaminated, shown for completeness):
arm            n        mean         MAX    greedy agree
REF           48   0.000e+00   0.000e+00 48/48 =  100.0%
REF2          48   0.000e+00   0.000e+00 48/48 =  100.0%
INT8FO        48   4.465e+00   3.254e+01 40/48 =   83.3%
INT8GFO       48   1.783e-03   1.695e-02 48/48 =  100.0%

SAMPLER-RELEVANT TAIL @ canonical (T=1.0, top_p=0.95, top_k=20), prefill positions only.
TV is the headline: under a drafter that tracks the target, speculative
decoding's expected accept rate is sum_t min(p_ref, p_arm), so TV IS the
expected accept-rate loss.
arm            n  setOverlap    TV mean     TV p95     TV max  truncKL p95  bndryChurn  drawDiff
REF           48     100.00%  0.000e+00  0.000e+00  0.000e+00    0.000e+00 0/ 48     0.00%
REF2          48     100.00%  0.000e+00  0.000e+00  0.000e+00    0.000e+00 0/ 48     0.00%
INT8FO        48      98.29%  4.575e-02  3.620e-01  7.058e-01          inf 11/ 48     4.43%
INT8GFO       48      98.18%  4.945e-02  2.842e-01  8.840e-01          inf 10/ 48     7.81%

SAMPLER-RELEVANT TAIL @ owner's k~40 (T=1.0, top_p=1.0, top_k=40), prefill positions only.
TV is the headline: under a drafter that tracks the target, speculative
decoding's expected accept rate is sum_t min(p_ref, p_arm), so TV IS the
expected accept-rate loss.
arm            n  setOverlap    TV mean     TV p95     TV max  truncKL p95  bndryChurn  drawDiff
REF           48     100.00%  0.000e+00  0.000e+00  0.000e+00    0.000e+00 0/ 48     0.00%
REF2          48     100.00%  0.000e+00  0.000e+00  0.000e+00    0.000e+00 0/ 48     0.00%
INT8FO        48      93.53%  4.584e-02  3.459e-01  6.971e-01          inf 41/ 48     6.25%
INT8GFO       48      95.49%  4.856e-02  2.723e-01  8.649e-01          inf 33/ 48     9.64%

FULL-VOCAB numbers below, kept and labelled: they are the right
instrument for GREEDY or beam decode, where every rank matters.

arm         first-prefill-position KL       median          max
REF                         0.000e+00    0.000e+00    0.000e+00
REF2                        0.000e+00    0.000e+00    0.000e+00
INT8FO                      5.704e-04    3.044e-05    3.246e-03
INT8GFO                     2.510e-04    4.335e-05    1.298e-03

PAIRED TV @ canonical (T=1.0, top_p=0.95, top_k=20), prefill positions.
Each row is A scored against B -- B is that ROW's reference, not the
table's.  TV is the expected speculative accept-rate loss, so the
column is read directly in percentage points.  The interval is a
cluster bootstrap over PROMPTS; the IQR is the raw spread over
positions with no floor applied to it.

  A       vs B        n  mean TV pp           95% CI pp    p25 pp    p50 pp    p75 pp    p95 pp    max pp
  REF2    REF        48      0.0000 [  0.0000,  0.0000]    0.0000    0.0000    0.0000    0.0000    0.0000

lm_head CUT -- the divergence entering the head vs leaving it,
prefill positions, relative rms.  'in' is the residual stream after
24 layers; 'out' is the centred logits.  amp = out/in: at 1.0 the
head passes the body's divergence through unchanged, above it the
head adds its own.

  A       vs B        n   rel rms IN  rel rms OUT      amp
  REF2    REF        48   0.0000e+00   0.0000e+00      nan

crest report: no activation statistics captured, skipped

[i8fo] folded GEMMs=0 of 38400 served  displaced (A fragments already clobbered)=0  unmatched (served the shipped way)=0
[i8f]  fused prep calls=39200  deferred-rotation flushes=0  unmatched=0
[i8fo] REFUSED: 0 of 38400 int8 calls served to ['INT8FO', 'INT8GFO'] were folded; the arm's wall clock is a mixture -- or, at zero, is the unfolded one and equal to INT8F by construction.
EXIT=1
