[i8] basket: 6 prompts x 16384 tok, domains ['code', 'longctx', 'math', 'sci', 'wiki']
compile-cache MISS: tkv-autotune — tkv decode/prefill autotune tables (~min/shape) (re)compiles this boot [/cache/tkv]
compile-cache MISS: xgrammar — xgrammar compiled grammars (re)compiles this boot [/cache/xgrammar]
compile-cache MISS: megacache-fx — torch Mega-Cache bundle (Inductor FX + AOTAutograd backend) (re)compiles this boot [/cache/arbi-serve/compile-cache/588802f088df2611.bin]
default-pool residency snapshot returned no live ranges — the unpooled-weight fold cannot see what it would move
VRAM at profile: this card's residency could NOT be split between this process and any other — NVML's per-process walk did not identify us — so the whole 14.26 GiB resident on device 0 is being booked as driver.cuda_context. Any co-tenant's bytes are inside that number and KV is sized that much smaller for this boot.
[TKV] Cold-boot kernel autotune starting: this is a real, expected one-time cost, not a hang — timing candidate decode-kernel configs for the shapes this deployment needs. Per-cell progress prints below as each one resolves; subsequent boots on this machine with an unchanged config reuse the cached table and skip this. See README.md's 'First boot on a new machine' section for real measured timings and how to skip this cost entirely with a pre-baked image (TKV_BAKE=1).
[TKV autotune] start: kv_cache shape=(1024, 256, 2048) data_ptr=6a0000000 in_capture=False (total_pages=1024) H_kv=4 num_sms=128 sw=0 buckets=(1024, 4096, 8192, 16384) batches=(1, 2) (capped at max_batch=1) reusable_cells=0 inherited=0 variants=['splitk'] tile_tokens=[4, 8, 16] min_blocks=[0, 3]
[TKV autotune] bucket=  1024 batch=  1: HELD splitk splits=32 tt=4 mb=0 (0.053ms) margin=0.051% sigma=1.101% z=0.11 n=19 inherited=False
[TKV autotune] bucket=  1024 batch=  2: HELD splitk splits=32 tt=4 mb=0 (0.053ms) margin=0.110% sigma=0.410% z=0.66 n=8 inherited=False
[TKV autotune] bucket=  4096 batch=  1: HELD splitk splits=128 tt=4 mb=0 (0.054ms) margin=0.209% sigma=0.504% z=1.12 n=7 inherited=False
[TKV autotune] bucket=  4096 batch=  2: HELD splitk splits=64 tt=4 mb=0 (0.053ms) margin=0.124% sigma=0.888% z=0.33 n=13 inherited=False
[TKV autotune] bucket=  8192 batch=  1: HELD splitk splits=128 tt=4 mb=0 (0.054ms) margin=0.250% sigma=1.780% z=0.37 n=18 inherited=False
[TKV autotune] bucket=  8192 batch=  2: HELD splitk splits=128 tt=4 mb=0 (0.054ms) margin=0.043% sigma=1.715% z=0.07 n=15 inherited=False
[TKV autotune] bucket= 16384 batch=  1: HELD splitk splits=256 tt=4 mb=0 (0.054ms) margin=0.578% sigma=0.547% z=2.46 n=2 inherited=False
[TKV autotune] bucket= 16384 batch=  2: HELD splitk splits=32 tt=4 mb=0 (0.153ms) margin=0.006% sigma=0.092% z=0.15 n=2 inherited=False
[TKV autotune] complete: 8 (batch, bucket) cells swept in 206.73s batch_invariant=0 tile_width_invariant=0
[TKV autotune-provenance] fp=f4e763cb2f31eb74 regime=swept cells=8 sweep_src=b5332b8dfc1c8389 tile_tokens=4 num_splits=32,64,128,256 min_blocks_per_sm=0 resolved=0 held=8 inherited=0
activation profile: the decode probe's FIRST call cost 8240380 allocator events / 1036 MiB / 206804.9 ms against 12102 / 1 MiB / 47.9 ms warm — a first-call kernel search inside the probe, not a step. The reported peak AND device time are the warm ones (2 passes, the last is what is reported); the cold numbers are logged so the difference is visible.
driver.modules_loaded baseline: SEEDED at 157286400 B (0.146 GiB) → /root/.cache/arbi-serve/budget-cache/84169d322e956268.modules.json. This boot ran UNGATED — it booked its own bracketed growth, so an unregistered pool or a raw cudaMalloc would be inside that number rather than on driver.residual. Every later boot at this configuration is held to it. Expected ONCE per configuration; if it repeats, the budget cache is not persisting (point ARBI_SERVE_BUDGET_CACHE_DIR at a durable volume) and the guard is inert.
  return isinstance(obj, torch.Tensor)
  if not _in_default_pool(obj.data_ptr()):
[i8] trellis->path map: 401 linears
JIT compile AFTER serving-ready [cpp_ext]: exl3_i8_gemm_k4_cb2 cached .so load — a live request paid this compile's latency. This is a boot-warmup coverage gap: extend warmup to pre-compile this kernel/specialization. Counter jit_compile_serving (must-not-fire) at GET /v1/admin/flag_truth.
[i8] prefix cache: server default resolved=True, per-request cache_enabled=False; any cache hit below is refused
JIT compile AFTER serving-ready [cpp_ext]: arbi_serve_exl3_a_prep_v3 fresh ninja/nvcc build — a live request paid this compile's latency. This is a boot-warmup coverage gap: extend warmup to pre-compile this kernel/specialization. Counter jit_compile_serving (must-not-fire) at GET /v1/admin/flag_truth.
[i8] warmup done
  [0] code           REF      16384tok  wall=  5.03s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [0] code           REF2     16384tok  wall=  5.01s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [0] code           INT8FO   16384tok  wall=  3.19s  finish=stop  leg=0F/3200S acc=3200x16/0x32
  [0] code           INT8GFO  16384tok  wall=  3.68s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [1] code           REF      16384tok  wall=  4.82s  finish=stop  leg=3200F/0S acc=3200x16/0x32
  [1] code           REF2     16384tok  wall=  4.83s  finish=stop  leg=3200F/0S acc=3200x16/0x32
  [1] code           INT8FO   16384tok  wall=  3.28s  finish=stop  leg=0F/3200S acc=3200x16/0x32
  [1] code           INT8GFO  16384tok  wall=  3.48s  finish=stop  leg=0F/3200S acc=3200x16/0x32
  [2] wiki           REF      16384tok  wall=  5.07s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [2] wiki           REF2     16384tok  wall=  5.06s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [2] wiki           INT8FO   16384tok  wall=  3.50s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [2] wiki           INT8GFO  16384tok  wall=  3.72s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [3] math           REF      16384tok  wall=  5.08s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [3] math           REF2     16384tok  wall=  5.08s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [3] math           INT8FO   16384tok  wall=  3.51s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [3] math           INT8GFO  16384tok  wall=  3.71s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [4] longctx        REF      16384tok  wall=  5.09s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [4] longctx        REF2     16384tok  wall=  5.10s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [4] longctx        INT8FO   16384tok  wall=  3.52s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [4] longctx        INT8GFO  16384tok  wall=  3.73s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [5] sci            REF      16384tok  wall=  5.12s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [5] sci            REF2     16384tok  wall=  5.11s  finish=length  leg=3200F/0S acc=3200x16/0x32
  [5] sci            INT8FO   16384tok  wall=  3.52s  finish=length  leg=0F/3200S acc=3200x16/0x32
  [5] sci            INT8GFO  16384tok  wall=  3.72s  finish=length  leg=0F/3200S acc=3200x16/0x32

[i8] kernel census: int8 calls=38400 rows=78643200 reconstructs SKIPPED=38400 fallback (non-4bpw) calls=0
[i8] shapes served by the kernel (K,N)->calls: {(5120, 1024): 3072, (5120, 6144): 4608, (5120, 10240): 4608, (5120, 12288): 1536, (5120, 17408): 12288, (6144, 5120): 6144, (17408, 5120): 6144}
[i8] excluded projections: none  excluded calls=0

  linear class                                           served   excl  int8 vs fp32 legB vs fp32 int8 vs legB
  model.layers.0.linear_attn.in_proj_qkv                     98      0             -            -            -
  model.layers.0.linear_attn.in_proj_z                       98      0             -            -            -
  model.layers.9.linear_attn.out_proj                        98      0             -            -            -
  model.layers.9.mlp.down_proj                               98      0             -            -            -
  model.layers.9.mlp.gate_proj                               98      0             -            -            -
  model.layers.9.mlp.up_proj                                 98      0             -            -            -
[i8] per-arm census (int8 kernel calls): {'REF': 0, 'REF2': 0, 'INT8FO': 19200, 'INT8GFO': 19200}

ARM RECEIPTS -- what each arm actually executed, from the engine's
own per-call census.  fused/standalone is the reconstruct variant;
acc16/acc32 is the cuBLAS compute type leg B asked hgemm for.
  arm       legB calls          rows     fused  standalone     acc16     acc32   fp32pin
  REF            19200      39321600     19200           0     19200         0         0
  REF2           19200      39321600     19200           0     19200         0         0
  INT8FO         19200      39321600         0       19200     19200         0         0
  INT8GFO        19200      39321600         0       19200     19200         0         0
  all arms match their expected leg/accumulator signature
Traceback (most recent call last):
(KL post-processing refused: two prompts stopped at EOS, so the walk lengths differ; the --ignore-eos rerun is receipt_e2e_single_bank_v2.txt)
