[transformers] Token indices sequence length is longer than the specified maximum sequence length for this model (317092 > 262144). Running this sequence through the model will result in indexing errors
[i8] basket: 10 prompts x 16384 tok, domains ['code', 'longctx', 'math', 'sci', 'wiki']
compile-cache MISS: xgrammar — xgrammar compiled grammars (re)compiles this boot [/cache/xgrammar]
[TKV autotune-provenance] fp=57911e26f4e47f74 regime=cache-hit cells=8 sweep_src=2772549a0ef713db tile_tokens=4,8,16 num_splits=32,128,256 min_blocks_per_sm=0,3 resolved=0 held=0 inherited=0
/opt/venv/lib/python3.12/site-packages/torch/__init__.py:1172: FutureWarning: `torch.distributed.reduce_op` is deprecated, please use `torch.distributed.ReduceOp` instead
  return isinstance(obj, torch.Tensor)
JIT compile AFTER serving-ready [cpp_ext]: exl3_i8_gemm cached .so load — a live request paid this compile's latency. This is a boot-warmup coverage gap: extend warmup to pre-compile this kernel/specialization. Counter jit_compile_serving (must-not-fire) at GET /v1/admin/flag_truth.
[i8] prefix cache: server default resolved=True, per-request cache_enabled=False; any cache hit below is refused
[i8] warmup done
  [0] code           REF     16384tok  wall=  5.15s  finish=stop
  [0] code           REF2    16384tok  wall=  5.13s  finish=stop
  [0] code           STD     16384tok  wall=  5.41s  finish=stop
  [0] code           INT8    16384tok  wall=  4.13s  finish=stop
  [1] code           REF     16384tok  wall=  4.69s  finish=stop
  [1] code           REF2    16384tok  wall=  4.69s  finish=stop
  [1] code           STD     16384tok  wall=  4.97s  finish=stop
  [1] code           INT8    16384tok  wall=  3.66s  finish=stop
  [2] wiki           REF     16384tok  wall=  5.19s  finish=length
  [2] wiki           REF2    16384tok  wall=  5.19s  finish=length
  [2] wiki           STD     16384tok  wall=  5.45s  finish=length
  [2] wiki           INT8    16384tok  wall=  4.14s  finish=length
  [3] wiki           REF     16384tok  wall=  5.18s  finish=length
  [3] wiki           REF2    16384tok  wall=  5.19s  finish=length
  [3] wiki           STD     16384tok  wall=  5.49s  finish=length
  [3] wiki           INT8    16384tok  wall=  4.18s  finish=length
  [4] math           REF     16384tok  wall=  5.22s  finish=length
  [4] math           REF2    16384tok  wall=  5.18s  finish=length
  [4] math           STD     16384tok  wall=  5.46s  finish=length
  [4] math           INT8    16384tok  wall=  4.14s  finish=length
  [5] math           REF     16384tok  wall=  5.18s  finish=length
  [5] math           REF2    16384tok  wall=  5.19s  finish=length
  [5] math           STD     16384tok  wall=  5.45s  finish=length
  [5] math           INT8    16384tok  wall=  4.30s  finish=length
  [6] longctx        REF     16384tok  wall=  5.18s  finish=length
  [6] longctx        REF2    16384tok  wall=  5.19s  finish=length
  [6] longctx        STD     16384tok  wall=  5.45s  finish=length
  [6] longctx        INT8    16384tok  wall=  4.14s  finish=length
  [7] longctx        REF     16384tok  wall=  5.19s  finish=length
  [7] longctx        REF2    16384tok  wall=  5.20s  finish=length
  [7] longctx        STD     16384tok  wall=  5.46s  finish=length
  [7] longctx        INT8    16384tok  wall=  4.14s  finish=length
  [8] sci            REF     16384tok  wall=  5.19s  finish=length
  [8] sci            REF2    16384tok  wall=  5.19s  finish=length
  [8] sci            STD     16384tok  wall=  5.45s  finish=length
  [8] sci            INT8    16384tok  wall=  4.14s  finish=length
  [9] sci            REF     16384tok  wall=  5.30s  finish=length
  [9] sci            REF2    16384tok  wall=  5.21s  finish=length
  [9] sci            STD     16384tok  wall=  5.47s  finish=length
  [9] sci            INT8    16384tok  wall=  4.15s  finish=length

[i8] kernel census: int8 calls=32000 rows=65536000 reconstructs SKIPPED=32000 fallback (non-4bpw) calls=0
[i8] shapes served by the kernel (K,N)->calls: {(5120, 1024): 2560, (5120, 6144): 3840, (5120, 10240): 3840, (5120, 12288): 1280, (5120, 17408): 10240, (6144, 5120): 5120, (17408, 5120): 5120}

arm        prefill KL       median          max    decode KL  greedy agree
REF         0.000e+00    0.000e+00    0.000e+00    0.000e+00 145/145 = 100.0%
REF2        0.000e+00    0.000e+00    0.000e+00    0.000e+00 145/145 = 100.0%
STD         3.166e-04    3.795e-05    1.511e-03    7.916e-01 133/145 =  91.7%
INT8        9.963e-04    1.667e-04    5.568e-03    1.048e+00 130/145 =  89.7%
