[ab] server http://172.17.0.1:8100 model /models/Qwen3.8-27B-exl3-4.0bpw
[transformers] Token indices sequence length is longer than the specified maximum sequence length for this model (317092 > 262144). Running this sequence through the model will result in indexing errors
  [0] code       prompt=8244 out=512 proposed=710 accepted=372 rate=0.5239 tok/s= 80.07 length
  [1] code       prompt=8243 out=512 proposed=815 accepted=348 rate=0.4270 tok/s= 72.27 length
  [2] code       prompt=8243 out=512 proposed=855 accepted=341 rate=0.3988 tok/s= 69.55 length
  [3] code       prompt=8244 out=512 proposed=870 accepted=340 rate=0.3908 tok/s= 68.72 length
  [4] wiki       prompt=8244 out=512 proposed=760 accepted=362 rate=0.4763 tok/s= 76.17 length
  [5] wiki       prompt=8244 out=512 proposed=885 accepted=338 rate=0.3819 tok/s= 67.69 length
  [6] wiki       prompt=8244 out=512 proposed=815 accepted=348 rate=0.4270 tok/s= 72.15 length
  [7] wiki       prompt=8244 out=512 proposed=725 accepted=367 rate=0.5062 tok/s= 78.79 length
  [8] math       prompt=8244 out=512 proposed=735 accepted=367 rate=0.4993 tok/s= 78.04 length
  [9] math       prompt=8244 out=512 proposed=680 accepted=379 rate=0.5574 tok/s= 82.48 length
  [10] math       prompt=8244 out=512 proposed=830 accepted=346 rate=0.4169 tok/s= 71.02 length
  [11] math       prompt=8244 out=512 proposed=770 accepted=362 rate=0.4701 tok/s= 75.23 length
  [12] longctx    prompt=8244 out=512 proposed=900 accepted=334 rate=0.3711 tok/s= 66.71 length
  [13] longctx    prompt=8244 out=512 proposed=880 accepted=335 rate=0.3807 tok/s= 67.93 length
  [14] longctx    prompt=8244 out=512 proposed=805 accepted=352 rate=0.4373 tok/s= 72.72 length
  [15] longctx    prompt=8244 out=512 proposed=860 accepted=340 rate=0.3953 tok/s= 69.15 length
  [16] sci        prompt=8244 out=191 proposed=315 accepted=127 rate=0.4032 tok/s= 50.57 stop
  [17] sci        prompt=8244 out=512 proposed=670 accepted=378 rate=0.5642 tok/s= 83.46 length
  [18] sci        prompt=8244 out=208 proposed=315 accepted=145 rate=0.4603 tok/s= 55.00 stop
  [19] sci        prompt=8245 out=403 proposed=590 accepted=285 rate=0.4831 tok/s= 71.83 stop
[ab] pooled accept rate 0.4441 over 20 prompts -> /cache/ab/G.json

## pins
    runtime flags: prefix_cache=on; non-default: cudagraph_kv_pages_buckets=(), exl3_int8_gemm=True, exl3_int8_group_scale=True, serve_flat_cache_dir=/cache/flat-weights, serve_grammar_cache_dir=/cache/xgrammar, serve_calibration_cache=/cache/calibrations, serve_budget_cache_dir=/cache/arbi-serve/budget-cache, serve_calibration_dir=/cal, allow_backend_drift=True, serve_mtp_k_cache=/cache/arbi-serve/mtp-k, tkv_bits=4, tkv_calibration_file=/cal/qwen3.8-27b-exl3-4.0bpw_k4v4_mtp_20260816.json, TKV_CACHE_DIR=/cache/tkv, TKV_MTP_MAX_BLOCK_M=12
20:13:16 INFO    exl3 int8 leg pinned 408 linear(s) over 8 geometr(ies): 10240x5120xK4=shape111, 17408x5120xK4=shape111, 5120x10240xK4=shape111, 5120x1024xK4=shape111, 5120x12288xK4=shape111, 5120x17408xK4=shape111, 5120x6144xK4=shape111, 6144x5120xK4=shape111. Each is ONE frozen launch shape, so the reduction order does not follow the row count. Declined: 5120x248320xK6 (widest call site is 48 rows, below shape 111's TILE_M of 128)
20:13:16 INFO    EXL3 int8 prefill leg ARMED at post-load: 408 of 409 bound linears pinned (0 newly here) over 8 geometr(ies) 10240x5120xK4=shape111, 17408x5120xK4=shape111, 5120x10240xK4=shape111, 5120x1024xK4=shape111, 5120x12288xK4=shape111, 5120x17408xK4=shape111, 5120x6144xK4=shape111, 6144x5120xK4=shape111; scratch 39.4 MiB in capture.io_buffers. Declined: 5120x248320xK6 (widest call site is 48 rows, below shape 111's TILE_M of 128). One frozen launch shape per geometry, so the reduction order does not follow the row count. Whether it actually SERVED is the counter exl3_int8_prefill_gemm, not this line.
20:13:19 INFO    EXL3 int8 prefill leg ARMED at post-drafter: 408 of 409 bound linears pinned (0 newly here) over 8 geometr(ies) 10240x5120xK4=shape111, 17408x5120xK4=shape111, 5120x10240xK4=shape111, 5120x1024xK4=shape111, 5120x12288xK4=shape111, 5120x17408xK4=shape111, 5120x6144xK4=shape111, 6144x5120xK4=shape111; scratch 0.0 MiB in capture.io_buffers. Declined: 5120x248320xK6 (widest call site is 48 rows, below shape 111's TILE_M of 128). One frozen launch shape per geometry, so the reduction order does not follow the row count. Whether it actually SERVED is the counter exl3_int8_prefill_gemm, not this line.
20:13:26 INFO    tkv prefill prewarm: 2 kernel(s) compiled in 1.07s across 1 codec + 0 bypass geometry group(s); 0 variant(s) skipped; prefill kernel cache now holds 3 key(s). A post-ready cute compile after this is a warmup-coverage gap (jit_compile_serving), not an accepted residual.
20:13:27 INFO    prefix grouper: disabled at boot (ARBI_PREFIX_GROUPING=off) — live-flippable via POST /v1/admin/config_override
  driver.modules_loaded is 1756 MiB CAPTURE-driven of 1914 MiB (92%) — the capture phases' own bracketed deltas minus the 314 MiB already metered as driver.cudagraph_exec. Capture does not CREATE these module loads, it front-loads them: a first request would pay them otherwise. So the captured ladder's true cost is capture.cudagraphs + capture.io_buffers + driver.cudagraph_exec + this share — spread across two groups, which is why the CUDA-graph group under-states it.
