[ab] server http://172.17.0.1:8100 model /models/Qwen3.8-27B-exl3-4.0bpw
[transformers] Token indices sequence length is longer than the specified maximum sequence length for this model (317092 > 262144). Running this sequence through the model will result in indexing errors
  [0] code       prompt=8244 out=512 proposed=755 accepted=361 rate=0.4781 tok/s= 68.62 length
  [1] code       prompt=8243 out=512 proposed=940 accepted=323 rate=0.3436 tok/s= 59.16 length
  [2] code       prompt=8243 out=512 proposed=835 accepted=346 rate=0.4144 tok/s= 64.24 length
  [3] code       prompt=8244 out=512 proposed=750 accepted=361 rate=0.4813 tok/s= 69.00 length
  [4] wiki       prompt=8244 out=512 proposed=910 accepted=329 rate=0.3615 tok/s= 60.35 length
  [5] wiki       prompt=8244 out=512 proposed=730 accepted=370 rate=0.5068 tok/s= 70.27 length
  [6] wiki       prompt=8244 out=512 proposed=820 accepted=351 rate=0.4280 tok/s= 64.61 length
  [7] wiki       prompt=8244 out=512 proposed=845 accepted=345 rate=0.4083 tok/s= 63.43 length
  [8] math       prompt=8244 out=512 proposed=830 accepted=346 rate=0.4169 tok/s= 64.36 length
  [9] math       prompt=8244 out=512 proposed=745 accepted=364 rate=0.4886 tok/s= 69.07 length
  [10] math       prompt=8244 out=512 proposed=880 accepted=336 rate=0.3818 tok/s= 61.59 length
  [11] math       prompt=8244 out=420 proposed=610 accepted=298 rate=0.4885 tok/s= 64.62 stop
  [12] longctx    prompt=8244 out=512 proposed=940 accepted=324 rate=0.3447 tok/s= 58.85 length
  [13] longctx    prompt=8244 out=512 proposed=805 accepted=351 rate=0.4360 tok/s= 64.86 length
  [14] longctx    prompt=8244 out=512 proposed=835 accepted=346 rate=0.4144 tok/s= 63.98 length
  [15] longctx    prompt=8244 out=512 proposed=825 accepted=347 rate=0.4206 tok/s= 64.53 length
  [16] sci        prompt=8244 out=210 proposed=315 accepted=147 rate=0.4667 tok/s= 46.05 stop
  [17] sci        prompt=8244 out=446 proposed=650 accepted=317 rate=0.4877 tok/s= 65.07 stop
  [18] sci        prompt=8244 out=306 proposed=505 accepted=206 rate=0.4079 tok/s= 52.60 stop
  [19] sci        prompt=8245 out=512 proposed=665 accepted=379 rate=0.5699 tok/s= 74.31 length
[ab] pooled accept rate 0.4310 over 20 prompts -> /cache/ab/OFF.json

## pins
19:59:55 INFO    tkv prefill prewarm: 2 kernel(s) compiled in 1.09s across 1 codec + 0 bypass geometry group(s); 0 variant(s) skipped; prefill kernel cache now holds 3 key(s). A post-ready cute compile after this is a warmup-coverage gap (jit_compile_serving), not an accepted residual.
19:59:57 INFO    prefix grouper: disabled at boot (ARBI_PREFIX_GROUPING=off) — live-flippable via POST /v1/admin/config_override
  driver.modules_loaded is 1756 MiB CAPTURE-driven of 1914 MiB (92%) — the capture phases' own bracketed deltas minus the 314 MiB already metered as driver.cudagraph_exec. Capture does not CREATE these module loads, it front-loads them: a first request would pay them otherwise. So the captured ladder's true cost is capture.cudagraphs + capture.io_buffers + driver.cudagraph_exec + this share — spread across two groups, which is why the CUDA-graph group under-states it.
