[ab] server http://172.17.0.1:8100 model /models/Qwen3.8-27B-exl3-4.0bpw
[transformers] Token indices sequence length is longer than the specified maximum sequence length for this model (317092 > 262144). Running this sequence through the model will result in indexing errors
  [0] code       prompt=8244 out=512 proposed=735 accepted=369 rate=0.5020 tok/s= 78.99 length
  [1] code       prompt=8243 out=512 proposed=895 accepted=332 rate=0.3709 tok/s= 68.09 length
  [2] code       prompt=8243 out=512 proposed=840 accepted=344 rate=0.4095 tok/s= 71.48 length
  [3] code       prompt=8244 out=512 proposed=805 accepted=351 rate=0.4360 tok/s= 73.91 length
  [4] wiki       prompt=8244 out=512 proposed=920 accepted=328 rate=0.3565 tok/s= 66.36 length
  [5] wiki       prompt=8244 out=512 proposed=760 accepted=359 rate=0.4724 tok/s= 77.05 length
  [6] wiki       prompt=8244 out=512 proposed=810 accepted=350 rate=0.4321 tok/s= 73.50 length
  [7] wiki       prompt=8244 out=512 proposed=875 accepted=336 rate=0.3840 tok/s= 69.21 length
  [8] math       prompt=8244 out=512 proposed=680 accepted=380 rate=0.5588 tok/s= 84.24 length
  [9] math       prompt=8244 out=512 proposed=685 accepted=378 rate=0.5518 tok/s= 83.62 length
  [10] math       prompt=8244 out=512 proposed=815 accepted=348 rate=0.4270 tok/s= 73.26 length
  [11] math       prompt=8244 out=512 proposed=735 accepted=364 rate=0.4952 tok/s= 79.33 length
  [12] longctx    prompt=8244 out=512 proposed=865 accepted=339 rate=0.3919 tok/s= 69.89 length
  [13] longctx    prompt=8244 out=512 proposed=855 accepted=340 rate=0.3977 tok/s= 70.57 length
  [14] longctx    prompt=8244 out=512 proposed=870 accepted=341 rate=0.3920 tok/s= 69.65 length
  [15] longctx    prompt=8244 out=512 proposed=910 accepted=331 rate=0.3637 tok/s= 67.23 length
  [16] sci        prompt=8244 out=378 proposed=480 accepted=283 rate=0.5896 tok/s= 79.34 stop
  [17] sci        prompt=8244 out=512 proposed=565 accepted=402 rate=0.7115 tok/s= 96.05 length
  [18] sci        prompt=8244 out=239 proposed=355 accepted=168 rate=0.4732 tok/s= 60.70 stop
  [19] sci        prompt=8245 out=256 proposed=335 accepted=192 rate=0.5731 tok/s= 67.36 stop
[ab] pooled accept rate 0.4486 over 20 prompts -> /cache/ab/T.json

## pins
20:09:38 INFO    exl3 int8 leg pinned 408 linear(s) over 8 geometr(ies): 10240x5120xK4=shape27, 17408x5120xK4=shape27, 5120x10240xK4=shape27, 5120x1024xK4=shape27, 5120x12288xK4=shape27, 5120x17408xK4=shape27, 5120x6144xK4=shape27, 6144x5120xK4=shape27. Each is ONE frozen launch shape, so the reduction order does not follow the row count. Declined: 5120x248320xK6 (widest call site is 48 rows, below shape 27's TILE_M of 128)
20:09:38 INFO    EXL3 int8 prefill leg ARMED at post-load: 408 of 409 bound linears pinned (0 newly here) over 8 geometr(ies) 10240x5120xK4=shape27, 17408x5120xK4=shape27, 5120x10240xK4=shape27, 5120x1024xK4=shape27, 5120x12288xK4=shape27, 5120x17408xK4=shape27, 5120x6144xK4=shape27, 6144x5120xK4=shape27; scratch 38.3 MiB in capture.io_buffers. Declined: 5120x248320xK6 (widest call site is 48 rows, below shape 27's TILE_M of 128). One frozen launch shape per geometry, so the reduction order does not follow the row count. Whether it actually SERVED is the counter exl3_int8_prefill_gemm, not this line.
20:09:41 INFO    EXL3 int8 prefill leg ARMED at post-drafter: 408 of 409 bound linears pinned (0 newly here) over 8 geometr(ies) 10240x5120xK4=shape27, 17408x5120xK4=shape27, 5120x10240xK4=shape27, 5120x1024xK4=shape27, 5120x12288xK4=shape27, 5120x17408xK4=shape27, 5120x6144xK4=shape27, 6144x5120xK4=shape27; scratch 0.0 MiB in capture.io_buffers. Declined: 5120x248320xK6 (widest call site is 48 rows, below shape 27's TILE_M of 128). One frozen launch shape per geometry, so the reduction order does not follow the row count. Whether it actually SERVED is the counter exl3_int8_prefill_gemm, not this line.
20:09:49 INFO    tkv prefill prewarm: 2 kernel(s) compiled in 1.10s across 1 codec + 0 bypass geometry group(s); 0 variant(s) skipped; prefill kernel cache now holds 3 key(s). A post-ready cute compile after this is a warmup-coverage gap (jit_compile_serving), not an accepted residual.
20:09:51 INFO    prefix grouper: disabled at boot (ARBI_PREFIX_GROUPING=off) — live-flippable via POST /v1/admin/config_override
  driver.modules_loaded is 1756 MiB CAPTURE-driven of 1914 MiB (92%) — the capture phases' own bracketed deltas minus the 314 MiB already metered as driver.cudagraph_exec. Capture does not CREATE these module loads, it front-loads them: a first request would pay them otherwise. So the captured ladder's true cost is capture.cudagraphs + capture.io_buffers + driver.cudagraph_exec + this share — spread across two groups, which is why the CUDA-graph group under-states it.
