[ab] server http://172.17.0.1:8100 model /models/Qwen3.8-27B-exl3-4.0bpw
[transformers] Token indices sequence length is longer than the specified maximum sequence length for this model (317092 > 262144). Running this sequence through the model will result in indexing errors
  [0] code       prompt=8244 out=512 proposed=1092 accepted=355 rate=0.3251 tok/s= 79.94 length
  [1] code       prompt=8243 out=512 proposed=1162 accepted=349 rate=0.3003 tok/s= 76.61 length
  [2] code       prompt=8243 out=512 proposed=1085 accepted=360 rate=0.3318 tok/s= 80.58 length
  [3] code       prompt=8244 out=512 proposed=1085 accepted=361 rate=0.3327 tok/s= 80.62 length
  [4] wiki       prompt=8244 out=512 proposed=966 accepted=375 rate=0.3882 tok/s= 87.84 length
  [5] wiki       prompt=8244 out=512 proposed=1050 accepted=363 rate=0.3457 tok/s= 82.60 length
  [6] wiki       prompt=8244 out=512 proposed=931 accepted=380 rate=0.4082 tok/s= 90.16 length
  [7] wiki       prompt=8244 out=512 proposed=1015 accepted=366 rate=0.3606 tok/s= 84.69 length
  [8] math       prompt=8244 out=512 proposed=917 accepted=382 rate=0.4166 tok/s= 91.23 length
  [9] math       prompt=8244 out=512 proposed=714 accepted=409 rate=0.5728 tok/s=108.53 length
  [10] math       prompt=8244 out=512 proposed=1015 accepted=366 rate=0.3606 tok/s= 84.57 length
  [11] math       prompt=8244 out=512 proposed=826 accepted=395 rate=0.4782 tok/s= 98.13 length
  [12] longctx    prompt=8244 out=512 proposed=1120 accepted=352 rate=0.3143 tok/s= 78.55 length
  [13] longctx    prompt=8244 out=512 proposed=1036 accepted=366 rate=0.3533 tok/s= 83.26 length
  [14] longctx    prompt=8244 out=512 proposed=1064 accepted=359 rate=0.3374 tok/s= 81.50 length
  [15] longctx    prompt=8244 out=512 proposed=973 accepted=372 rate=0.3823 tok/s= 87.13 length
  [16] sci        prompt=8244 out=140 proposed=294 accepted=98 rate=0.3333 tok/s= 48.44 stop
  [17] sci        prompt=8244 out=512 proposed=959 accepted=374 rate=0.3900 tok/s= 86.62 length
  [18] sci        prompt=8244 out=403 proposed=791 accepted=289 rate=0.3654 tok/s= 78.91 stop
  [19] sci        prompt=8245 out=307 proposed=490 accepted=237 rate=0.4837 tok/s= 82.17 stop
[ab] pooled accept rate 0.3717 over 20 prompts -> /cache/ab/DF_T.json

## pins / drafter (server log)
20:49:25 INFO    exl3 int8 leg pinned 400 linear(s) over 7 geometr(ies): 17408x5120xK4=shape27, 5120x10240xK4=shape27, 5120x1024xK4=shape27, 5120x12288xK4=shape27, 5120x17408xK4=shape27, 5120x6144xK4=shape27, 6144x5120xK4=shape27. Each is ONE frozen launch shape, so the reduction order does not follow the row count. Declined: 5120x248320xK6 (widest call site is 64 rows, below shape 27's TILE_M of 
20:49:25 INFO    EXL3 int8 prefill leg ARMED at post-load: 400 of 401 bound linears pinned (0 newly here) over 7 geometr(ies) 17408x5120xK4=shape27, 5120x10240xK4=shape27, 5120x1024xK4=shape27, 5120x12288xK4=shape27, 5120x17408xK4=shape27, 5120x6144xK4=shape27, 6144x5120xK4=shape27; scratch 34.0 MiB in capture.io_buffers. Declined: 5120x248320xK6 (widest call site is 64 rows, below shape 27's TILE
20:49:28 INFO    EXL3 int8 prefill leg ARMED at post-drafter: 400 of 448 bound linears pinned (0 newly here) over 7 geometr(ies) 17408x5120xK4=shape27, 5120x10240xK4=shape27, 5120x1024xK4=shape27, 5120x12288xK4=shape27, 5120x17408xK4=shape27, 5120x6144xK4=shape27, 6144x5120xK4=shape27; scratch 0.0 MiB in capture.io_buffers. Declined: 17408x5120xK6 (shape 27 cannot run at K=6 cb=1: this extension w
20:50:09 INFO    cudaGraphInstantiate MEASURED: decode=12.38 MiB/graph x16, dflash=6.25 MiB/graph x8 (24 exec(s), 248.0 MiB driver total) — persisting; retires the modelled per-graph constant.
20:50:09 INFO    graph_pool budget cache: persisted 1480589312 B (1412.0 MiB) + measured cudaGraphInstantiate decode=12.38 MiB/graph, dflash=6.25 MiB/graph → /cache/arbi-serve/budget-cache/dec74d8777b09bf8.json
20:50:11 INFO    MTP↔tkv wiring verified: fused MTP-verify attention on 16/16 layers
