[ab] server http://172.17.0.1:8100 model /models/Qwen3.8-27B-exl3-4.0bpw
[transformers] Token indices sequence length is longer than the specified maximum sequence length for this model (317092 > 262144). Running this sequence through the model will result in indexing errors
  [0] code       prompt=8244 out=512 proposed=1169 accepted=346 rate=0.2960 tok/s= 74.74 length
  [1] code       prompt=8243 out=512 proposed=1127 accepted=352 rate=0.3123 tok/s= 77.13 length
  [2] code       prompt=8243 out=512 proposed=1057 accepted=362 rate=0.3425 tok/s= 80.78 length
  [3] code       prompt=8244 out=512 proposed=910 accepted=381 rate=0.4187 tok/s= 90.09 length
  [4] wiki       prompt=8244 out=512 proposed=994 accepted=370 rate=0.3722 tok/s= 84.57 length
  [5] wiki       prompt=8244 out=512 proposed=1162 accepted=346 rate=0.2978 tok/s= 69.46 length
  [6] wiki       prompt=8244 out=512 proposed=952 accepted=378 rate=0.3971 tok/s= 86.43 length
  [7] wiki       prompt=8244 out=512 proposed=1064 accepted=360 rate=0.3383 tok/s= 79.94 length
  [8] math       prompt=8244 out=512 proposed=847 accepted=397 rate=0.4687 tok/s= 93.94 length
  [9] math       prompt=8244 out=411 proposed=553 accepted=332 rate=0.6004 tok/s= 98.97 stop
  [10] math       prompt=8244 out=512 proposed=1120 accepted=351 rate=0.3134 tok/s= 77.05 length
  [11] math       prompt=8244 out=512 proposed=875 accepted=387 rate=0.4423 tok/s= 92.29 length
  [12] longctx    prompt=8244 out=512 proposed=1078 accepted=358 rate=0.3321 tok/s= 79.37 length
  [13] longctx    prompt=8244 out=512 proposed=1029 accepted=364 rate=0.3537 tok/s= 82.45 length
  [14] longctx    prompt=8244 out=512 proposed=1092 accepted=355 rate=0.3251 tok/s= 78.93 length
  [15] longctx    prompt=8244 out=512 proposed=1050 accepted=362 rate=0.3448 tok/s= 80.85 length
  [16] sci        prompt=8244 out=303 proposed=602 accepted=217 rate=0.3605 tok/s= 69.91 stop
  [17] sci        prompt=8244 out=471 proposed=791 accepted=357 rate=0.4513 tok/s= 91.07 stop
  [18] sci        prompt=8244 out=310 proposed=714 accepted=207 rate=0.2899 tok/s= 64.09 stop
  [19] sci        prompt=8245 out=366 proposed=693 accepted=267 rate=0.3853 tok/s= 76.98 stop
[ab] pooled accept rate 0.3628 over 20 prompts -> /cache/ab/DF_G.json

## pins / drafter (server log)
20:52:46 INFO    exl3 int8 leg pinned 400 linear(s) over 7 geometr(ies): 17408x5120xK4=shape111, 5120x10240xK4=shape111, 5120x1024xK4=shape111, 5120x12288xK4=shape111, 5120x17408xK4=shape111, 5120x6144xK4=shape111, 6144x5120xK4=shape111. Each is ONE frozen launch shape, so the reduction order does not follow the row count. Declined: 5120x248320xK6 (widest call site is 64 rows, below shape 111's TI
20:52:46 INFO    EXL3 int8 prefill leg ARMED at post-load: 400 of 401 bound linears pinned (0 newly here) over 7 geometr(ies) 17408x5120xK4=shape111, 5120x10240xK4=shape111, 5120x1024xK4=shape111, 5120x12288xK4=shape111, 5120x17408xK4=shape111, 5120x6144xK4=shape111, 6144x5120xK4=shape111; scratch 35.1 MiB in capture.io_buffers. Declined: 5120x248320xK6 (widest call site is 64 rows, below shape 11
20:52:49 INFO    EXL3 int8 prefill leg ARMED at post-drafter: 400 of 448 bound linears pinned (0 newly here) over 7 geometr(ies) 17408x5120xK4=shape111, 5120x10240xK4=shape111, 5120x1024xK4=shape111, 5120x12288xK4=shape111, 5120x17408xK4=shape111, 5120x6144xK4=shape111, 6144x5120xK4=shape111; scratch 0.0 MiB in capture.io_buffers. Declined: 17408x5120xK6 (shape 111 cannot run at K=6 cb=1: this ext
20:53:30 INFO    cudaGraphInstantiate MEASURED: decode=12.38 MiB/graph x16, dflash=6.25 MiB/graph x8 (24 exec(s), 248.0 MiB driver total) — persisting; retires the modelled per-graph constant.
20:53:30 INFO    graph_pool budget cache: persisted 1501560832 B (1432.0 MiB) + measured cudaGraphInstantiate decode=12.38 MiB/graph, dflash=6.25 MiB/graph → /cache/arbi-serve/budget-cache/dec74d8777b09bf8.json
20:53:32 INFO    MTP↔tkv wiring verified: fused MTP-verify attention on 16/16 layers
