arm=branch tag=b4 ready_s=71 at=2026-09-13T22:23:41Z
== 22:23:41 ==
source_digest: 9e68117a78a7480d
server_info:   {'max_context': 186624, 'max_batch': 8}
kv metrics:
arbi_serve_kv_pool_pages_total{backend="paged_kv:tkv-k4v4",kind="paged_kv"} 730.0
arbi_serve_kv_pool_pages_total{backend="",kind="gdn"} 730.0
--- boot lines ---
22:22:35 INFO    max_context=auto → resolved to model max_position_embeddings=262144 (the model's trained context ceiling).
22:22:53 INFO    serving torch-caching overhang: 0.65 GiB (MEASURED at this config (budget cache))
22:23:26 INFO    dflash pre-sweep: captured 8/8 draft graphs (pool sealed; graph pool now 382 MiB reserved)
22:23:26 INFO    dflash realize pre-sweep: captured 24/24 realization graphs (pool sealed; graph pool now 382 MiB reserved)
22:23:31 INFO    serving grow floor: 316 MiB (base 4 [MEASURED] + arena-regrow 43 [MEASURED: admission-bound 504 MiB against the 604 MiB the pool already holds (serving mark UNMEASURED; the residency is the boot probes')] + max[spec-reserve 263 (=add(device_slots) of verify-tail 61 + verify-logits 62 [64 rows, measured 62, bound 61], dflash-draft 141), tkv-bypass-scratch 0, tkv-prefill-staging 9] + logprobs-tile 0 + media-encode 0 [stochastic_on=True rejection_flag=True measured=61 analytic=409] + token2wav-vocoder 0 + driver-growth 6 [MEASURED_CACHED]) vs gmu_floor 0 MiB vs residency-gate unpaid-graph-exec 0 MiB (predicted 234 MiB for this configuration's capture ladder, 234 MiB already metered onto the card as driver.cudagraph_exec -> 0 MiB still unpaid; 0 member record(s)) vs serving-consumption 0 MiB [MEASURED] -> 316 MiB [tts-codec-peak 0 + tts-kv-growth 0 (5000 ticks)]
22:23:31 INFO    serving grow floor: 1021 MiB (base 4 [MEASURED] + arena-regrow 748 [MEASURED: admission-bound 504 MiB against the 604 MiB the pool already holds (serving mark UNMEASURED; the residency is the boot probes')] + max[spec-reserve 263 (=add(device_slots) of verify-tail 61 + verify-logits 62 [64 rows, measured 62, bound 61], dflash-draft 141), tkv-bypass-scratch 0, tkv-prefill-staging 9] + logprobs-tile 0 + media-encode 0 [stochastic_on=True rejection_flag=True measured=61 analytic=409] + token2wav-vocoder 0 + driver-growth 6 [MEASURED_CACHED]) vs gmu_floor 0 MiB vs residency-gate unpaid-graph-exec 0 MiB (predicted 234 MiB for this configuration's capture ladder, 234 MiB already metered onto the card as driver.cudagraph_exec -> 0 MiB still unpaid; 0 member record(s)) vs serving-consumption 752 MiB [MEASURED] -> 1021 MiB [tts-codec-peak 0 + tts-kv-growth 0 (5000 ticks)]
22:23:33 INFO    serving grow floor: 387 MiB (base 4 [MEASURED] + arena-regrow 114 [MEASURED: admission-bound 504 MiB against the 604 MiB the pool already holds (serving mark UNMEASURED; the residency is the boot probes')] + max[spec-reserve 263 (=add(device_slots) of verify-tail 61 + verify-logits 62 [64 rows, measured 62, bound 61], dflash-draft 141), tkv-bypass-scratch 0, tkv-prefill-staging 9] + logprobs-tile 0 + media-encode 0 [stochastic_on=True rejection_flag=True measured=61 analytic=409] + token2wav-vocoder 0 + driver-growth 6 [MEASURED_CACHED]) vs gmu_floor 0 MiB vs residency-gate unpaid-graph-exec 0 MiB (predicted 234 MiB for this configuration's capture ladder, 234 MiB already metered onto the card as driver.cudagraph_exec -> 0 MiB still unpaid; 0 member record(s)) vs serving-consumption 118 MiB [MEASURED] -> 387 MiB [tts-codec-peak 0 + tts-kv-growth 0 (5000 ticks)]
22:23:34 INFO    serving grow floor: 387 MiB (base 4 [MEASURED] + arena-regrow 114 [MEASURED: admission-bound 504 MiB against the 604 MiB the pool already holds (a drive's mark: the boot probes had 798 MiB, 194 MiB of it residue the KV grow now reads free)] + max[spec-reserve 263 (=add(device_slots) of verify-tail 61 + verify-logits 62 [64 rows, measured 62, bound 61], dflash-draft 141), tkv-bypass-scratch 0, tkv-prefill-staging 9] + logprobs-tile 0 + media-encode 0 [stochastic_on=True rejection_flag=True measured=61 analytic=409] + token2wav-vocoder 0 + driver-growth 6 [MEASURED_CACHED]) vs gmu_floor 0 MiB vs residency-gate unpaid-graph-exec 0 MiB (predicted 234 MiB for this configuration's capture ladder, 234 MiB already metered onto the card as driver.cudagraph_exec -> 0 MiB still unpaid; 0 member record(s)) vs serving-consumption 118 MiB [MEASURED] -> 387 MiB [tts-codec-peak 0 + tts-kv-growth 0 (5000 ticks)]
22:23:34 WARNING max_context=auto: served context NARROWED to 183808 tokens (from the 262144-token sizing bound). This is the REAL served-context statement — the KV pool is realized: 730 servable pages × 256 tokens/page, less the 11-page capture-sentinel prefix and a 0-page allocator margin. The earlier sizing bound was a forecast and came out high. Pin --max-context explicitly to refuse loud instead.
22:23:34 INFO    graph_pool budget cache: persisted 1509949440 B (1440.0 MiB) + measured cudaGraphInstantiate decode=11.88 MiB/graph, dflash=1.38 MiB/graph → /cache/arbi-serve/budget-cache/04f389ba71121f05.json
22:23:34 INFO    Boot: KV cache ready at 3.03 GiB (13% of 23.52 GiB) — 730 pages × 256 tok/page = 186880 servable tokens, free 0.39 GiB (2% of 23.52 GiB)
22:23:34 INFO    max_context=auto: widened served context to 186624 tokens to use the full realized KV pool (730 servable pages).
22:23:35 INFO    Boot: KV cache ready at 3.03 GiB (13% of 23.52 GiB) — 730 pages × 256 tok/page = 186880 servable tokens, free 0.39 GiB (2% of 23.52 GiB)
