arm=branch tag=b2 ready_s=223 at=2026-09-13T21:48:18Z
== 21:48:18 ==
source_digest: 9e68117a78a7480d
server_info:   {'max_context': 186624, 'max_batch': 8}
kv metrics:
arbi_serve_kv_pool_pages_total{backend="paged_kv:tkv-k4v4",kind="paged_kv"} 730.0
arbi_serve_kv_pool_pages_total{backend="",kind="gdn"} 730.0
--- boot lines ---
21:44:50 INFO    max_context=auto → resolved to model max_position_embeddings=262144 (the model's trained context ceiling).
21:45:28 INFO    serving torch-caching overhang: 0.65 GiB (MEASURED at this config (budget cache))
21:45:31 WARNING max_context=auto: PROVISIONAL SIZING bound — the model ceiling (262144 tokens) exceeds what the reserved KV address space can hold (VA serving ceiling 1000 pages), so the max_context-scaled buffers are sized to 255744 tokens and the boot CONTINUES. The KV pool does NOT exist yet: this number sizes the pre-resize buffers (piecewise/verify block tables, drafter KV slabs); it is NOT the served-context statement. The length requests are admitted to is decided AFTER the KV pool is realized, by finalize_served_max_context, and may land above or below this bound. Pin --max-context to refuse loud instead.
21:48:02 INFO    dflash pre-sweep: captured 8/8 draft graphs (pool sealed; graph pool now 382 MiB reserved)
21:48:02 INFO    dflash realize pre-sweep: captured 24/24 realization graphs (pool sealed; graph pool now 382 MiB reserved)
21:48:07 INFO    serving grow floor: 316 MiB (base 4 [MEASURED] + arena-regrow 43 [MEASURED: admission-bound 504 MiB against the 604 MiB the pool already holds (serving mark UNMEASURED; the residency is the boot probes')] + max[spec-reserve 263 (=add(device_slots) of verify-tail 61 + verify-logits 62 [64 rows, measured 62, bound 61], dflash-draft 141), tkv-bypass-scratch 0, tkv-prefill-staging 9] + logprobs-tile 0 + media-encode 0 [stochastic_on=True rejection_flag=True measured=61 analytic=409] + token2wav-vocoder 0 + driver-growth 6 [MEASURED_CACHED]) vs gmu_floor 0 MiB vs residency-gate unpaid-graph-exec 0 MiB (predicted 234 MiB for this configuration's capture ladder, 234 MiB already metered onto the card as driver.cudagraph_exec -> 0 MiB still unpaid; 0 member record(s)) vs serving-consumption 0 MiB [MEASURED] -> 316 MiB [tts-codec-peak 0 + tts-kv-growth 0 (5000 ticks)]
21:48:07 INFO    serving grow floor: 1019 MiB (base 4 [MEASURED] + arena-regrow 746 [MEASURED: admission-bound 504 MiB against the 604 MiB the pool already holds (serving mark UNMEASURED; the residency is the boot probes')] + max[spec-reserve 263 (=add(device_slots) of verify-tail 61 + verify-logits 62 [64 rows, measured 62, bound 61], dflash-draft 141), tkv-bypass-scratch 0, tkv-prefill-staging 9] + logprobs-tile 0 + media-encode 0 [stochastic_on=True rejection_flag=True measured=61 analytic=409] + token2wav-vocoder 0 + driver-growth 6 [MEASURED_CACHED]) vs gmu_floor 0 MiB vs residency-gate unpaid-graph-exec 0 MiB (predicted 234 MiB for this configuration's capture ladder, 234 MiB already metered onto the card as driver.cudagraph_exec -> 0 MiB still unpaid; 0 member record(s)) vs serving-consumption 750 MiB [MEASURED] -> 1019 MiB [tts-codec-peak 0 + tts-kv-growth 0 (5000 ticks)]
21:48:09 INFO    serving grow floor: 387 MiB (base 4 [MEASURED] + arena-regrow 114 [MEASURED: admission-bound 504 MiB against the 604 MiB the pool already holds (serving mark UNMEASURED; the residency is the boot probes')] + max[spec-reserve 263 (=add(device_slots) of verify-tail 61 + verify-logits 62 [64 rows, measured 62, bound 61], dflash-draft 141), tkv-bypass-scratch 0, tkv-prefill-staging 9] + logprobs-tile 0 + media-encode 0 [stochastic_on=True rejection_flag=True measured=61 analytic=409] + token2wav-vocoder 0 + driver-growth 6 [MEASURED_CACHED]) vs gmu_floor 0 MiB vs residency-gate unpaid-graph-exec 0 MiB (predicted 234 MiB for this configuration's capture ladder, 234 MiB already metered onto the card as driver.cudagraph_exec -> 0 MiB still unpaid; 0 member record(s)) vs serving-consumption 118 MiB [MEASURED] -> 387 MiB [tts-codec-peak 0 + tts-kv-growth 0 (5000 ticks)]
21:48:10 INFO    serving grow floor: 387 MiB (base 4 [MEASURED] + arena-regrow 114 [MEASURED: admission-bound 504 MiB against the 604 MiB the pool already holds (a drive's mark: the boot probes had 798 MiB, 194 MiB of it residue the KV grow now reads free)] + max[spec-reserve 263 (=add(device_slots) of verify-tail 61 + verify-logits 62 [64 rows, measured 62, bound 61], dflash-draft 141), tkv-bypass-scratch 0, tkv-prefill-staging 9] + logprobs-tile 0 + media-encode 0 [stochastic_on=True rejection_flag=True measured=61 analytic=409] + token2wav-vocoder 0 + driver-growth 6 [MEASURED_CACHED]) vs gmu_floor 0 MiB vs residency-gate unpaid-graph-exec 0 MiB (predicted 234 MiB for this configuration's capture ladder, 234 MiB already metered onto the card as driver.cudagraph_exec -> 0 MiB still unpaid; 0 member record(s)) vs serving-consumption 118 MiB [MEASURED] -> 387 MiB [tts-codec-peak 0 + tts-kv-growth 0 (5000 ticks)]
21:48:10 WARNING max_context=auto: served context NARROWED to 183808 tokens (from the 255744-token sizing bound). This is the REAL served-context statement — the KV pool is realized: 730 servable pages × 256 tokens/page, less the 11-page capture-sentinel prefix and a 0-page allocator margin. The earlier sizing bound was a forecast and came out high. Pin --max-context explicitly to refuse loud instead.
21:48:10 INFO    graph_pool budget cache: REPLACING the cached 1627389952 B (1552.0 MiB) with this boot's 1509949440 B (1440.0 MiB) — the reading is the cuMem per-tag counter (capture.cudagraphs + capture.io_buffers), which is exact for this cache key, so the monotonic ratchet does not apply.
21:48:10 INFO    graph_pool budget cache: persisted 1509949440 B (1440.0 MiB) + measured cudaGraphInstantiate decode=11.75 MiB/graph, dflash=1.44 MiB/graph → /cache/arbi-serve/budget-cache/04f389ba71121f05.json
21:48:10 INFO    Boot: KV cache ready at 3.03 GiB (13% of 23.52 GiB) — 730 pages × 256 tok/page = 186880 servable tokens, free 0.39 GiB (2% of 23.52 GiB)
21:48:10 INFO    max_context=auto: widened served context to 186624 tokens to use the full realized KV pool (730 servable pages).
21:48:11 INFO    Boot: KV cache ready at 3.03 GiB (13% of 23.52 GiB) — 730 pages × 256 tok/page = 186880 servable tokens, free 0.39 GiB (2% of 23.52 GiB)
