dasllama-server connecting
gen tok/s
active / max streams
queued
tokens generated
admitted / finished / evicted
avg decode batch
prefix cache: tokens / pages
ttft last / avg (ms)
weights + kv (GB)
mtp accepted / drafted
asr ready·active·pending
§ 01live
generation tok/s · last 2 min
prefill tok/s · last 2 min
stream slots · last 2 min
prefilldecodequeued (count)
§ 02models
§ 03model catalog
curated, sha-pinned downloads — one at a time, resume on retry
modelsizectxneeds
loading the catalog…
§ 04streams
no active streams — POST a chat completion and watch it flow
§ 05prefix cache
cached conversationtokpageshitsageidle
nothing cached yet — finish a >1-page prompt and re-send it to watch TTFT collapse
§ 06chat
talk to the served model — the conversation lives in this page only
§ 09config
§ 10config editor
namepathbackendeffectivectxquantdefault
rows save as the [[models]] roster — blank cells inherit the flat defaults below; ⚙ opens per-model overrides
§ 11controls
gc runs at the next lifecycle safe point · shutdown drains accepted work, then exits
§ 12benchmark
quiesced only — refuses while streams are active · pp512 and tg128 on the served model, three reps each, the text, audio and speech routes holding meanwhile (seconds on a small model)
§ 13sidecar exchange
loading…