vram · resident weights — one model's GPU state at a time
§ 03model catalog
curated, sha-pinned downloads — one at a time, resume on retry
model
size
ctx
needs
loading the catalog…
§ 04streams
no active streams — POST a chat completion and watch it flow
nothing finished yet
§ 05prefix cache
cached conversation
tok
pages
hits
age
idle
nothing cached yet — finish a >1-page prompt and re-send it to watch TTFT collapse
§ 06chat
talk to the served model — the conversation lives in this page only
§ 07asr studio
record or upload a clip, then run — speech spans overlay the waveform, segments play back on click
rtf per job · audio s / wall s
asr jobs
no jobs yet
§ 08speech studio
0 / 4096 characterstype a line and press speak — the waveform plays once, click it to replay
§ 09config
…
§ 10config editor
name
path
backend
effective
ctx
quant
default
rows save as the [[models]] roster — blank cells inherit the flat defaults below; ⚙ opens per-model overrides
§ 11controls
gc runs at the next lifecycle safe point · shutdown drains accepted work, then exits
§ 12benchmark
quiesced only — refuses while streams are active · pp512 and tg128 on the served model, three reps each, the text, audio and speech routes holding meanwhile (seconds on a small model)
pp512 tok/s
tg128 tok/s
the full bench record (hardware, command lines, shas) — paste into a benchmark submission