—
Changing these reloads the model. Nothing else is affected.
The model keeps what it has already read, so a follow-up question does not read the whole conversation again. Measured on a document plus two questions: 11.2s to the first word, then 1.1s and 0.8s.
What the engine has been doing. It reconfigures itself while it runs -- handing memory back, forgetting prompts, rebuilding the pool -- and each of those changes how fast the next reply arrives.
Point a coding agent at this model. It runs on this Mac; nothing you type or any file it reads leaves the machine.
Every tool an agent asks this model to call shows up here.
Measured over every reply this server has produced.