Pick a model and press Start Server.
Model
Required. Any MLX model id or a path on disk.
KV compression
Network
⚠ 0.0.0.0 exposes the server to your local network with no authentication.
Default generation
API endpoints
Available once the server is running.
Memory
measuredAvailable once the server is running.
Server log
No output yet.
Methods
| Method | Type | Status | Compression info |
|---|
About VeloxQuant
Run local AI models on your Mac, using less memory.
What this app does
VeloxQuant lets you run a language model entirely on your own Mac — nothing you type ever leaves your computer. This panel is the easy way to do it: pick a model, press Start, and you have a private AI server ready to chat with in a couple of minutes.
Good to know
Any number on this panel that's measured comes straight from your Mac while the server is running, so you can trust it.
Compression numbers are a little different: they show how well your conversation data compresses, not memory your Mac has actually freed up yet — that's on the way in a future update. We'd rather tell you that plainly than let you assume something we haven't built yet.