Stopped

Pick a model and press Start Server.

Model

Required. Any MLX model id or a path on disk.

KV compression

Network

Default generation

API endpoints

Available once the server is running.

    Memory

    Available once the server is running.

    Server log

    No output yet.

    Methods

    MethodTypeStatusCompression info

    About VeloxQuant

    Run local AI models on your Mac, using less memory.

    What this app does

    VeloxQuant lets you run a language model entirely on your own Mac — nothing you type ever leaves your computer. This panel is the easy way to do it: pick a model, press Start, and you have a private AI server ready to chat with in a couple of minutes.

    Good to know

    Any number on this panel that's measured comes straight from your Mac while the server is running, so you can trust it.

    Compression numbers are a little different: they show how well your conversation data compresses, not memory your Mac has actually freed up yet — that's on the way in a future update. We'd rather tell you that plainly than let you assume something we haven't built yet.