magnitudedev/magnitude: Open supply inference engine for brokers that optimizes itself on your precise {hardware}. Compiles and tunes its kernels in your machine, so open fashions run as much as 2x sooner than llama.cpp. Works on Apple Silicon, NVIDIA, AMD, or nothing however a CPU. · GitHub

Run open fashions as quick as your {hardware} permits
Magnitude is an open supply inference engine for brokers that optimizes itself on your precise {hardware}. It compiles and tunes its kernels in your machine, so open fashions run as much as 2x sooner than llama.cpp. One click on connects the agent you already use (Pi, OpenCode, Hermes, Codex, and extra). Works on Apple Silicon, NVIDIA, AMD, or nothing however a CPU.
Download Magnitude for macOS, Windows, or Linux
⭐ Help us attain extra builders and develop the Magnitude neighborhood. Star this repo!
demo-9-29.mp4
- Download Magnitude, set up it, and open the app.
- Choose a really helpful mannequin in Discover and obtain it.
- Connect your agent in Connections and begin utilizing it.
The desktop app consists of the magnitude CLI. No separate set up is required.
- Up to 2x sooner than llama.cpp: 92% sooner decode on Metal, 19% on CUDA
- Tuned in your machine: kernels are tuned in your {hardware} earlier than a mannequin runs
- Built for the very best fashions: hand-optimized kernels for fashionable open-weight households
- Memory that flexes: 27% much less reminiscence per agent, freed when brokers cease
- Fast concurrent classes: classes share prefix caches to stop slowdown
- Works along with your agent: one click on to attach Pi, OpenCode, Hermes, Codex, and extra
- Free, non-public, open supply: no token prices, nothing leaves your machine, Apache 2.0
An open supply inference engine that optimizes itself on your {hardware}. It ships as a desktop app that runs open fashions and connects them to the agent you already use.
They ship kernels precompiled for broad lessons of {hardware}. Magnitude compiles and tunes its kernels in your precise machine earlier than a mannequin runs, so that they suit your precise chip. See the benchmarks against llama.cpp.
Any Apple Silicon, NVIDIA, or AMD GPU, or nothing however a CPU. There is not any mounted minimal. Smaller machines run smaller fashions, and extra reminiscence allows you to run bigger ones.
macOS, Linux, and Windows.
See the complete listing at magnitude.dev/models. We write optimized kernels for the preferred open-weight households, which is how we beat generalist engines.
One click on connects Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. Anything else works by way of the OpenAI-compatible API.
Yes. Prompts, information, and fashions keep in your machine. No web wanted as soon as a mannequin is downloaded.
