magnitudedev/magnitude: Open supply inference engine for brokers that optimizes itself on your precise {hardware}. Compiles and tunes its kernels in your machine, so open fashions run as much as 2x sooner than llama.cpp. Works on Apple Silicon, NVIDIA, AMD, or nothing however a CPU. · GitHub


Run open fashions as quick as your {hardware} permits

Download Magnitude
Documentation
Discord
Follow Magnitude on Twitter
GitHub Repo stars

Magnitude is an open supply inference engine for brokers that optimizes itself on your precise {hardware}. It compiles and tunes its kernels in your machine, so open fashions run as much as 2x sooner than llama.cpp. One click on connects the agent you already use (Pi, OpenCode, Hermes, Codex, and extra). Works on Apple Silicon, NVIDIA, AMD, or nothing however a CPU.

Download Magnitude for macOS, Windows, or Linux

⭐ Help us attain extra builders and develop the Magnitude neighborhood. Star this repo!


demo-9-29.mp4


  1. Download Magnitude, set up it, and open the app.
  2. Choose a really helpful mannequin in Discover and obtain it.
  3. Connect your agent in Connections and begin utilizing it.

The desktop app consists of the magnitude CLI. No separate set up is required.

  • Up to 2x sooner than llama.cpp: 92% sooner decode on Metal, 19% on CUDA
  • Tuned in your machine: kernels are tuned in your {hardware} earlier than a mannequin runs
  • Built for the very best fashions: hand-optimized kernels for fashionable open-weight households
  • Memory that flexes: 27% much less reminiscence per agent, freed when brokers cease
  • Fast concurrent classes: classes share prefix caches to stop slowdown
  • Works along with your agent: one click on to attach Pi, OpenCode, Hermes, Codex, and extra
  • Free, non-public, open supply: no token prices, nothing leaves your machine, Apache 2.0

Up to 2x sooner than llama.cpp

Magnitude vs llama.cpp: 9% faster prefill and 92% faster decode on Metal, 23% faster prefill and 19% faster decode on CUDA

An open supply inference engine that optimizes itself on your {hardware}. It ships as a desktop app that runs open fashions and connects them to the agent you already use.

How is it sooner than llama.cpp, Ollama, or LM Studio?

They ship kernels precompiled for broad lessons of {hardware}. Magnitude compiles and tunes its kernels in your precise machine earlier than a mannequin runs, so that they suit your precise chip. See the benchmarks against llama.cpp.

Any Apple Silicon, NVIDIA, or AMD GPU, or nothing however a CPU. There is not any mounted minimal. Smaller machines run smaller fashions, and extra reminiscence allows you to run bigger ones.

What working methods does it help?

macOS, Linux, and Windows.

Which fashions does it help?

See the complete listing at magnitude.dev/models. We write optimized kernels for the preferred open-weight households, which is how we beat generalist engines.

Which brokers work with it?

One click on connects Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. Anything else works by way of the OpenAI-compatible API.

Yes. Prompts, information, and fashions keep in your machine. No web wanted as soon as a mannequin is downloaded.



Source link