Vibra-Ingenn/Janus: Janus is a API router for AI fashions written in Go and has a Vulkan Model runner · GitHub


Janus is a single Go binary that runs .gguf fashions in your machine (GPU or CPU) and exposes an OpenAI-compatible API. No Python, no Docker, no Ollama required.

Use it your means: name it from the command line (curl, PowerShell, scripts), wire it into Cursor / Cline / any OpenAI consumer — similar native fashions, no matter workflow suits you.


  • Local inference — llama.cpp through Vulkan (AMD / Intel / NVIDIA) or CPU fallback
  • OpenAI-compatible API — /v1/chat/completions, /v1/fashions
  • Hot-swap fashions — change .gguf with out restarting
  • Thinking mannequin help — reasoning break up into reasoning_content
  • Chat template auto-detection — makes use of the template from GGUF metadata
  • Zero dependencies — one .exe on Windows, no Python, no Docker

Platform What you want
Windows (main) Windows 10/11, Go 1.22+, Vulkan-capable GPU really useful
Linux Go 1.22+, Vulkan or CPU
macOS Go 1.22+, CPU backend (Vulkan varies by {hardware})

Disk: plan for the mannequin dimension (usually 2–8 GB per mannequin) plus ~50 MB for Janus + llama.dll.


git clone https://github.com/Vibra-Ingenn/Janus.git
cd Janus
.construct.ps1

construct.ps1 downloads pre-built llama.cpp Vulkan DLLs and compiles distjanus.exe.

Put a .gguf file within the fashions folder. Use the included downloader:

go construct -o distmodelget.exe .cmdmodelget
.distmodelget.exe -repo meta-llama/Llama-3.2-3B-Instruct -file Llama-3.2-3B-Instruct-Q8_0.gguf -out .fashions

Or obtain any GGUF from Hugging Face.

Edit .env:

INFERENCE_BACKEND=vulkan
JANUS_MODEL_PATH=./fashions/Llama-3.2-3B-Instruct-Q8_0.gguf
JANUS_MAX_TOKENS=4096

Variable Default Meaning
INFERENCE_BACKEND vulkan vulkan, cpu, or openrouter
JANUS_MODEL_PATH (required) Path to your .gguf file
JANUS_GPU_LAYERS -1 -1 = all layers on GPU, 0 = CPU solely
JANUS_VRAM_CEILING_MB 9216 VRAM price range trace (MiB)
JANUS_MAX_TOKENS 4096 Max tokens per reply
JANUS_LISTEN_ADDR 127.0.0.1:8990 Bind handle

Opens http://127.0.0.1:8990 in your browser.

curl http://127.0.0.1:8990/well being

Quick begin (Linux / macOS)

git clone https://github.com/Vibra-Ingenn/Janus.git
cd Janus
go mod tidy
go construct -o dist/janus ./cmd/janus
cp .env.instance .env
# edit .env — set JANUS_MODEL_PATH and INFERENCE_BACKEND=cpu if no Vulkan
./dist/janus

On Linux you want libllama.so subsequent to the binary or on LD_LIBRARY_PATH.


curl http://127.0.0.1:8990/v1/chat/completions 
  -H "Content-Type: software/json" 
  -d '{
    "mannequin": "native",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Base URL: http://127.0.0.1:8990/v1

Method Path Description
GET /well being Liveness examine (?deep=true for particulars)
GET /v1/fashions Model checklist
POST /v1/chat/completions Chat (streaming supported)
POST /fashions/load Hot-swap mannequin
GET /fashions/checklist Available .gguf information
GET /engine/standing VRAM and backend data

curl http://127.0.0.1:8990/v1/chat/completions 
  -H "Content-Type: software/json" 
  -d '{"mannequin":"native","stream":true,"messages":[{"role":"user","content":"Tell me a joke"}]}'

Connect to Cursor / Cline / different purchasers

Base URL:  http://127.0.0.1:8990/v1
API Key:   (go away clean)

Pitfalls (discovered the arduous means)

Problem What’s happening Fix
“It constructed however my adjustments aren’t there” On Windows, Go cannot overwrite a operating .exe. Stop all janus.exe in Task Manager, then rebuild.
“Address already in use” A leftover course of holds port 8990. Task Manager → finish all janus.exe.
Server begins, chat fails JANUS_MODEL_PATH incorrect or no .gguf in fashions/. Set the trail in .env, put the file in fashions/.
“native engine failed to begin” Missing llama.dll or GPU driver difficulty. Run .construct.ps1. Update GPU drivers, or set INFERENCE_BACKEND=cpu.
Wrong URL Default is http://127.0.0.1:8990, not 8080. Bookmark 8990.
First reply takes endlessly Model loading into VRAM — regular. Wait 10–60s; smaller quants (This autumn) load quicker.


cmd/janus/          Main server (OpenAI-compatible API)
cmd/modelget/       Hugging Face mannequin downloader
inside/engine/    llama.cpp Vulkan/CPU backend
inside/bridge/    DLL loader and FFI bindings
inside/singleton/ Single-instance guard
fashions/             Put .gguf information right here (not dedicated)
dist/               janus.exe + llama.dll after construct

Get-Process -Name "janus" -ErrorAction SilentlyContinue | Stop-Process -Force
go check ./...
go construct -o distjanus.exe .cmdjanus
.distjanus.exe

Or simply .run.ps1. For a full construct together with llama DLLs, use .construct.ps1.

Contributions welcome — see CONTRIBUTING.md.


MIT — see LICENSE.



Source link