Vibra-Ingenn/Janus: Janus is a API router for AI fashions written in Go and has a Vulkan Model runner · GitHub

Janus is a single Go binary that runs .gguf fashions in your machine (GPU or CPU) and exposes an OpenAI-compatible API. No Python, no Docker, no Ollama required.
Use it your means: name it from the command line (curl, PowerShell, scripts), wire it into Cursor / Cline / any OpenAI consumer — similar native fashions, no matter workflow suits you.
- Local inference — llama.cpp through Vulkan (AMD / Intel / NVIDIA) or CPU fallback
- OpenAI-compatible API —
/v1/chat/completions,/v1/fashions - Hot-swap fashions — change
.ggufwith out restarting - Thinking mannequin help —
reasoning break up intoreasoning_content - Chat template auto-detection — makes use of the template from GGUF metadata
- Zero dependencies — one
.exeon Windows, no Python, no Docker
| Platform | What you want |
|---|---|
| Windows (main) | Windows 10/11, Go 1.22+, Vulkan-capable GPU really useful |
| Linux | Go 1.22+, Vulkan or CPU |
| macOS | Go 1.22+, CPU backend (Vulkan varies by {hardware}) |
Disk: plan for the mannequin dimension (usually 2–8 GB per mannequin) plus ~50 MB for Janus + llama.dll.
git clone https://github.com/Vibra-Ingenn/Janus.git
cd Janus
.construct.ps1
construct.ps1 downloads pre-built llama.cpp Vulkan DLLs and compiles distjanus.exe.
Put a .gguf file within the fashions folder. Use the included downloader:
go construct -o distmodelget.exe .cmdmodelget
.distmodelget.exe -repo meta-llama/Llama-3.2-3B-Instruct -file Llama-3.2-3B-Instruct-Q8_0.gguf -out .fashions
Or obtain any GGUF from Hugging Face.
Edit .env:
INFERENCE_BACKEND=vulkan
JANUS_MODEL_PATH=./fashions/Llama-3.2-3B-Instruct-Q8_0.gguf
JANUS_MAX_TOKENS=4096
| Variable | Default | Meaning |
|---|---|---|
INFERENCE_BACKEND |
vulkan |
vulkan, cpu, or openrouter |
JANUS_MODEL_PATH |
(required) | Path to your .gguf file |
JANUS_GPU_LAYERS |
-1 |
-1 = all layers on GPU, 0 = CPU solely |
JANUS_VRAM_CEILING_MB |
9216 |
VRAM price range trace (MiB) |
JANUS_MAX_TOKENS |
4096 |
Max tokens per reply |
JANUS_LISTEN_ADDR |
127.0.0.1:8990 |
Bind handle |
Opens http://127.0.0.1:8990 in your browser.
curl http://127.0.0.1:8990/well being
git clone https://github.com/Vibra-Ingenn/Janus.git
cd Janus
go mod tidy
go construct -o dist/janus ./cmd/janus
cp .env.instance .env
# edit .env — set JANUS_MODEL_PATH and INFERENCE_BACKEND=cpu if no Vulkan
./dist/janus
On Linux you want libllama.so subsequent to the binary or on LD_LIBRARY_PATH.
curl http://127.0.0.1:8990/v1/chat/completions
-H "Content-Type: software/json"
-d '{
"mannequin": "native",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Base URL: http://127.0.0.1:8990/v1
| Method | Path | Description |
|---|---|---|
| GET | /well being |
Liveness examine (?deep=true for particulars) |
| GET | /v1/fashions |
Model checklist |
| POST | /v1/chat/completions |
Chat (streaming supported) |
| POST | /fashions/load |
Hot-swap mannequin |
| GET | /fashions/checklist |
Available .gguf information |
| GET | /engine/standing |
VRAM and backend data |
curl http://127.0.0.1:8990/v1/chat/completions
-H "Content-Type: software/json"
-d '{"mannequin":"native","stream":true,"messages":[{"role":"user","content":"Tell me a joke"}]}'
Base URL: http://127.0.0.1:8990/v1
API Key: (go away clean)
| Problem | What’s happening | Fix |
|---|---|---|
| “It constructed however my adjustments aren’t there” | On Windows, Go cannot overwrite a operating .exe. |
Stop all janus.exe in Task Manager, then rebuild. |
| “Address already in use” | A leftover course of holds port 8990. | Task Manager → finish all janus.exe. |
| Server begins, chat fails | JANUS_MODEL_PATH incorrect or no .gguf in fashions/. |
Set the trail in .env, put the file in fashions/. |
| “native engine failed to begin” | Missing llama.dll or GPU driver difficulty. |
Run .construct.ps1. Update GPU drivers, or set INFERENCE_BACKEND=cpu. |
| Wrong URL | Default is http://127.0.0.1:8990, not 8080. |
Bookmark 8990. |
| First reply takes endlessly | Model loading into VRAM — regular. | Wait 10–60s; smaller quants (This autumn) load quicker. |
cmd/janus/ Main server (OpenAI-compatible API)
cmd/modelget/ Hugging Face mannequin downloader
inside/engine/ llama.cpp Vulkan/CPU backend
inside/bridge/ DLL loader and FFI bindings
inside/singleton/ Single-instance guard
fashions/ Put .gguf information right here (not dedicated)
dist/ janus.exe + llama.dll after construct
Get-Process -Name "janus" -ErrorAction SilentlyContinue | Stop-Process -Force
go check ./...
go construct -o distjanus.exe .cmdjanus
.distjanus.exe
Or simply .run.ps1. For a full construct together with llama DLLs, use .construct.ps1.
Contributions welcome — see CONTRIBUTING.md.
MIT — see LICENSE.
