Niko1221/Strata: Qwen3.8-Flash-Next on any shopper {hardware}: one-click set up for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optionally available picture enter. · GitHub
English · 简体中文 · 日本語 · Deutsch · Français · Español · Português
Run a 125-billion-parameter AI mannequin by yourself gaming PC
NVIDIA or AMD graphics card (12 GB or extra) · Windows or Linux · free and open supply

A voxel pagoda backyard, 1 shot immediate operating on an RTX 5070 with Strata (IQ3_S, 128K context) ·
full video (49 s)
Strata runs Qwen3.8-Flash-Next on a traditional PC. This is a
massive, sensible AI mannequin that often wants a server. It chats, writes code, reads photos and works along with your apps
and coding brokers. Nothing leaves your PC.
We measured it on two abnormal gaming PCs. A token is about ¾ of a phrase.
- Writes solutions: how briskly the reply seems in a brief chat. 60 tokens per second is quicker than you possibly can learn.
- Reads your immediate: how briskly it takes in what you ship (right here a 32K-token doc, code or chat historical past).
| NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAM | AMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM | ||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
|
|
NVIDIA: Q2_0 with engine 0.1.36, the opposite rows with 0.1.26 (4K solutions, 32K prompts). The full tables are in
DETAILS.md. A card with extra VRAM is quicker: an RTX 3090 (24 GB) ought to write
about 100-140 tokens per second. Long chats and different playing cards: speed of each model,
community results.
Strata is free. If it runs nicely in your PC, a espresso retains the work on it going.
| Graphics card | NVIDIA GeForce RTX 20, 30, 40 or 50 collection, or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700 or RX 6800 / 6900 collection. It wants 12 GB of VRAM or extra. |
| RAM | 32 GB or extra. Your RAM decides which model matches. 64 GB runs each dimension. |
| Disk | About 80 GB free. Use an SSD in case you can: the primary begin is far sooner. |
| System | Windows 10 / 11 or Linux, and a present graphics driver from NVIDIA or AMD. |
The installer units up every thing else. Two or three playing cards can share the mannequin (multi-GPU).
Experimental, written and examined by group members on their very own machines:
- Older graphics playing cards (Tesla P40 / V100, GTX 10, Radeon VII / MI50, RX 6700 XT, RX 5500 XT): Older GPUs.
- Intel Arc, constructed from supply on Linux: Intel Arc.
- Older processors with out AVX2: they work, however slowly. Older CPUs.
The full checklist: docs/INSTALL.md.
Do you utilize an AI coding assistant (Claude Code, Cursor, Codex, GitHub Copilot, …)? Paste this into it:
Set up Strata on this PC for me: https://github.com/Niko1221/Strata - observe docs/AI_SETUP.md in that repository.
It checks your graphics card, RAM and disk and picks the mannequin that matches. Then it installs and begins it and tells
you the right way to join your apps. AI instruments can even set up, begin and cease Strata by its
MCP server.
Download Strata and unzip it (or git clone it).
Windows: double-click START-HERE.bat. Linux: run ./setup.sh within the Strata folder.
The steps are the identical for NVIDIA and AMD. The installer finds your card and units up the correct engine for it. It
asks you a couple of questions:
- which mannequin and which dimension,
- how a lot context (how a lot textual content the mannequin retains in thoughts),
- whether or not it ought to learn photos.
Press Enter every time for the really useful reply. Then it downloads the mannequin (about 70 GB) and begins it. If the
obtain stops, run it once more: it continues the place it left off. Your browser opens the Strata app at
http://127.0.0.1:8080.
While the mannequin begins, your PC will be gradual or cease responding for 1-3 minutes (longest the primary time).
Strata masses 35-55 GB into your RAM and locks a part of it for the graphics card. This is regular. Wait, and do not
shut the window. The window exhibits what Strata is doing.
Next time, run START-HERE.bat (or ./setup.sh) once more. It begins instantly and downloads nothing twice. Close
its window to cease the mannequin. UPDATE.bat (./replace.sh) updates Strata with out beginning it. Updating, Docker,
a number of playing cards, the place the recordsdata go and each possibility: docs/INSTALL.md.
The installer recommends one on your RAM. The similar mannequin is available in a number of sizes, compressed kind of. Smaller
sizes are sooner. Larger sizes are a bit smarter.
| Your RAM | Take | Why |
|---|---|---|
| 32 GB | Coder | it matches 32 GB, and it’s made for code (with a 24 GB card, Q2_0 and IQ2_XS run too) |
| 48 GB | IQ2_XS (or Q2_0, the quickest) | the bigger sizes don’t match |
| 64 GB | IQ2_XS (really useful), or IQ3_XXS / IQ3_S | each dimension matches; IQ3_S is the most effective and the slowest |
| 96 GB or extra | IQ3_S, or Unsloth’s UD-IQ4_XS (~4-bit) | room for the most important sizes with every thing else open |
- Coder: a coding model with half of the specialists eliminated. It reaches 91% of the total
mannequin’s SWE-bench Verified rating (measured by its authors) and matches 32 GB of RAM. It is weaker outdoors code,
together with Chinese and different CJK textual content (#438). For these, take Q2_0, IQ2_XS or IQ3_S, which maintain each knowledgeable. - Swift 1.5: a fine-tune that thinks for a a lot shorter time earlier than it solutions. You
get the reply sooner, at about the identical high quality. - Unsloth UD-IQ4_XS: Unsloth’s ~4-bit model, between IQ3_S and
UD-Q4_K_XL in high quality. A 94 GB obtain. With lower than ~80 GB of RAM, Strata reads a part of it from the SSD
whereas it solutions, so it’s slower there (an NVMe SSD helps). - Unsloth UD-Q4_K_XL (experimental): the closest to the total
mannequin. But Strata reads most of it from the SSD whereas it solutions, so it writes solely 7-8.5 tokens/s on a 64 GB PC. - OrcaRouter’s Uncensored IQ3_XXS: you set it up by hand. It is
not within the installer’s menu.
Sizes, downloads and what matches the place: docs/MODELS.md. To add one other mannequin later, run
SETUP.bat (Linux: ./setup.sh --setup).

The Strata app’s Monitor (left) whereas a coding agent writes the pagoda backyard from the video (proper)
- In the browser: open
http://127.0.0.1:8080. It has Chat, a reside Monitor of the mannequin and your
GPU/CPU/RAM, and About with the settings and addresses. - Your apps and coding brokers: add an “OpenAI-compatible” supplier with the bottom URL
http://127.0.0.1:8080/v1. Any API key and any mannequin identify work.- Apps that use Anthropic’s API:
http://127.0.0.1:8080/v1/messages(Claude Code:
ANTHROPIC_BASE_URL=http://127.0.0.1:8080). - Codex CLI and different apps that use the OpenAI Responses API:
/v1/responses
(setup).
- Apps that use Anthropic’s API:
- Thinking: select off, low, medium or excessive within the chat menu or in your app’s “reasoning effort”. Off is the
quickest. High is greatest for arduous questions. - Pictures: say sure to “Images?” in setup. Then click on Picture within the chat, or connect photos in your app.
AMD playing cards learn photos on Linux by the processor; on Windows they can not but. - From your telephone or one other PC:
START-HERE.bat --setup --host 0.0.0.0 --api-key. Always set a key. - One request at a time: by default Strata solutions one request, and the others wait. To reply a number of without delay,
set"parallel": 2(BATCHING.md). On a 12 GB card this makes every reply slower. - Long prompts: Strata reads the primary message of a chat in full, about 1 minute per 30,000 tokens. Follow-up
messages begin in seconds.
More: where your chats are stored, the API.
- My PC froze the primary time Strata began. This is regular whereas it masses the mannequin. Wait, and do not shut the
window. Still frozen after 10 minutes? Restart the PC, shut different packages and take a look at once more, or decide a smaller dimension. - It stopped whereas downloading or putting in. Run
START-HERE.bat(or./setup.sh) once more. It continues the place
it stopped. - It’s very gradual and the disk gentle retains blinking, or it says “the engine stopped unexpectedly”. Your PC does
not have sufficient free RAM. Close different packages (browsers use loads), or decide a smaller dimension (Q2_0 or IQ2_XS). - It says port 8080 is already in use. Strata is already operating. Look for its window.
More issues and their fixes: docs/TROUBLESHOOTING.md. Still caught? Open an
issue and fix strata- from the Strata folder. Found a
safety drawback? Report it privately: SECURITY.md.
Models like this one often run on servers with lots of of gigabytes of graphics reminiscence. Your graphics card has
12-24 GB. Strata makes the mannequin match by sharing the work throughout your complete PC. Think of a kitchen: the issues
you utilize on a regular basis keep on the counter, and the remainder waits within the pantry.
- The mannequin is a workforce of 24,576 small specialists (“specialists”). Each phrase wants solely 10 of them.
- Your graphics card retains the few thousand specialists which are used most frequently. Your RAM holds all of them,
and your processor works on the remainder on the similar time. Your SSD holds an enormous lookup desk.
- Guess, then verify: a small helper guesses the following few phrases. The huge mannequin checks them abruptly. You get
the identical reply, 1.6-1.8x sooner. - Long texts are learn in huge items (as much as 8,192 tokens at a time), at over 1,000 tokens per second.
The longer clarification: docs/HOW_IT_WORKS.md. Every half and its numbers:
the details and the paper.
The mannequin is Qwen3.8-Flash-Next by the Qwen workforce. It was
compressed by ISTA-DASLab, UkisAI (Swift 1.5)
and Unsloth. Strata makes use of components of llama.cpp / ggml. All credit:
docs/HOW_IT_WORKS.md. Strata is open supply below the MIT License. A number of
components and each mannequin have their very own licenses (which ones).
Strata is free and open supply. If it’s helpful to you, you possibly can assist its improvement:

