Niko1221/Strata: Qwen3.8-Flash-Next on any shopper {hardware}: one-click set up for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, non-obligatory picture enter. · GitHub


English · 简体中文 · 日本語 · Deutsch · Français · Español · Português

Run a 125-billion-parameter AI mannequin by yourself gaming PC
NVIDIA or AMD graphics card (12 GB or extra) · Windows or Linux · free and open supply

A voxel pagoda garden that Strata's model wrote, running in the browser
A voxel pagoda backyard, 1 shot immediate working on an RTX 5070 with Strata (IQ3_S, 128K context) ·
full video (49 s)

Strata runs Qwen3.8-Flash-Next on a standard PC. This is a
massive, sensible AI mannequin that often wants a server. It chats, writes code, reads footage and works together with your apps
and coding brokers. Nothing leaves your PC.

We measured it on two odd gaming PCs. A token is about ¾ of a phrase.

  • Writes solutions: how briskly the reply seems in a brief chat. 60 tokens per second is quicker than you’ll be able to learn.
  • Reads your immediate: how briskly it takes in what you ship (right here a 32K-token doc, code or chat historical past).

NVIDIA: RTX 5070 (12 GB), Ryzen 5 7600, 64 GB RAM AMD: RX 9070 XT (16 GB), Ryzen 9 3900X, 47 GB RAM
Size Writes solutions Reads your immediate
Q2_0 94 tokens/s 2,650 tokens/s
IQ2_XS 79 tokens/s 2,090 tokens/s
IQ3_XXS 62 tokens/s 1,750 tokens/s
IQ3_S 53 tokens/s 1,620 tokens/s
Coder 55 tokens/s 2,180 tokens/s
Size Writes solutions Reads your immediate
Q2_0 60 tokens/s 1,160 tokens/s
IQ2_XS 52 tokens/s 1,110 tokens/s
Coder 44 tokens/s 1,420 tokens/s

NVIDIA: Q2_0 with engine 0.1.36, the opposite rows with 0.1.26 (4K solutions, 32K prompts). The full tables are in
DETAILS.md. A card with extra VRAM is quicker: an RTX 3090 (24 GB) ought to write
about 100-140 tokens per second. Long chats and different playing cards: speed of each model,
community results.

Buy Me A Coffee
Strata is free. If it runs nicely in your PC, a espresso retains the work on it going.

Graphics card NVIDIA GeForce RTX 20, 30, 40 or 50 sequence, or AMD Radeon RX 7900 XT / XTX, RX 7800 XT / 7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700 or RX 6800 / 6900 sequence. It wants 12 GB of VRAM or extra.
RAM 32 GB or extra. Your RAM decides which model matches. 64 GB runs each dimension.
Disk About 80 GB free. Use an SSD for those who can: the primary begin is way sooner.
System Windows 10 / 11 or Linux, and a present graphics driver from NVIDIA or AMD.

The installer units up every thing else. Two or three playing cards can share the mannequin (multi-GPU).

Experimental, written and examined by group members on their very own machines:

  • Older graphics playing cards (Tesla P40 / V100, GTX 10, Radeon VII / MI50, RX 6700 XT, RX 5500 XT): Older GPUs.
  • Intel Arc, constructed from supply on Linux: Intel Arc.
  • Older processors with out AVX2: they work, however slowly. Older CPUs.

The full checklist: docs/INSTALL.md.

Do you utilize an AI coding assistant (Claude Code, Cursor, Codex, GitHub Copilot, …)? Paste this into it:

Set up Strata on this PC for me: https://github.com/Niko1221/Strata - comply with docs/AI_SETUP.md in that repository.

It checks your graphics card, RAM and disk and picks the mannequin that matches. Then it installs and begins it and tells
you easy methods to join your apps. AI instruments may set up, begin and cease Strata by way of its
MCP server.

Download Strata and unzip it (or git clone it).
Windows: double-click START-HERE.bat. Linux: run ./setup.sh within the Strata folder.

The steps are the identical for NVIDIA and AMD. The installer finds your card and units up the suitable engine for it. It
asks you a number of questions:

  • which mannequin and which dimension,
  • how a lot context (how a lot textual content the mannequin retains in thoughts),
  • whether or not it ought to learn footage.

Press Enter every time for the advisable reply. Then it downloads the mannequin (about 70 GB) and begins it. If the
obtain stops, run it once more: it continues the place it left off. Your browser opens the Strata app at
http://127.0.0.1:8080.

While the mannequin begins, your PC may be sluggish or cease responding for 1-3 minutes (longest the primary time).
Strata hundreds 35-55 GB into your RAM and locks a part of it for the graphics card. This is regular. Wait, and do not
shut the window. The window reveals what Strata is doing.

Next time, run START-HERE.bat (or ./setup.sh) once more. It begins immediately and downloads nothing twice. Close
its window to cease the mannequin. UPDATE.bat (./replace.sh) updates Strata with out beginning it. Updating, Docker,
a number of playing cards, the place the information go and each possibility: docs/INSTALL.md.

Which mannequin ought to I choose?

The installer recommends one to your RAM. The identical mannequin is available in a number of sizes, compressed roughly. Smaller
sizes are sooner. Larger sizes are a bit smarter.

Your RAM Take Why
32 GB Coder it matches 32 GB, and it’s made for code (with a 24 GB card, Q2_0 and IQ2_XS run too)
48 GB IQ2_XS (or Q2_0, the quickest) the bigger sizes don’t match
64 GB IQ2_XS (advisable), or IQ3_XXS / IQ3_S each dimension matches; IQ3_S is the most effective and the slowest
96 GB or extra IQ3_S, or Unsloth’s UD-IQ4_XS (~4-bit) room for the biggest sizes with every thing else open

  • Coder: a coding model with half of the specialists eliminated. It reaches 91% of the total
    mannequin’s SWE-bench Verified rating (measured by its authors) and matches 32 GB of RAM. It is weaker outdoors code,
    together with Chinese and different CJK textual content (#438). For these, take Q2_0, IQ2_XS or IQ3_S, which preserve each professional.
  • Swift 1.5: a fine-tune that thinks for a a lot shorter time earlier than it solutions. You
    get the reply sooner, at about the identical high quality.
  • Unsloth UD-IQ4_XS: Unsloth’s ~4-bit model, between IQ3_S and
    UD-Q4_K_XL in high quality. A 94 GB obtain. With lower than ~80 GB of RAM, Strata reads a part of it from the SSD
    whereas it solutions, so it’s slower there (an NVMe SSD helps).
  • Unsloth UD-Q4_K_XL (experimental): the closest to the total
    mannequin. But Strata reads most of it from the SSD whereas it solutions, so it writes solely 7-8.5 tokens/s on a 64 GB PC.
  • OrcaRouter’s Uncensored IQ3_XXS: you set it up by hand. It is
    not within the installer’s menu.

Sizes, downloads and what matches the place: docs/MODELS.md. To add one other mannequin later, run
SETUP.bat (Linux: ./setup.sh --setup).

The Strata app's Monitor tab next to a coding agent
The Strata app’s Monitor (left) whereas a coding agent writes the pagoda backyard from the video (proper)

  • In the browser: open http://127.0.0.1:8080. It has Chat, a dwell Monitor of the mannequin and your
    GPU/CPU/RAM, and About with the settings and addresses.
  • Your apps and coding brokers: add an “OpenAI-compatible” supplier with the bottom URL
    http://127.0.0.1:8080/v1. Any API key and any mannequin title work.

    • Apps that use Anthropic’s API: http://127.0.0.1:8080/v1/messages (Claude Code:
      ANTHROPIC_BASE_URL=http://127.0.0.1:8080).
    • Codex CLI and different apps that use the OpenAI Responses API: /v1/responses
      (setup).
  • Thinking: select off, low, medium or excessive within the chat menu or in your app’s “reasoning effort”. Off is the
    quickest. High is greatest for laborious questions.
  • Pictures: say sure to “Images?” in setup. Then click on Picture within the chat, or connect footage in your app.
    AMD playing cards learn footage on Linux by way of the processor; on Windows they cannot but.
  • From your telephone or one other PC: START-HERE.bat --setup --host 0.0.0.0 --api-key . Always set a key.
  • One request at a time: by default Strata solutions one request, and the others wait. To reply a number of without delay,
    set "parallel": 2 (BATCHING.md). On a 12 GB card this makes every reply slower.
  • Long prompts: Strata reads the primary message of a chat in full, about 1 minute per 30,000 tokens. Follow-up
    messages begin in seconds.

More: where your chats are stored, the API.

  • My PC froze the primary time Strata began. This is regular whereas it hundreds the mannequin. Wait, and do not shut the
    window. Still frozen after 10 minutes? Restart the PC, shut different applications and take a look at once more, or choose a smaller dimension.
  • It stopped whereas downloading or putting in. Run START-HERE.bat (or ./setup.sh) once more. It continues the place
    it stopped.
  • It’s very sluggish and the disk gentle retains blinking, or it says “the engine stopped unexpectedly”. Your PC does
    not have sufficient free RAM. Close different applications (browsers use quite a bit), or choose a smaller dimension (Q2_0 or IQ2_XS).
  • It says port 8080 is already in use. Strata is already working. Look for its window.

More issues and their fixes: docs/TROUBLESHOOTING.md. Still caught? Open an
issue and fasten strata-.log from the Strata folder. Found a
safety downside? Report it privately: SECURITY.md.

Models like this one often run on servers with tons of of gigabytes of graphics reminiscence. Your graphics card has
12-24 GB. Strata makes the mannequin match by sharing the work throughout your entire PC. Think of a kitchen: the issues
you utilize on a regular basis keep on the counter, and the remainder waits within the pantry.

The model's 24,576 experts: the busiest on the graphics card, all of them in RAM, a lookup table on the SSD

  • The mannequin is a crew of 24,576 small specialists (“specialists”). Each phrase wants solely 10 of them.
  • Your graphics card retains the few thousand specialists which are used most frequently. Your RAM holds all of them,
    and your processor works on the remainder on the identical time. Your SSD holds an enormous lookup desk.

A small helper guesses the next words; the big model checks them all at once and keeps the right ones

  • Guess, then examine: a small helper guesses the subsequent few phrases. The huge mannequin checks them all of sudden. You get
    the identical reply, 1.6-1.8x sooner.
  • Long texts are learn in huge items (as much as 8,192 tokens at a time), at over 1,000 tokens per second.

The longer clarification: docs/HOW_IT_WORKS.md. Every half and its numbers:
the details and the paper.

The mannequin is Qwen3.8-Flash-Next by the Qwen crew. It was
compressed by ISTA-DASLab, UkisAI (Swift 1.5)
and Unsloth. Strata makes use of components of llama.cpp / ggml. All credit:
docs/HOW_IT_WORKS.md. Strata is open supply below the MIT License. Just a few
components and each mannequin have their very own licenses (which ones).

Strata is free and open supply. If it’s helpful to you, you’ll be able to help its growth:

Buy Me A Coffee



Source link