Strata is an open-source inference engine built to run Qwen3.8-Flash-Next, a 125B-parameter MoE model, on an ordinary gaming PC. Holding the whole model in VRAM takes about 75GB even at 2-bit; with Strata, one 12GB graphics card plus 64GB of RAM runs it, and the author's RTX 5070 measured 53 to 94 tokens/s. With the experts in RAM in both cases, it is more than twice as fast as llama.cpp. At eBay prices on October 4, an RTX 5070 build with 64GB of RAM costs about $2,300, about $1,500 less than a Ryzen AI Max+ 395 128GB, and is more than twice as fast. Data as of October 4, 2026, for Strata v0.1.39.
What Is Strata? Run Qwen3.8-Flash-Next on a 12GB GPU (2026)
Translated from the 简体中文 edition · Read the original
Figures in the text and FAQ are as of Oct 4, 2026. Tables use the latest published prices; collection dates may differ by device.

What is Strata?
An inference engine written only for Qwen3.8-Flash-Next, open-sourced on GitHub in September 2026 under the MIT license.
What sets it apart from general-purpose tools like llama.cpp and Ollama: it runs only Qwen3.8-Flash-Next and its Coder and Swift derivatives, and every step is tailored to this model's architecture. It supports Windows and Linux, comes with a web chat interface once installed, and offers OpenAI- and Anthropic-compatible APIs, so it can plug into coding tools like Claude Code.
Why is a 12GB GPU enough?
Because only a small part of the model is used for every token.
For every token it writes, the model has to read all the parameters that token uses, and only VRAM reads fast enough. Each of Qwen3.8-Flash-Next's 48 layers holds two kinds of weights:
- The shared part: attention, the router and so on. Every token uses it; 3.5GB across the whole model.
- Experts: 512 per layer, 24,576 in all, 34GB. Each token picks only 10 per layer, 2% of them.
On top of that there is a 28.8GB lookup table, and each token reads only a dozen or so rows of it. The table holds 51B parameters, which is why this site's model database lists the total as 176B.

The model is 66GB, but each token reads only about 4.2GB, and 3.5GB of that is the shared part. So Strata puts only the shared part and the roughly 4,500 most-used experts on the GPU, all the experts in RAM, and the lookup table on the SSD; which rows a token needs can be worked out in advance, so the SSD does not slow things down. The model has not shrunk. The bulk has moved from VRAM to RAM, which is why you need 64GB of RAM.
Strata vs llama.cpp: what's different?
Against holding the whole model in VRAM, Strata cuts the VRAM requirement from about 75GB to 12GB, more than 80 percent less. Against llama.cpp with the experts in RAM, both need little VRAM and the difference is speed: on the same machine Strata is more than twice as fast.

How much VRAM the whole model needs, calculated live from this site's model database:
Amazon.jp
| Model | 2-bit | 4-bit | Cheapest card that runs it | ||
|---|---|---|---|---|---|
| 8K Context | 32K Context | 8K Context | 32K Context | ||
| Qwen3.8 Flash Next176B · A6B | 75 GBOffload: 4 GB VRAM + 73 GB RAM | 76 GBOffload: 4 GB VRAM + 73 GB RAM | 105 GBOffload: 4 GB VRAM + 103 GB RAM | 106 GBOffload: 5 GB VRAM + 103 GB RAM | |
Strata is faster in three ways.
1. It picks the most-used experts for the GPU
llama.cpp splits by layer: a 12GB card has about 6GB left after the shared part, each layer's experts take about 0.7GB, so only about 8 layers fit, a sixth of the 48. Strata picks the roughly 4,500 most-used experts from all layers and fits them into the same 6GB, swapping them as the conversation goes; in measurements, 72% of the experts each token needed were already on the GPU. The CPU reads only about a third as much from RAM as with the per-layer split.
2. The GPU and CPU compute at the same time

In llama.cpp the GPU waits while the CPU computes the experts. Strata splits the 10 experts picked in each layer: the GPU computes the roughly 7 it holds, and at the same time the CPU computes the other 3 or so where they sit in RAM. Missing experts are not copied to the GPU, because PCIe 4.0 x16 manages only about 26 GB/s, slower than the CPU reading RAM directly (41 to 52 GB/s).
3. It guesses first and checks after, for several tokens per pass

The model ships with a draft layer (an MTP prediction layer) that guesses up to the next 3 tokens; the main model checks them in a single pass and keeps the correct ones. The most expensive part of a pass is reading the 3.5GB shared part, and that happens once however many positions are checked. In measurements this lifted speed from 47–57 tokens/s to 82–92, with output identical to running without guessing.
What are the trade-offs?
- Lower precision: it uses 2- to 3-bit quantized versions; Q2_0 is the fastest and IQ3_S the closest to the original. How quantization works: What Is a Quantized Model.
- One model only: for any other model you still need llama.cpp or Ollama.
- One user at a time: it handles one request at a time, so it is not suited to serving a team.
- The Coder version is weak outside code: it drops half the experts, and users report garbled answers in languages other than English. For general use, pick Q2_0, IQ2_XS or IQ3_S.
Which GPUs can run Qwen3.8-Flash-Next?
NVIDIA RTX 20 series or newer and mainstream AMD RX 7000 and 9000 cards, with 12GB of VRAM or more, plus 64GB of RAM.
- NVIDIA: GeForce RTX 20, 30, 40 and 50 series with 12GB of VRAM or more.
- AMD: RX 7900 XT/XTX, 7800 XT, 7700 XT, 9060 XT, 9070/9070 XT and Radeon AI PRO R9700; the RX 6800/6900 series installs but is not officially validated.
- Experimental: community-built experimental versions exist for Tesla P40, V100, Instinct MI50 and Intel Arc; they need a separate install or your own build (notes).
Pick by VRAM size, not by generation. Each extra 1GB of VRAM holds about 700 more experts and leaves the CPU a little less to do:

The RTX 5060 Ti has 4GB more VRAM but less bandwidth (448 vs 672 GB/s), so it runs about as fast as the 5070; the RTX 3090 doubles the VRAM and has more bandwidth too, an estimated 50 percent faster. The RTX 5090 row comes from a community report. This site has not re-tested any of these numbers.
Current prices of the supported cards this site tracks:
Amazon.jp
| Model | VRAM GB | Bandwidth GB/s | INT8 TOPS | Power W | Amazon.jp | Price per GB |
|---|---|---|---|---|---|---|
| 12 | 360 | 102 | 170 | $540 | $45 | |
| Radeon RX 7800 XT | 16 | 624 | 75 | 263 | $822 | $51 |
| Radeon RX 9070 XT | 16 | 640 | 389 | 304 | $937 | $59 |
| 16 | 448 | 190 | 180 | $949 | $59 | |
| 12 | 672 | 247 | 250 | $1,056 | $88 | |
| 16 | 288 | 177 | 165 | $1,323 | $83 | |
| 16 | 896 | 352 | 300 | $1,457 | $91 | |
| 24 | 936 | 285 | 350 | $2,647 | $110 | |
| 24 | 1,008 | 661 | 450 | $4,752 | $198 | |
| 32 | 1,792 | 838 | 575 | $6,962 | $218 |
Grey prices carry over the last valid quote; no new listings were seen recently.
- Cheapest that runs it: RTX 3060 12GB, about $450 on eBay, but nobody has tested it yet; its memory bandwidth is only a little over half the RTX 5070's, so it will be slower.
- Clearly faster: RTX 3090 24GB. What to check when buying one is in RTX 3090 vs RTX 5070 Ti for local LLMs.
RAM decides which version you can run (model notes):

Buy two 32GB sticks, not four 16GB: DDR5 often has to clock down with all four slots filled. Experts that are not on the GPU are computed by the CPU, so a DDR5 platform and a CPU with AVX-512 (AMD Ryzen 7000 and 9000 series) are faster. DDR4 has about half the bandwidth of DDR5; it runs, but will not reach the official numbers.
How much does Strata save?
Comparing complete builds: an RTX 5070 Strata build with 64GB of RAM costs about $2,300, about $1,500 less than a Ryzen AI Max+ 395 128GB, and is more than twice as fast at 2-bit. DGX Spark costs about $9,000, about four times as much, and is still slower than the Strata build.

The strength of the 395 and DGX Spark is that they can also run 4-bit and other models. If all you want is Qwen3.8-Flash-Next, the Strata build is the better deal.
How to install Strata
Install the GPU driver, download the project, then double-click START-HERE.bat on Windows or run ./setup.sh on Linux and press Enter through the prompts. The installer sets up Python, the inference engine and the model by itself (install guide).
- Install the driver: NVIDIA needs driver 580 or newer; AMD uses a recent Adrenalin on Windows and the system's built-in driver on Linux.
- Free up disk space: the model is about 70 to 80GB; Q2_0 on a CPU with AVX-512 needs about 40GB more.
- Download the project: download the zip from GitHub and extract it, or
git cloneit. - Run the installer: double-click
START-HERE.baton Windows or run./setup.shon Linux. It detects your GPU and recommends a version for your RAM; press Enter to accept. - Wait for the download and startup: the model is about 70GB, and the download resumes if interrupted. It starts by itself when done; open
http://127.0.0.1:8080in a browser to chat. Startup reads 35 to 55GB into RAM, and the computer may stutter for 1 to 3 minutes. That is normal.
After that, double-clicking START-HERE.bat starts it directly, loading in 30 to 90 seconds; closing the window stops it.
- Connect other software: add an "OpenAI-compatible" provider with the address
http://127.0.0.1:8080/v1; the key and model name can be anything. For Claude Code, setANTHROPIC_BASE_URL=http://127.0.0.1:8080. - Multiple GPUs: two or three NVIDIA cards can work together, no NVLink needed (multi-GPU notes, and this site's multi-GPU guide).
Data sources: the model architecture and Strata's speed and memory requirements come from the Strata repository's README, docs, paper and community reports, and the llama.cpp comparison comes from issue #28; this site has not re-tested any of them. Speeds for the Ryzen AI Max+ 395 and DGX Spark are this site's estimates. Hand-written prices are eBay median prices on October 4, 2026 (CPU, board and power supply are this site's reference prices); live tables are generated from this site's database.
FAQ
How is Strata different from Ollama and LM Studio?
Ollama and LM Studio are general-purpose tools that run any model. Strata runs only Qwen3.8-Flash-Next and is built around its architecture: it picks the most-used experts for the GPU, runs the GPU and CPU at the same time, and uses the model's built-in prediction layer to produce several tokens per pass. For this model Strata is much faster; for any other model it does not work.
I only have 32GB of RAM. Can I use it?
Only the Coder version, which drops half the experts; it is weak outside code, and users report garbled answers in languages other than English. With a 24GB GPU, Q2_0 and IQ2_XS also fit. 48GB runs the two smallest versions (Q2_0 and IQ2_XS), and 64GB runs every version.
On a tight budget, should I add RAM or upgrade the GPU first?
RAM first. Without enough RAM the model will not start; the GPU only affects speed. Once you have 64GB, upgrade the GPU, and judge it mainly by VRAM size.
Does compressing it to 2 or 3 bits make the model dumber?
It loses some. Q2_0 is the fastest and IQ3_S the closest to the original. ISTA-DASLab, which made the quantized versions, says IQ3_S matches the original model on public benchmarks; this site has not re-tested that. If quality matters, run IQ3_S with 64GB of RAM.
Does it work on AMD GPUs and Macs?
AMD GPUs, yes: RX 7900 XT/XTX, 7800 XT, 7700 XT, 9060 XT, 9070/9070 XT and Radeon AI PRO R9700, with 12GB of VRAM or more; on Windows, image input does not work yet. Macs are not supported; Strata has only Windows and Linux versions.



