Strata, an open-source inference engine published on GitHub, claims to run a 125-billion-parameter model on a single consumer graphics card with 12 GB of memory. According to aiweekly.co on October 4, 2026, the model is Qwen3.8-Flash-Next, the card is a 12 GB GPU backed by 32 GB of system memory, and output speed is 60 to 94 tokens per second. This explainer checks what the project itself says, separates reported numbers from derived arithmetic, and lists the caveats that matter before treating it as a general result.
Qwen
NVIDIAWhat Strata claims
The Strata README describes the model as "a team of 24,576 small specialists" of which each word needs only 10. It lists an NVIDIA RTX 5070 at 94 tokens per second of output and 2,650 tokens per second of input at Q2_0 quantization, and 53 tokens per second at IQ3_S. aiweekly.co adds an AMD RX 9070 XT at 60 tokens per second and labels every figure self-reported and without independent validation. The project is MIT licensed, and the model carries separate Qwen terms, according to a fork of the repository.

Why a 125B mixture-of-experts model can fit
A mixture-of-experts (MoE) model splits its feed-forward layers into many experts and routes each token to only a few. The README's figure of 10 active experts out of 24,576 means about 0.04% of the experts run per token (derived: 10 divided by 24,576). The README does not state the model's active parameter count, which also includes shared layers such as attention, so the real per-token compute cannot be derived from these numbers alone.
The practical effect is that a token touches a small slice of the weights, so the rest can sit somewhere slower without being read every step.
Expert offloading across GPU, RAM and SSD
Strata's README assigns each memory tier a job. The graphics card "keeps the few thousand experts that are used most often," system RAM "holds all of them," and the SSD "holds a big lookup table." Hardware requirements listed there are an NVIDIA RTX 20, 30, 40 or 50 series or AMD Radeon RX 7900 or 9000 series card with 12 GB or more, 32 GB or more of RAM and about 80 GB of free disk, per the jhohertz fork's README summary. A lavx news write-up reports that the first load can freeze the machine for one to three minutes.
This is a caching design. If the experts that are used most often sit in GPU memory, most tokens can avoid the slow path, and a miss means fetching an expert from RAM. The README does not publish hit rates, so how often a miss occurs under real workloads is unknown from the sources.
The model is a team of 24,576 small specialists, and each word needs only 10 of them.
Strata README
Quantization bit widths, derived
Weight memory is parameters times bits per weight divided by eight. For 125 billion parameters, that is roughly 31 GB at 2 bits, 47 GB at 3 bits, 62.5 GB at 4 bits, 125 GB at 8 bits and 250 GB at 16 bits (all derived, weights only). The README's quantization options (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S) range from about 30 GB to about 70 GB per the fork summary, which is consistent with the 2-bit to 3-bit rows once format overhead is included.
Weight memory for a 125B-parameter model by bit width
Derived: 125 billion parameters times bits per weight, divided by 8, in gigabytes. The dashed line marks 12 GB of GPU memory.
Weights only. Real quantization formats add scaling metadata, and the KV cache and runtime buffers need extra memory on top.
Even the 2-bit row is far above 12 GB, which is why the engine needs system RAM at all. The model fits the machine, not the card. Lower bit widths also trade quality for size. The README's own numbers show the trade: 94 tokens per second at Q2_0 against 53 at IQ3_S on the same card.
A bandwidth sanity check, derived
If this were a dense 125B model at 2 bits, every token would require reading about 31 GB of weights (derived). At 94 tokens per second that is roughly 2.9 TB per second of memory traffic (derived), a very large figure that makes reading every weight on every token implausible on consumer hardware. Reaching the reported speed therefore requires reading only a small fraction of weights per token, which MoE routing and a hot-expert cache provide. The README also credits a speculative drafter, in which "a small helper guesses the next few words" and the big model checks them at once, for a 1.6 to 1.8 times speedup.
How it compares with llama.cpp, Ollama and vLLM
Strata says it uses parts of llama.cpp and ggml, and the README gives no head-to-head benchmark against them. The lavx article places it as a step beyond llama.cpp, Ollama, LM Studio and Jan, but that is the outlet's framing, not a measurement. The comparison that can be made safely is architectural: Strata's headline feature is a tiered placement of experts across GPU, RAM and SSD plus built-in speculation, packaged with a one-click installer and an OpenAI and Anthropic-compatible API on localhost.
Caveats before you rely on it
- Benchmark conditions. The README states "4K answers, 32K prompts" for its figures, so shorter or longer workloads may differ.
- Independent testing exists but differs. A hands-on note used an RTX 5090 with 32 GB, 128 GB of RAM and IQ3_S, and measured 114 tokens per second on a 400-token code task. That is a different, larger machine, so it does not confirm the 12 GB claim.
- Quality loss. The same author reported 96.7% on 60 basic-to-applied coding questions and 68.8% on 16 hard ones, which is one person's small test, not a standard benchmark.
- Setup friction. The author found that updating required re-running setup to fetch new engine binaries, so inspect what an installer downloads before running it.
- Forks. Several repositories carry near-identical READMEs, so check you are reading the original.
What to watch
Independent runs on genuine 12 GB cards, published expert hit rates, quality scores at Q2_0 versus higher bit widths, and long-context behavior will decide whether this is a general technique or a favorable configuration. Teams that want to compare local and hosted models on their own tasks can do so in a model-agnostic workspace such as Metir. For related open-weight coverage, see Alibaba's Qwen 3.8 Max.
Sources:
- aiweekly.co: Strata runs 125B Qwen3.8-Flash-Next on a 12GB gaming GPU
- Strata repository on GitHub
- Strata fork README (jhohertz)
- lavx news: Strata runs 125B-parameter Qwen model on consumer GPUs
- zephel01: two days with Strata on a single GPU plus RAM
Image credits
- Hero and in-body: "2023 Gigabyte GeForce RTX 4070 Aero OC 12GB (3).jpg" by Jacek Halicki, licensed CC BY-SA 4.0, via Wikimedia Commons. It shows a 12 GB consumer graphics card as a class example and does not depict the hardware Strata was benchmarked on.