Skip to content
Rush Commerce
Tools & Teardowns3 min read

Strata runs a 125B Qwen model on a 12GB gaming GPU

Strata, an MIT-licensed inference engine, runs 125B Qwen3.8-Flash-Next on a 12GB GPU at 94 tokens/s. What local AI on a gaming PC means for your AI bill.

A 125-billion-parameter model on a gaming PC used to be a joke. Strata, an MIT-licensed open-source inference engine, now runs Qwen3.8-Flash-Next (125B) on a single consumer GPU with 12 GB of VRAM. The project reports 94 tokens per second of output on an RTX 5070. For a small business, local AI on hardware you already own just moved from "toy model" to "real model."

What actually happened

The facts, from the Strata README:

  • Hardware floor: 12 GB of VRAM or more, 32 GB of system RAM, about 80 GB of free disk. Windows 10/11 or Linux.
  • GPUs: NVIDIA GeForce RTX 20 through 50 series, plus a list of AMD Radeon cards including the RX 7900 XT/XTX and RX 9070/9070 XT.
  • Speed: on an RTX 5070 at Q2_0 quantization, about 94 tokens/s for generation and 2,650 tokens/s for input processing.
  • Quantizations: Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, a Coder variant, and Unsloth variants up to UD-Q4_K_XL.
  • API: serves OpenAI-compatible and Anthropic-compatible endpoints at http://127.0.0.1:8080/v1.
  • Install: double-click START-HERE.bat on Windows, or run ./setup.sh on Linux.

How does a 125B model fit in 12 GB? It is a mixture-of-experts model. Each token uses only a small set of experts, so Strata keeps the hot experts on the GPU and keeps the rest in system RAM.

The caveats are real. Q2 is aggressive quantization, and quality drops as you squeeze bits. The README warns your PC can be slow or stop responding for 1–3 minutes while the model loads. Third-party write-ups note it handles one request at a time, and independent benchmarks across hardware are still thin. This is a solo developer's project, not a vendor with a support contract.

Why it matters for your business

Local AI is now a budget line, not a research project. If a mid-range gaming GPU in an office PC can run a 125B model, then private drafting, document summaries and code help can run with zero per-token cost and zero data leaving the building. That matters for client files, HR documents and anything under an NDA.

The OpenAI-compatible endpoint is the real feature. Tools that talk to the OpenAI or Anthropic API can point at localhost instead. That is how you build a fallback: same code, swap the base URL. Your cloud vendor raises prices or has an outage, and the work keeps moving.

Do not put it in front of customers yet. One request at a time is fine for one person at a desk. It is not a customer-facing chatbot backend. Test quality on your own tasks before you trust a Q2 model with anything that ships.

Key takeaways

  • Strata runs 125B Qwen3.8-Flash-Next on a 12 GB GPU with 32 GB of RAM
  • Reported speed: about 94 tokens/s output on an RTX 5070 at Q2_0
  • It exposes OpenAI- and Anthropic-compatible APIs on localhost, so existing tools can switch to it
  • It serves one request at a time; use it for internal work, not customer traffic
  • Low-bit quantization costs quality; test on your real tasks first

Want AI that keeps working when your vendor changes the terms? We build vendor-agnostic systems that can swap between cloud APIs and a local model by changing one URL. See how we build or run the numbers on your AI spend.

Sources: Strata on GitHub, AlphaSignal.

  • #local-inference
  • #strata
  • #qwen
  • #open-source
  • #self-hosted-ai
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.