Strata runs a 125B Qwen model on a 12GB gaming GPU
Strata, an MIT-licensed inference engine, runs 125B Qwen3.8-Flash-Next on a 12GB GPU at 94 tokens/s. What local AI on a gaming PC means for your AI bill.
A 125-billion-parameter model on a gaming PC used to be a joke. Strata, an MIT-licensed open-source inference engine, now runs Qwen3.8-Flash-Next (125B) on a single consumer GPU with 12 GB of VRAM. The project reports 94 tokens per second of output on an RTX 5070. For a small business, local AI on hardware you already own just moved from "toy model" to "real model."
What actually happened
The facts, from the Strata README:
- Hardware floor: 12 GB of VRAM or more, 32 GB of system RAM, about 80 GB of free disk. Windows 10/11 or Linux.
- GPUs: NVIDIA GeForce RTX 20 through 50 series, plus a list of AMD Radeon cards including the RX 7900 XT/XTX and RX 9070/9070 XT.
- Speed: on an RTX 5070 at Q2_0 quantization, about 94 tokens/s for generation and 2,650 tokens/s for input processing.
- Quantizations: Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, a Coder variant, and Unsloth variants up to UD-Q4_K_XL.
- API: serves OpenAI-compatible and Anthropic-compatible endpoints at
http://127.0.0.1:8080/v1. - Install: double-click
START-HERE.baton Windows, or run./setup.shon Linux.
How does a 125B model fit in 12 GB? It is a mixture-of-experts model. Each token uses only a small set of experts, so Strata keeps the hot experts on the GPU and keeps the rest in system RAM.
The caveats are real. Q2 is aggressive quantization, and quality drops as you squeeze bits. The README warns your PC can be slow or stop responding for 1–3 minutes while the model loads. Third-party write-ups note it handles one request at a time, and independent benchmarks across hardware are still thin. This is a solo developer's project, not a vendor with a support contract.
Why it matters for your business
Local AI is now a budget line, not a research project. If a mid-range gaming GPU in an office PC can run a 125B model, then private drafting, document summaries and code help can run with zero per-token cost and zero data leaving the building. That matters for client files, HR documents and anything under an NDA.
The OpenAI-compatible endpoint is the real feature. Tools that talk to the OpenAI or Anthropic API can point at localhost instead. That is how you build a fallback: same code, swap the base URL. Your cloud vendor raises prices or has an outage, and the work keeps moving.
Do not put it in front of customers yet. One request at a time is fine for one person at a desk. It is not a customer-facing chatbot backend. Test quality on your own tasks before you trust a Q2 model with anything that ships.
Key takeaways
- Strata runs 125B Qwen3.8-Flash-Next on a 12 GB GPU with 32 GB of RAM
- Reported speed: about 94 tokens/s output on an RTX 5070 at Q2_0
- It exposes OpenAI- and Anthropic-compatible APIs on localhost, so existing tools can switch to it
- It serves one request at a time; use it for internal work, not customer traffic
- Low-bit quantization costs quality; test on your real tasks first
Want AI that keeps working when your vendor changes the terms? We build vendor-agnostic systems that can swap between cloud APIs and a local model by changing one URL. See how we build or run the numbers on your AI spend.
Sources: Strata on GitHub, AlphaSignal.
- #local-inference
- #strata
- #qwen
- #open-source
- #self-hosted-ai
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Poper Blocker adblocker scrapes your AI chats: 2M+ users
Researchers say the Poper Blocker browser extension sends ChatGPT, Claude and Gemini conversations to its server. Remove it and audit your team's extensions.
Read itPerplexity Decisions API: $0.04 per million, open weights too
Perplexity launched a Decisions API on pplx-decider-v1-27b at $0.04 per million input tokens, with Apache 2.0 weights. Price your ticket routing again.
Read it