Skip to content
Rush Commerce
Tools & Teardowns4 min read

Naive-N0.5-Flash: a 309B MIT model with 1M context

NaiveAI open-weighted a 309B MoE coding model under MIT with native 1M context and zero full-attention layers. What self-hostable long context actually costs.

A permissive license and a million tokens of context used to be a pick-one situation. NaiveAI published Naive-N0.5-Flash to Hugging Face under MIT — a 309B-parameter mixture of experts with native 1M-token context and, notably, no full-attention layers anywhere in the stack. For anyone weighing whether a self-hosted open-weights coding model can handle a whole repository, the architecture is the part worth reading.

What actually happened

The model card puts 309B total parameters against 15.5B active per token, which is the ratio that decides whether you can serve it at all. Weights are roughly 315GB in FP8, and inference needs FP8-capable Nvidia GPUs plus headroom beyond the weights themselves.

The attention design is the unusual claim. The 48-layer stack is 39 sliding-window attention layers and 9 DeepSeek Sparse Attention layers — no dense full attention at any depth. The sliding window is 128 tokens; the sparse layers select a top-2,048 tokens for backbone attention, with 4 KV groups for grouped-query attention. Full attention costs scale with the square of context length, which is why most long-context models get expensive exactly when you use the context. Drop it entirely and the cost curve flattens.

Training ran on top of Xiaomi's MiMo-V2.5 base for 3.25T tokens in three stages: 50B for indexer warmup, 3T on sparse attention training, 200B of learning-rate decay. NaiveAI's throughput figures are ~50 tokens/s per user in Standard mode and up to ~2,000 tokens/s in an Ultrafast mode, via an inference stack they call NaiveRT.

Two things we are not repeating as fact. NaiveAI says the model was developed under an "AI-centered R&D" process where models write the code and run the experiments — an interesting claim, unverifiable from the outside. And the throughput numbers are vendor figures on vendor hardware with no independent benchmark attached.

Why an MIT-licensed long-context model matters for your business

MIT is the whole story. Not "open weights" with a use-case carve-out, an acceptable-use appendix, or a revenue threshold that flips you into a commercial license. MIT means you fine-tune it, ship it inside a product, and owe nobody a conversation. That is rarer at this size than the phrase "open source AI" suggests.

315GB is not a laptop, and it is not a moonshot either. This is a rented multi-GPU box, not a data center. The real math is utilization: a self-hosted model you keep busy beats per-token API pricing, and one that idles does not. Run the arithmetic on your actual duty cycle before anyone orders hardware.

No full attention changes what long context is good for. Sliding-window plus sparse layers keeps long-context cost sane, but 128-token windows with 2,048-token sparse selection is a different retrieval profile than dense attention. Feed it your real repository and check whether it actually connects a function to its caller forty files away, rather than trusting a context-length number on a spec sheet.

Self-hosting buys you the thing no API sells. Weights on your disk cannot be deprecated, repriced, or have their safety behavior changed under you on a Tuesday. We have had to migrate clients off models that vanished with thirty days notice. An MIT checkpoint is a floor under your product.

Test it as a fallback before you test it as a primary. The cheapest way to find out whether an open model is good enough is to route your least critical workload through it for a week and log the failures. That is a config change in any system built with the model behind an interface.

Key takeaways

  • NaiveAI released Naive-N0.5-Flash on Hugging Face under the MIT license: 309B total parameters, 15.5B active, native 1M-token context
  • No full-attention layers: 39 sliding-window layers (128-token window) plus 9 DeepSeek Sparse Attention layers selecting a top 2,048 tokens
  • Built on Xiaomi's MiMo-V2.5 and trained a further 3.25T tokens across three stages
  • Weights are roughly 315GB in FP8 and require FP8-capable Nvidia GPUs with additional memory for inference
  • Vendor throughput figures: ~50 tokens/s per user standard, up to ~2,000 tokens/s in Ultrafast mode — not independently benchmarked
  • MIT with no use restrictions is the differentiator; self-hosted weights cannot be repriced or deprecated out from under your product
  • Validate long-context retrieval on your own repository before trusting the 1M figure

Is your product one vendor email away from a rewrite? We build AI systems with the model behind an interface and an open-weights fallback wired in from day one, so a price change or a deprecation notice is a config edit instead of a quarter of work. See how we keep AI systems portable, or tell us which model you are locked into.

Sources: Naive-N0.5-Flash model card, Hugging Face.

  • #open-weights
  • #llm
  • #long-context
  • #coding-agents
  • #self-hosting
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.