Skip to content
Rush Commerce
AI & Automation3 min read

Reflection Beam: 501B open weights, Apache 2.0, 23B active

Reflection's Beam open-weight model claims GLM-5.2-level reasoning at 3-4x less inference compute. Apache 2.0 weights land later in October. What to test.

Reflection AI announced Beam today, a 501-billion-parameter open-weight model it says matches the best Chinese open models at a fraction of the compute. Only 23 billion parameters fire per token, the license is Apache 2.0, and the weights ship later this month. For anyone who wants a self-hosted model that a procurement team will sign off on, the shortlist just got one name longer.

What actually happened

Reflection introduced Beam on October 5. The specs from its own post:

  • Architecture: sparse mixture-of-experts, 501B total parameters, 23B active, with interleaved local and global attention.
  • Context: 256K tokens in pretraining, extended to 1M tokens in midtraining. Text only.
  • Training: 23.8 trillion tokens in under four weeks on 6,144 Nvidia GB300 GPUs, then reinforcement learning across nearly a million tasks.
  • Reported scores: 80.9 on SWE-Bench Verified, 65.5 on SWE-Bench Pro v1, 80.1 on Terminal-Bench v2.1, 90.5 on GPQA Diamond.
  • Efficiency claim: comparable reasoning scores to Z.ai's GLM-5.2 while using 3–4x less inference compute.

The catch: you can't download it yet. Beam is in final red-teaming. Reflection says weights, a technical report, and a model card arrive "later this month" under Apache 2.0, with early access signup at its platform. TechCrunch frames it as a US-built answer to Chinese open models, backed by more than $7 billion in GB300 compute deals with SpaceX and Nebius.

Why an open-weight model like Beam matters for your business

Active parameters set your bill. A 23B-active MoE costs roughly what a 23B dense model costs per token to run, but you still need memory for all 501B weights. That means a multi-GPU node or a hosted provider, not a gaming PC. If you self-host, price the hardware for the total size and the throughput for the active size.

Origin is now a buying criterion. Enterprise buyers and public-sector customers ask where a model came from. A US-trained, Apache 2.0 model removes a question that GLM or Qwen invites, whether or not that question is fair. It joins Aleph Alpha's Kolibri on the short list of permissive Western options.

Every score above is self-reported. None of it is independently checked yet, and the weights aren't public. When they land, run your own eval set: your tickets, your code, your documents. Measure cost per finished task, not tokens.

Key takeaways

  • Beam is a 501B mixture-of-experts model with 23B active parameters and up to 1M tokens of context
  • Reflection reports 80.9 on SWE-Bench Verified and GLM-5.2-level reasoning at 3–4x less inference compute
  • Weights ship later in October under Apache 2.0; today there is only an early-access signup
  • You pay memory for 501B and compute for 23B, so plan hosting for both
  • All benchmarks are vendor-reported; test on your own workload before you switch

Open weights only help if your stack can swap models without a rewrite. We build vendor-agnostic AI systems with an eval harness wired in, so a new model like Beam gets scored on your real work the week it ships. See how we build it or run the numbers on self-hosting.

Sources: Reflection, TechCrunch.

  • #open-weight-models
  • #reflection-ai
  • #beam
  • #self-hosted-ai
  • #ai-costs
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.