Skip to content
Rush Commerce
AI & Automation2 min read

Aleph Alpha Kolibri: 78B Apache 2.0 model runs on one B200

Aleph Alpha released Kolibri, a 78B open-weight MoE model under Apache 2.0 with 3.46B active parameters and a 1M context. Here is what it costs to self-host.

Germany's Aleph Alpha shipped Kolibri, a 78-billion-parameter open-weight model under the Apache 2.0 license. It is a mixture-of-experts model with only 3.46B parameters active per token, a 1M-token context window, and a minimum footprint of one B200 or one H200. Aleph Alpha calls it a sovereign model. For a small business, the useful word is "permissive": you can download it, run it on your own hardware, and build on it with no usage fee.

What actually happened

Aleph Alpha announced Kolibri on October 3. The weights are on Hugging Face. Key specs, per the company:

  • Size: 78.1B total parameters, 3.46B active, 384 experts with 6 active per token
  • Context: up to 1,048,576 tokens; the model card recommends 262,144 or less for efficient serving
  • Languages: English and German, with about 21% German pre-training data
  • Training: 20 trillion pre-training tokens on 768 B200 GPUs over 21 days
  • Features: tool calling, reasoning modes from none to high, and training to say "I don't know"

The model card lists minimum hardware as 2× A100 80GB, 2× H100, 1× H200, 1× B200 or 1× B300, with about 78GB of FP8 weights. Quantized variants are available. It serves through vLLM with an OpenAI-compatible API. As of release, the model card says no commercial inference provider hosts it yet.

Aleph Alpha's reported scores include 84.3% on GPQA Diamond and 92.7% on HumanEval+. These are the vendor's own numbers. We have not seen independent results yet.

Why it matters for your business

Small active size means cheap tokens. With 3.46B active parameters, each token costs about as much compute as a small model, while the full 78B gives it more knowledge. That is the right shape for high-volume work like classification, extraction and support triage.

Apache 2.0 is the license you want. No user caps, no acceptable-use rider that a vendor can rewrite next quarter. You can fine-tune it and ship it inside a product.

No hosted API yet is a real cost. Today you run it yourself or you wait. One rented H200 is a reasonable test bed. Run your own prompts against it before you trust the benchmark table.

Key takeaways

  • Kolibri is a 78B mixture-of-experts model with 3.46B active parameters, under Apache 2.0
  • Context goes to 1M tokens; 262K or less is recommended for serving
  • Minimum hardware is one H200, B200 or B300, or two A100 80GB or H100 GPUs
  • It serves through vLLM with an OpenAI-compatible API, so swapping it in is a config change
  • Benchmarks are self-reported: test on your own workload first

Paying per token for work an open model could do? We build vendor-agnostic AI systems that can swap between hosted APIs and self-hosted open weights like Kolibri. See what we build, or run your numbers in our ROI calculator.

Sources: Aleph Alpha, Kolibri-1 model card (Hugging Face).

  • #aleph-alpha
  • #kolibri
  • #open-weight-models
  • #self-hosted-ai
  • #apache-2-0
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.