Skip to content
Rush Commerce
Software & Dev3 min read

JetBrains Mellum 2.1: open coding model for self-hosted agents

JetBrains Mellum 2.1 is a 12B open coding model (2.5B active) under Apache 2.0, built for fast sub-agents on your own GPUs. Here's where it fits.

JetBrains just released Mellum 2.1, an open coding model built for one job: running the small, fast steps inside a coding agent on your own hardware. It is a 12B mixture-of-experts model with only 2.5B parameters active per token, and it ships under Apache 2.0. That license matters more than the benchmark charts. You can run it, fine-tune it, and ship it inside a product with no usage fee.

What actually happened

Per the JetBrains AI blog:

  • Same architecture as Mellum 2, which JetBrains open-sourced in June. 12B total parameters, 2.5B active.
  • The change is post-training. Reinforcement learning became the main training stage, with millions of sandboxed runs across thousands of environments: software engineering, tool use, competitive programming, math, and science.
  • Biggest gain is agentic coding. JetBrains compares it with Mellum 2, Qwen3.5-9B, and Gemma 4 E4B on LiveCodeBench, SWE-Verified, BFCL, and others. The scores are in charts, so check them against your own tasks before you trust them.
  • Speed is the pitch. On one H200, JetBrains says it serves almost twice as many tokens as Qwen3.5-9B under heavy load. A multi-token prediction head makes single requests about 1.6x faster.
  • Weights are on Hugging Face now. GGUF builds for llama.cpp, Ollama, and LM Studio, and the speculative-decoding head for vLLM, are listed as "coming soon."

JetBrains positions it as a worker, not a lead: finding the root cause of a failing test, drafting a fix, checking it.

Why it matters for your business

Most agent steps don't need a frontier model. A coding agent spends many calls on small jobs: read this file, run this test, summarize this stack trace. Paying frontier-model prices for each call adds up fast. A good split is a frontier model that plans and a small open model that does the grunt work. Mellum 2.1 is built for that second role.

Your code stays in your building. If client contracts or compliance rules keep source code off third-party APIs, a 2.5B-active model on one GPU is a realistic way to put agents in your pipeline anyway.

Apache 2.0 is the real feature. No per-seat fee, no acceptable-use surprise, no vendor sunset. If the model works for you this quarter, it still works next year.

Wait for the GGUF builds if you're on a laptop. Today it's a Hugging Face download aimed at server inference. The Ollama and llama.cpp builds will make it practical on a dev machine.

Key takeaways

  • Mellum 2.1 is a 12B MoE (2.5B active) open coding model under Apache 2.0
  • RL-heavy post-training; JetBrains reports the largest gain in agentic coding
  • Claims about 2x Qwen3.5-9B throughput under load on one H200, and about 1.6x single-request speed with MTP
  • Weights on Hugging Face now; GGUF/Ollama and vLLM MTP support "coming soon"
  • Best fit: cheap, private sub-agents under a frontier planner

Paying frontier prices for every agent step? We build agent pipelines that send planning to a frontier model and routine steps to open models you host, so you pay less and your code stays in-house. Run the numbers on your workload, or see how we build it.

Sources: JetBrains AI Blog, JetBrains on Hugging Face.

  • #jetbrains
  • #mellum
  • #open-weights
  • #coding-agents
  • #self-hosted-llm
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.