SuperNinja Enterprise sells AI employees on a flat annual bill
NinjaTech bundled agents, inference, and reserved GPUs into one fixed annual fee priced in AI employees. Flat pricing is a ceiling, not a discount.
NinjaTech AI launched SuperNinja Enterprise today: an agent platform that runs in the customer's own cloud, on open-weight models, billed as a fixed annual fee with the GPU and inference capacity included in the same contract. It is sold in units of 100, 500, or 1,000 "AI employees." Flat-rate agent pricing is the trend worth tracking here — and the headcount metaphor is the part to be careful with.
What actually happened
Per NinjaTech's announcement and SiliconANGLE's coverage, the platform ships in three packages — 100, 500, or 1,000 agents — with more available on demand. It deploys on Microsoft Azure via Fireworks AI, on AWS, or air-gapped on customer hardware, with single-tenant capacity reserved per customer. Agents are driven from Slack and Microsoft Teams. Customers can still call Anthropic or OpenAI models when a task needs a frontier model.
The pitch is unmetered usage: no per-token charge. CEO Babak Pahlavan's line is "your CFO gets an AI bill that doesn't move," against a market where "most vendors keep AI on a usage meter." NinjaTech claims total cost roughly 10x lower by running open-weight models rather than frontier ones — that is a vendor claim, unaudited, and it depends entirely on which model handles which task. Specific prices were not disclosed. Implementation runs through Infosys, with Optimum HealthcareIT covering healthcare; NinjaTech took investment from the Infosys Innovation Fund in August 2026, so the services partner is also a shareholder. Pilots are offered before annual contracts.
Why flat-rate AI pricing matters for your business
Predictable billing is genuinely valuable. Token metering makes agent work impossible to budget, and every operator who has watched a retry loop bill for six hours knows why. But swap the framing: a fixed fee for reserved single-tenant capacity is not unlimited compute. It is a ceiling you prepaid. The per-token meter told you when you were wasting money. A flat bill hides that signal — a badly built agent that burns capacity on retries now costs you throughput instead of dollars, which is harder to see and slower to diagnose.
"AI employees" as a unit makes it worse, because it invites you to plan in headcount when the real constraint is task throughput on reserved GPUs. Two hundred agents idling and two hundred agents running long-horizon research are the same line item and completely different capacity loads. Before signing anything shaped like this, get three numbers in writing: what one agent is allowed to consume, what happens when you exceed the reserved capacity, and what the renewal looks like once your workload is inside their boundary. Then run the pilot with your ugliest workflow, not your cleanest demo.
Also price the exit. Open weights in your own cloud is a real portability advantage over a hosted meter — you can move the model. Whether you can move the orchestration, the skills, and the integrations is a different question, and it is the one that determines your leverage at renewal.
Key takeaways
- NinjaTech launched SuperNinja Enterprise: agents plus inference plus reserved GPU capacity on one fixed annual fee
- Sold in packages of 100, 500, or 1,000 agents; runs on Azure via Fireworks AI, on AWS, or air-gapped on customer hardware
- Open-weight models by default with optional calls to Anthropic or OpenAI models; the ~10x cost advantage is a vendor claim and prices were not disclosed
- Infosys is the implementation partner and also invested via its innovation fund in August 2026 — factor that into reference checks
- A flat fee on reserved single-tenant capacity is a prepaid ceiling: wasteful agents cost throughput instead of dollars, which hides the problem
- Before signing: per-agent consumption limits, overage behavior, renewal terms — and pilot your messiest workflow, not a demo
Per-seat AI pricing is a metaphor. Your cost per completed task is the number. We build agent workflows on infrastructure you own, instrument cost and throughput per task, and keep orchestration portable so renewal is a choice instead of a hostage negotiation. See how we build vendor-agnostic agent systems, or model the cost per task before you sign an annual deal.
Sources: NinjaTech AI press release, SiliconANGLE.
- #ai-agents
- #pricing
- #open-weight-models
- #vendor-risk
- #unit-economics
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Naive-N0.5-Flash: a 309B MIT model with 1M context
NaiveAI open-weighted a 309B MoE coding model under MIT with native 1M context and zero full-attention layers. What self-hostable long context actually costs.
Read itEmber-1 hits Kimi K3 quality with 40% fewer tokens
Fireworks trained a model to stop over-thinking: same coding benchmarks as Kimi K3, 71% fewer reasoning tokens. Cost per task, not cost per token.
Read it