Perplexity Decisions API: $0.04 per million, open weights too
Perplexity launched a Decisions API on pplx-decider-v1-27b at $0.04 per million input tokens, with Apache 2.0 weights. Price your ticket routing again.
Perplexity launched the Perplexity Decisions API this week. It runs on pplx-decider-v1-27b, a decision model that returns probabilities over answers you define instead of writing text. The price is $0.04 per million input tokens, and output is free. Perplexity also put the weights on Hugging Face under Apache 2.0. So you can rent it by the token or run it yourself. That combination changes the math on every classification call in your stack.
What actually happened
From the Decisions API docs:
- Endpoint:
POST https://api.perplexity.ai/v1/decisions. Input can be text, JSON, arrays or images (PNG, JPEG, WebP as base64). - Three question types: yes/no (a probability from 0 to 1), choice (top option plus the full distribution) and score (a probability-weighted level).
- Batching: 1 to 128 named questions per request.
- Price: $0.04 per million input tokens, free output, no per-request fee.
- Rate limit: 10 requests per second per organization on every plan.
From the model card:
- Base: Qwen3.8-27B, fine-tuned for decisions. Apache 2.0.
- Accuracy: 85.71% averaged over 11 benchmarks, against 84.51% for TypeSafe's Jev and 74.76% for the base Qwen model. It loses to Jev on the hard JevBench public set (70.30% vs 73.27%).
- Hardware to self-host: about 49 GB of GPU memory plus working space.
This is the third decision model we have covered in two days, after Strands Decider 2B and OpenAI's Decisions API. The category is getting crowded, and crowded categories get cheap.
Why the Perplexity Decisions API matters for your business
The price of a label just hit the floor. A support ticket of 500 tokens costs $0.00002 to route. Ten thousand tickets a month is twenty cents. If you still send "refund, sales or spam?" to a frontier chat model and parse the reply, you are paying for prose you throw away.
Hosted and self-hosted, same weights. This is the part we like. Start on the API. If volume grows, or a client will not let customer data leave their network, move the same model to your own GPU. Your prompts and thresholds do not change. That is an exit plan built into the vendor choice.
Watch the rate limit. Ten requests per second is fine for a shop's inbox. It is not fine for scoring a full product catalog in one burst. Use the 128-question batching, or self-host for bulk jobs.
Benchmarks are not your data. 85.71% on public tests still means mistakes. Label 200 of your own tickets, run them through, and set the confidence threshold where wrong answers stop costing you money.
Key takeaways
- Perplexity's Decisions API returns probabilities, not text, on pplx-decider-v1-27b
- $0.04 per million input tokens, output free, 10 requests per second per org
- Apache 2.0 weights on Hugging Face; self-hosting needs about 49 GB of GPU memory
- 85.71% across 11 benchmarks, slightly ahead of Jev, behind it on the hardest set
- Test on your own labeled data and gate actions on the confidence score
Still paying chat-model prices to sort tickets? We build agents that send simple decisions to a cheap decision model and keep the expensive model for the hard cases, on an API or on hardware you own. Estimate the savings or see how we build.
Sources: Perplexity Decisions API docs, Hugging Face: pplx-decider-v1-27b.
- #perplexity
- #decision-models
- #decisions-api
- #ai-agents
- #open-source
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Strands Decider 2B: AWS open-sources a 115ms decision model
AWS Strands Labs open-sourced Strands Decider 2B, a decision model that picks from fixed options with a confidence score. Use it to route tickets and gate tool calls.
Read itMeta Muse Gadgets: open ESP32 and Linux SDK for agent hardware
Meta open-sourced Muse Gadgets, an Apache 2.0 ESP32 firmware and Linux SDK that puts its Muse AI agent on cheap hardware. What it means for shop-floor devices.
Read it