New Relic AI Evaluation: score the transaction, not the call
New Relic AI Evaluation hits public preview in November with guardrail checks, RAG scoring, and cost-to-quality views. Build your golden dataset first.
Most teams that ship an AI feature test it once, in a notebook, and then hope. New Relic wants that test to run forever, attached to production traffic. On October 6 the company announced New Relic AI Evaluation, a new piece of its AI Observability product that scores model and agent output inside the same traces you already use to debug slow checkouts. The tool is not out yet. The work it assumes you have done is something you can start this week.
What actually happened
Per New Relic's announcement on Business Wire, AI Evaluation goes to public preview in November 2026. The feature list:
- Real-time guardrail checks for prompt injection, jailbreaks, PII leaks, toxicity, and bias.
- Quality scores attached to distributed traces, so a bad answer and a slow database call show up in the same view.
- RAG scoring for faithfulness and answer relevancy.
- Cost-to-quality analysis that flags models whose output is not worth their token bill.
- Pre-built evaluators, a side-by-side prompt playground, versioned golden datasets built from traces or synthetic data, and A/B tests across prompts, models, and configs.
Chief Product Officer Brian Emerson's pitch is that point tools grade single LLM calls, while New Relic grades the whole application transaction. New Relic did not publish pricing or a list of supported frameworks.
Why AI evaluation matters for your business
The useful idea here is not the vendor. It is the shape of the loop: the same test set you use before launch keeps scoring live traffic after launch. A support bot that answered refund questions right in September can drift in October because a model version changed under you, or because someone edited the prompt on a Friday. Without a score on live output, you find out from an angry customer.
The cost-to-quality view is the one owners will feel. If a cheaper model scores the same on your real questions, you are paying a premium for nothing. You cannot know that without a scored test set.
What we tell clients, whether or not they buy New Relic:
- Build the golden dataset now. Fifty real questions with known-good answers, pulled from your own tickets or orders. That file is portable to any eval tool.
- Log model name and prompt version on every call. You cannot compare what you did not record.
- Score cost and quality together. A model choice is a price decision.
- Guardrails for PII first. A leaked customer email is a worse day than a wrong answer.
Key takeaways
- New Relic AI Evaluation enters public preview in November 2026; no pricing yet
- It attaches quality and guardrail scores to the same traces as the rest of your app
- Cost-to-quality analysis targets models that cost more than their output is worth
- Your golden dataset of real questions is the asset; it moves with you between tools
- Record model and prompt version on every call so drift has a cause you can find
Shipping an AI feature without a test set? We build the eval harness first — golden dataset, logging, cost and quality scores — then wire it to whatever observability stack you already pay for. See how we build or check what a cheaper model could save you.
Sources: New Relic via Business Wire: New Relic Introduces AI Evaluation.
- #ai-evaluation
- #observability
- #ai-agents
- #new-relic
- #llm-testing
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
GitHub stacked pull requests GA: review AI code in small pieces
GitHub stacked pull requests are now GA on all plans, with the gh stack CLI and merge queue support. Why small PRs matter more when agents write code.
Read itGitHub secret scanning now catches Lovable and Supabase keys
GitHub secret scanning added detectors for Lovable, Supabase, and Pydantic tokens. If you vibe-coded an app into a repo, here's what to check today.
Read it