Autoheal raises $7.9M for agents that fix your agents
Autoheal's Evaluator scores agent runs against CI failures and incidents; its Healer opens PRs to change prompts and models. The eval loop is the product.
The hard part of running coding agents was never building one. It is knowing which of your twenty agents is quietly making things worse. Autoheal AI raised $7.9 million in seed funding announced September 28, led by Innovation Endeavors, to sell exactly that: a pair of meta-agents that grade your worker agents and then submit pull requests to fix them. The agent evaluation loop is the product, and it is worth understanding whether or not you ever buy it.
What actually happened
Per the company's funding announcement, the round included Emergent Ventures, U&I Ventures, Darkmode Ventures, Batch Ventures and Param Hansa Values.
The architecture is two background agents. An Evaluator scores every worker agent's runs using signals the team already produces — code review comments, CI failures, incident reports. A Healer takes the low scorers and opens pull requests against them, changing model selection, prompts, tools and skills. Those changes are version-controlled in git and require engineer approval, which is the detail that makes it an engineering system rather than a black box that edits itself.
The founding team came out of Harness, which they scaled past $200M ARR, plus Microsoft Azure, ThoughtSpot and AppDynamics. CEO Sid Choudhury frames the thesis as managing agents as code, overseen by continuously learning meta-agents. Named customers include Nomura, AvidXchange and Empiric Earth.
What we are not treating as proven: "self-improving software factory" is positioning, and no independent benchmark of the heal loop's effect on defect rate or token spend is attached to the announcement. Seed-stage claims about enterprise deployments are the company's own.
Why the agent evaluation loop matters for your business
The signal they are using is one you already have. CI failures, review comments and incident tickets are not exotic telemetry. They are sitting in your GitHub and your pager right now, untagged by which agent produced the change. Adding an agent identifier to every commit and every PR is an afternoon of work, and it turns your existing failure data into an agent scorecard. That is 80% of the idea and it costs nothing.
Prompts belong in git or you cannot debug them. The reason Autoheal can open a pull request against an agent is that the agent's prompt, tools and model choice are files. If yours live in a vendor console, a database row or somebody's Notion page, you have no diff, no blame, no rollback and no way to correlate a quality drop with a change. Move them into the repo first.
Most teams cannot answer "is this agent net positive." We see agents ship more code and more rework at the same time, and nobody measures the second half. Pick two numbers before you scale agent usage: percentage of agent PRs merged without human rewrite, and change failure rate on agent-authored commits. If you are not tracking those, adding agents is adding unmeasured risk.
Human approval is the line that keeps this sane. An agent that rewrites another agent's prompt and merges it is a system that can drift overnight with no author. Autoheal gating heals behind engineer approval is the right default, and it is the same default you should enforce in anything you build yourself.
You do not need a platform to start. Log every agent run with its prompt version, model ID and outcome. Join that to CI. Review the worst performers weekly and change one variable at a time. If that loop earns its keep manually, then buy the tool that automates it — with evidence rather than a pitch.
Key takeaways
- Autoheal raised $7.9M seed led by Innovation Endeavors, announced September 28, 2026
- An Evaluator agent scores worker agents on review comments, CI failures and incidents; a Healer agent opens PRs changing prompts, tools and model selection
- Heal changes are version-controlled in git and require engineer approval
- Founders came from Harness (scaled past $200M ARR), Microsoft Azure, ThoughtSpot and AppDynamics; customers named include Nomura and AvidXchange
- No independent benchmark of the heal loop's effect on defect rate or token spend was published with the announcement
- Build the cheap version first: tag commits with the agent that produced them, keep prompts in git, and track merge-without-rewrite rate and change failure rate
Can you tell which of your AI agents is actually earning its tokens? We instrument agent workflows so every run is tied to a prompt version, a model ID and a real outcome — then cut the ones that lose. See how we measure automation, or run the numbers on your own agent spend.
Sources: Autoheal funding announcement, GlobeNewswire, SiliconANGLE.
- #ai-agents
- #evals
- #devops
- #funding
- #coding-agents
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Naive-N0.5-Flash: a 309B MIT model with 1M context
NaiveAI open-weighted a 309B MoE coding model under MIT with native 1M context and zero full-attention layers. What self-hostable long context actually costs.
Read itEmber-1 hits Kimi K3 quality with 40% fewer tokens
Fireworks trained a model to stop over-thinking: same coding benchmarks as Kimi K3, 71% fewer reasoning tokens. Cost per task, not cost per token.
Read it