Skip to content
Rush Commerce
Software & Dev4 min read

Blitzy's coding agents walk a graph instead of grepping

Blitzy maps codebases into a Neo4j graph so agents traverse real relationships instead of searching. The retrieval layer, not the model, is what fixes context.

Your coding agent is not bad at code. It is bad at finding the right code. That is the argument Blitzy made at the Neo4j Graph Summit on September 28: the bottleneck in autonomous software work is retrieval, not reasoning. Blitzy runs thousands of coding agents in parallel against enterprise codebases, and instead of vector search or a pile of grep calls, it reverse-engineers the repository into a knowledge graph the agents traverse.

What actually happened

Blitzy builds a dynamic graph in Neo4j that maps the relationships already present in software — modules to files, files to classes, classes to functions, functions to the objects and variables they touch. An agent that needs to change a function does not search for the string. It starts at the node and walks the edges to everything that calls it, everything it calls, and everything downstream that breaks if the signature changes.

The constraint that forces this is context size. Blitzy's director of engineering, Neeraj Deshmukh, put the working agent context at roughly 200,000 to 300,000 tokens — call it 20,000 to 30,000 lines of code. An enterprise repository is orders of magnitude larger, so the question is never "can the model hold it" but "which 25,000 lines does it need." As Deshmukh described it, a graph tells you exactly what you will reach because you choose the starting point. There is a quieter benefit too: a malformed Cypher query returns nothing rather than something plausible, so a bad traversal fails loud instead of hallucinating a file.

The guardrails around it are conventional and worth noting. Humans approve an Agent Action Plan before code is written, generated code is tested immediately, and agents review other agents against the spec. Blitzy raised $200 million at a $1.4 billion valuation in May 2026 and reports 84.95% on SWE-Bench Pro as of June — a vendor-reported benchmark figure, not an independent one.

Why the retrieval layer matters for your codebase

You are not running thousands of parallel agents on a monorepo. You are running Claude Code or Codex on a 40,000-line application, and it keeps rewriting a helper that already exists three directories over. Same failure, smaller blast radius.

The transferable lesson is that context engineering beats model selection for this class of problem. Before you upgrade to a more expensive model, make the codebase legible. A few things that cost a day and pay for months: keep a real dependency map in the repo rather than in someone's memory; write a CLAUDE.md or equivalent that names the actual entry points, the modules that must not be touched casually, and where the shared utilities live; keep module boundaries tight enough that a change has a knowable blast radius. If your agent has to read twelve files to learn that all database access goes through one module, write that down instead.

The human-approval step deserves copying verbatim. Blitzy gates on a plan before code, not a diff after it. Reviewing a plan takes two minutes and catches the wrong approach; reviewing a 900-line diff takes an hour and catches typos. Make your agent state what it intends to change and why, and approve that.

Key takeaways

  • Blitzy maps codebases into a Neo4j knowledge graph so agents traverse module, file, function, and variable relationships instead of searching
  • Working agent context runs about 200K–300K tokens — roughly 20,000–30,000 lines — so selection matters more than capacity
  • A malformed graph query returns nothing, which fails loudly instead of returning a plausible wrong file
  • Blitzy raised $200M at a $1.4B valuation in May 2026 and reports 84.95% on SWE-Bench Pro (vendor-reported)
  • Improve retrieval before you upgrade models: dependency map, documented entry points, tight module boundaries
  • Gate on the plan, not the diff — approving intent catches wrong approaches that a 900-line review will not

Agents are only as good as the map you hand them. We restructure codebases so AI-assisted work is predictable: real module boundaries, documented entry points, agent instructions that reflect how the system actually works, and plan-first review gates. See what we have built, or tell us where your coding agent keeps getting lost.

Sources: SiliconANGLE, Blitzy.

  • #coding-agents
  • #knowledge-graph
  • #ai-agents
  • #context-engineering
  • #developer-tools
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.