Skip to content
Rush Commerce
Software & Dev3 min read

GLM-5.3 found 2,436 bugs in code you already run

Z.ai's GLM-5.3 tops the open-weights coding benchmarks and found 2,436 vulnerabilities in open-source projects — 1,097 critical. Your patch cadence is the story.

Z.ai shipped GLM-5.3 today and the coding numbers are what everyone will quote. Ignore them for a second. The number that should change how you run next week is this: since GLM-5.2, the company says the model found 2,436 vulnerabilities across 269 open-source projects, and 1,097 of them are rated critical or high. Those projects are your dependency tree. This is an open-weights coding model, and the bug-finding came along for the ride.

What actually happened

Per Z.ai's release, GLM-5.3 uses the same base model as GLM-5.2. Every gain came from scaled-up post-training — no new pretraining run. The coding jumps are steep: Terminal-Bench 3.0 went from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents' Last Exam (CLI) from 23.8 to 28.5, per benchmark tables reported by MarkTechPost.

The security side moved further. CyberGym went from 77.2% to 84.5%; ExploitBench from 24.4% to 54.4%. Z.ai's own framing is the uncomfortable part: it says the model began reasoning across multiple stages of exploitation and forming coherent full exploit chains — a capability it describes as arriving faster than its training intended, per Unite.AI.

Two details worth holding onto. First, the findings are tracked in a public disclosure ledger at cvd.z.ai — 53 CVEs assigned at launch, 2,383 still under embargo. Second, the weights aren't out. Z.ai says roughly two weeks, pending safety evaluation and hardening. Today it's API and GLM Coding Plan only.

Why it matters for your business

That embargo queue is the whole story. Twenty-three hundred findings are sitting in coordinated disclosure, and they become public CVEs on a schedule you don't set. The projects named — Linux kernel, WebKit, FreeBSD — are not exotic. They're the boring layers under your VPS, your container base image, and your customers' browsers.

Your patch cadence was almost certainly tuned for a world where finding bugs was slow, expensive, and mostly human. That world ended quietly. "It's been stable for twenty years" was never evidence of safety; it was evidence that nobody had looked hard. Now something looks hard, continuously, for the price of an API call.

The practical move is unglamorous and we keep repeating it because it keeps working: know exactly what you run (a real dependency inventory, not a vibe), automate the path from advisory to deployed patch so it takes hours instead of a sprint, and stop triaging by CVSS score instead of exploitation. The volume is going up. The only thing you control is how fast you can move.

The weights delay is worth noting too. If your cost plan depends on self-hosting an open model, remember that "open" arrived with a release date the vendor controls.

Key takeaways

  • GLM-5.3 shipped August 14 on the same base model as 5.2 — all gains from scaled post-training, not new pretraining
  • Coding: Terminal-Bench 3.0 4.6 → 28.3, DeepSWE v1.1 46.2 → 66.9. Security: CyberGym 77.2% → 84.5%, ExploitBench 24.4% → 54.4%
  • Z.ai reports 2,436 vulnerabilities found across 269 open-source projects since 5.2, 1,097 critical or high, with 2,383 still under embargo
  • Weights are held back roughly two weeks for safety hardening — "open-weights" is still a vendor-controlled release date
  • Assume your dependency tree is being audited at machine speed. Inventory what you run and automate the patch path.

Is your patch path a person remembering to check? We build the dependency inventory and the automated update pipeline that turns a new CVE into a deploy, not a fire drill. See how we build systems you own or tell us what's running in production.

Sources: Z.ai, MarkTechPost, Unite.AI.

  • #glm
  • #open-weights
  • #vulnerabilities
  • #ai-security
  • #patching
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.