Skip to content
Rush Commerce
Software & Dev3 min read

LMCache CVE-2026-105192: unpatched RCE in self-hosted LLMs

LMCache CVE-2026-105192 (CVSS 9.8) lets one network message run code on self-hosted LLM servers. No patch yet. Check your bind address and port 5555 today.

If you run your own LLM inference with vLLM and LMCache, check one setting today. JFrog disclosed CVE-2026-105192 on October 7: a CVSS 9.8 bug that lets anyone who can reach one network port run code on the server, with no login. There is no fixed release yet. The fix for now is configuration, and it takes five minutes.

What actually happened

LMCache is an open-source key-value cache that sits next to inference engines like vLLM. It makes repeat prompts cheaper and faster. Per JFrog's advisory, the problem is in LMCache's multiprocess mode:

  • That mode opens a ZeroMQ socket, port 5555 by default, with no authentication.
  • Messages on it are decoded with Python's pickle, which runs code by design. One crafted message is enough.
  • The code runs as the LMCache process user. JFrog says the official container images run as root.
  • The vulnerable code shipped in v0.3.9 and is still in v0.5.5, the v0.5.6 release candidates through rc3, and the dev branch as of October 7.

The 9.8 score applies when the socket is bound to a routable address with --host, for example 0.0.0.0. The default localhost bind is not reachable from other machines. LMCache used only inside a vLLM process does not open the port at all.

This is the second critical RCE in the vLLM stack this year. CVE-2026-22778, a malicious-video-URL bug in vLLM itself, was fixed in vLLM 0.14.1.

Why it matters for your business

Self-hosting moves the patch job to you. Running open-weight models on your own GPUs keeps customer data in-house and cuts per-token cost. It also means nobody patches the inference stack for you. An API vendor would have fixed this already.

Check three things now. Is LMCache in multiprocess mode? Is --host set to anything other than localhost? Can any other machine reach port 5555? If yes to all three, you are exposed. Bind to localhost or a trusted cluster network, and firewall the port. JFrog notes a firewall cuts exposure but does not stop a host that is already allowed to connect.

Stop running inference as root. A root container turns a cache bug into a full server takeover. Run the process as a normal user.

Keep an inventory. You can't patch what you forgot you deployed. Write down every model server, its version, and its open ports.

Key takeaways

  • CVE-2026-105192 is an unauthenticated RCE in LMCache's multiprocess ZeroMQ transport, CVSS 9.8
  • No fixed release as of October 7; v0.3.9 through v0.5.5 and the v0.5.6 candidates are affected
  • Exposure needs multiprocess mode plus a routable --host bind; default localhost is not remotely reachable
  • Bind to localhost, firewall port 5555, and do not run the container as root
  • Separately, confirm vLLM is on 0.14.1 or later for CVE-2026-22778

Running models on your own hardware? We set up self-hosted inference with locked-down network binds, non-root containers, and a written inventory, so a disclosure like this is a checklist, not a fire drill. See how we build it, or send us your current stack for a review.

Sources: JFrog Security Research, Orca Security.

  • #lmcache
  • #vllm
  • #cve-2026-105192
  • #self-hosted-llm
  • #ai-security
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.