The LMCache vulnerability (CVE-2026-105192) lets anyone who can reach the right network port run code on your LLM server, and there is no patch yet. It only hits setups where LMCache is exposed beyond localhost, but plenty of multi-node vLLM deployments are set up exactly that way. The fix for now is a firewall rule, not an upgrade.
If you run vLLM with LMCache on Kubernetes or across several machines, check your config today. Everyone else can relax a little.
Quick Facts
- CVE-2026-105192, rated CVSS 9.8 (critical), unauthenticated remote code execution.
- Affects LMCache 0.3.9 through 0.5.5, plus the 0.5.6 release candidates.
- Disclosed by JFrog on October 7, 2026, with a public proof-of-concept already online.
- No fixed release exists as of October 8.
- Risk is highest when LMCache’s multiprocess service listens on a routable address.
Those five lines are the whole story in short. The sections below explain what is happening and what to do about it.
What LMCache does and why this bug matters
LMCache is a caching layer that sits next to vLLM and stores the key-value (KV) cache the model builds while reading a prompt. Reusing that cache means repeated or shared prompts are answered faster and cheaper. That is why it shows up in production inference stacks.
The trouble is in its multiprocess mode, where several processes or nodes share one cache service. That service talks over ZeroMQ, by default on port 5555. Servers like this tend to sit deep inside a cluster, where people assume nobody untrusted can reach them.
How the LMCache vulnerability works
According to Cyber Security News, the ZeroMQ transport accepts messages with no authentication at all. No passwords, no CURVE encryption, no message checks.
Worse, the service unpacks incoming data with MessagePack and, for one custom extension type, calls Python’s pickle.loads on it. Pickle can execute code while it loads, so a crafted message runs the attacker’s commands before LMCache ever checks what it received.
JFrog researcher Yuval Moravchick found the flaw, and a working proof-of-concept needs just one message to port 5555. ByteIota reports that the official LMCache container images run as root. That turns a successful hit into control of the whole container.
Are you actually exposed?
Not every LMCache user is at risk. The socket binds to 127.0.0.1 by default, so a single-node setup on localhost can’t be reached from outside. The danger starts when someone passes --host with a routable address, such as 0.0.0.0, to let other nodes connect.
Here is a quick way to sort your own setup.
| Your setup | Exposure | What to do |
|---|---|---|
| Single node, LMCache bound to 127.0.0.1 | Low. Not reachable from other machines. | Keep it that way and watch for a patch. |
Multi-node or Kubernetes with --host 0.0.0.0 |
High. Any host that can reach port 5555 can attack it. | Lock down port 5555 now. |
| LMCache reachable from the public internet | Critical. Assume it can be hit at any time. | Block access immediately and check logs. |
| LMCache not used, or multiprocess mode off | None from this bug. | No action needed. |
Per ByteIota, the official vLLM production-stack guides recommend the routable --host flag for multi-node use. If you copied a guide, you may be in the high-risk row without knowing it.
How to protect your servers before a patch ships
With no fixed version, the goal is simply to keep untrusted traffic away from the service. These steps come from the reports above and JFrog’s advice.
- Block port 5555 for anything untrusted, using firewall rules or a Kubernetes NetworkPolicy.
- Switch back to
127.0.0.1if you don’t truly need multi-node sharing. - Allow only your vLLM worker pods to talk to the LMCache pods.
- Run the process as a non-root user wherever your setup allows it.
Even a tight network rule is a stopgap, not a cure. Anyone already inside your network, or any compromised pod, can still reach the port unless you restrict it by identity as well as by address.
This is the same pattern we saw with the Mistral Vibe vulnerability: AI tooling that runs code or loads data from places it shouldn’t trust. Our guide to the MALFEX npm malware packages covers a related supply-chain angle if you also ship JavaScript.
What to watch for next
JFrog’s longer-term advice is to drop pickle.loads from any path that handles network data, add real transport authentication, and refuse routable binding unless authentication is on. Those are changes only the LMCache maintainers can make.
Also worth knowing: ByteIota says a GitHub user reported six more LMCache security issues on October 6. None have CVEs or maintainer confirmation yet, so treat them as unverified for now. Check the project’s GitHub page and JFrog’s advisory for a fixed release.
Frequently Asked Questions
Is there a patch for CVE-2026-105192?
No. As of October 8, 2026, JFrog and other reports say no fixed LMCache release exists, and the 0.5.6 release candidates still contain the unsafe code.
Does this affect vLLM itself?
The flaw is in LMCache, not vLLM. But if you use LMCache’s multiprocess mode alongside vLLM, your inference servers are the ones at risk.
I only run LMCache on one machine. Do I need to worry?
Probably not about remote attacks, as long as it stays bound to 127.0.0.1. Still, check your launch command to be sure.
How do I check if port 5555 is exposed?
Scan the LMCache host from another machine on your network, or review your Kubernetes services and firewall rules. If anything outside your vLLM workers can connect, close it.
The takeaway
If you’re on a single machine, you can wait for the patch. If you run LMCache across nodes, treat this as urgent: close port 5555 today, and don’t undo the rule until a fixed version is out and tested. A public exploit plus a root-running container is not something to leave for the next sprint.


Leave a Reply