No Fix Yet for LMCache Flaw That Allows Unauthenticated Remote Code Execution

A critical security issue in LMCache, an open-source caching layer for large language model serving systems like vLLM, remains without a patch. The vulnerability affects multiprocess mode, where the cache operates as a separate server that LLM workers contact through ZeroMQ. An unauthenticated attacker on the network could execute code on the cache server.
LMCache is an open-source caching layer for large language model serving systems such as vLLM. In multiprocess mode, it runs as a separate server that LLM workers reach through ZeroMQ. The critical flaw allows an unauthenticated attacker on the network to execute code on the cache server. No patch is currently available. The situation highlights how auxiliary components in AI serving stacks can become security-critical.
If exploited, this issue could affect organizations running LMCache in multiprocess mode and the users relying on their LLM services. An attacker on the same network might compromise the cache server, potentially disrupting model serving or exposing connected systems. Until a fix arrives, operators may need to limit network exposure, monitor for unusual activity, or avoid affected configurations. The broader lesson is that supporting infrastructure for AI services can carry security risks even when the main models are not directly targeted.