Tensormesh and AMD Collaborate to Boost AI Inference Efficiency with Advanced KV Cache Optimisation

27 July 2026 | NEWS

The collaboration combines Tensormesh’s KV cache technology with AMD’s GPU virtual memory capabilities to double AI model density, reduce inference latency by up to 7x, and improve GPU utilisation for enterprise AI workloads.

Tensormesh, the company pioneering caching-accelerated inference optimisation for enterprise AI, announced a collaboration with AMD through which the Tensormesh KV cache solution and AMD virtual memory offering will work together to allow more models to be served on fewer GPUs while retaining high KV cache hit rates and throughput even with oversubscribed high-bandwidth memory (HBM). Tensormesh is working with AMD, leveraging its GPU technology and AMD Live Context Virtualisation components, and is tested using Dell servers with 8x AMD/ATI accelerators (MI355) GPUs and Dell storage. Tensormesh integrated LMCache coordinates KV cache management.

For users, running more models on the same set of GPUs means lower costs. LMCache users can now reuse the infrastructure they have built to expand their GPUs' capacity virtually. In addition, customers can reuse KV cache chunks stored for short-term memory virtualisation later for prefix or non-prefix KV cache matching, maximising system efficiency.

The Results

This new approach is far more efficient than building GPUs with more memory, which has led to the industry’s current memory supply crisis and increased the number of GPUs. Enterprises running AI over large document sets can now have a better experience and a lower bill. Testing showed:

  • Near-instant responses at scale: On a 300 GB document set (Kimi-K2.6), time-to-first-token dropped from 3.4 seconds to under half a second, a nearly 7x improvement, thanks to reusing cached KV data from DRAM and NFS instead of recomputing it from scratch.
  • No slowdown as usage grows: Output throughput held steady at ~48 tokens per second regardless of workload size, while unoptimized inference throughput fell by nearly 40% under the same load.
  • Twice the model density: The same hardware doubled the model density.

“This powerful new collaboration builds on AMD’s recent strategic investment in Tensormesh and expands the capabilities of AMD GPUs,” explained Junchen Jiang, Tensormesh CEO and LMCache co-creator. “Together, we’re greatly enhancing the memory that inference engines can access for model weights and KV cache, using all of the memory resources on each node.”

“We recognize Tensormesh and LMCache as KV cache management leaders,” said Anush Elangovan, AMD’s vice president of AI software. “And we’re thrilled to announce AMD’s breakthrough in virtual GPU memory management, amplified by LMCache and Tensormesh.”