Offload LLM KV Caches to Lustre on GKE for 60% Savings
Quick answer
Offload LLM KV caches to Google Cloud Managed Lustre on GKE for 60% fewer GPU hours and 50% TCO savings. A step-by-step guide for Llama-3.3-70B inference.
When your LLM’s KV cache grows bigger than a capybara’s appetite for waterweed, it’s time to think about offloading. Google Cloud’s new recipe pairs GKE with Managed Lustre to serve long-context windows without burning GPU hours. The result? Over 50% TCO savings and a 60% reduction in GPU-hour requirements for Llama-3.3-70B inference on a six-node A3 Mega cluster.
Why Offload to a Parallel Filesystem?
Pooling node-local SSDs sounds good, but it forces your compute cluster to handle data distribution and replication. That’s like asking a capybara to build a dam—possible, but not their strong suit. Instead, Google Cloud’s approach uses Managed Lustre as a dedicated, high-performance external cache tier. It’s a shared parallel filesystem that acts as a centralized attention cache, eliminating host-level capacity limits and networking overhead.
The Benchmark Results
With a 95% cache hit rate, the setup delivers impressive numbers:
- Model: Llama-3.3-70B
- Context: 50k token prompts, 256 token questions, 512 token outputs
- Savings: 60% fewer GPU hours, over 50% TCO reduction
And if you combine Lustre with CPU RAM offload, you get a 40% improvement in Time to First Token (TTFT) and 30% lower end-to-end latency. That’s like a caiman trying to keep up with a capybara in a swamp race—no contest.
Architecture Overview
The solution uses three main components:
- GKE GPU Nodes: Dedicated accelerators for model execution and tensor-parallel operations.
- Managed Lustre: A shared, high-bandwidth parallel filesystem that caches prefilled attention states.
- PVC Evictor: A garbage collector that removes least-recently-used cache chunks to keep storage healthy.
For a deeper dive into how GKE and cloud services compare, check out our Google Cloud review and Vercel review.
Getting Started
Ready to build your own? The guide walks you through creating a GKE cluster, provisioning Lustre storage, deploying vLLM with the llmd-fs-connector, and setting up the PVC Evictor. You’ll need a Hugging Face token and the right IAM permissions. The deployment supports models like Qwen3.5-35B-A3B and Gemma-4-31B-it.
One pro tip: Qwen3.5 needs a block size of 528 to avoid fragmentation, while Gemma 4 works fine with the default 256. And for the evictor, deploy one replica per 72 TB of Lustre capacity to keep things running smoothly.
When you’re done, don’t forget to clean up—those GPU nodes can get pricey. The guide includes a cleanup script that deletes the cluster and PVC, automatically removing the underlying Lustre storage.
For more on related tools, see our Supabase review and Neon Database review.
Original announcement published on Google Cloud.