Show HN: Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload
By arnav__1 · 2026-07-26 · 20 points · 0 comments
https://github.com/openlake-project/openlake
Hey HN, we’re the developers of OpenLake, an open source storage engine for offloading LLM KV caches from GPU memory into a shared tier of RAM and NVMe. We built OpenLake because KV caches are outgrowing GPU memory. A single 256K token conversation on Gemma 4 31B produces approx…
Open the full discussion on BetterNews