Show HN: RunNburn – Run a 295B Moe from a 98GB GGUF on a 64GB RAM Desktop
By coderredlab · 2026-07-30 · 2 points · 0 comments
https://github.com/coderredlab/runNburn
runNburn is an Apache-2.0 Rust inference engine for quantized GGUF models that are too big for your fast memory. The core idea: weights stay file-backed (mmap), host residency stays under an explicit byte budget (--ram-budget), and GPU caches are sized from detected free/to…
Open the full discussion on BetterNews