Show HN: I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp
By federicoTXTS · 2026-07-29 · 1 points · 0 comments
https://github.com/FedericoTs/quantprobe
Democratisation of local AI is key. I've been working on pushing the limits of commercial hardware, squeezing any extra bit possible. My Scientific Agentic AI hareness helped me to reallocate every single bit of it. I rewrote the Kernel, I went down the CUDA rabbit hole until I…
Open the full discussion on BetterNews