GLM-5.3-Flash at 3.3 tok/s
By marcobambini · 2026-08-27 · 1 points · 0 comments
A few months ago, I created the WARP engine (formerly WASTE) to run Kimi K3, the complete 2.78-trillion-parameter model, on macOS. GLM-5.3-Flash shares many architectural similarities with Kimi K3, so I added support for it as well. It requires as little as 5.14 GB of RAM to run…
Open the full discussion on BetterNews