I've tested some local LLMs on prosumer hardware, here are some findings
By felineflock · 2026-08-23 · 1 points · 0 comments
I have been benchmarking local LLMs on a Mac M4 Pro 24 GB RAM using LM Studio. I've tested mostly with 4-bit quantization, both MLX and GGUF, from 4b to 35b models, with speeds of 3 to 40 tokens/second. Results briefly: - fast small model -> extraction/classification -…
Open the full discussion on BetterNews