New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode
By medicis123 · 2026-07-22 · 3 points · 0 comments
Hi All, We are so excited to share the numbers and benchmark reports on our new inference server built specifically to run multi-model agentic workflows on DGX Spark clusters. We ran LlamaBench tests and also our own simulated traffic test on a 2 DGX Spark cluster setup and got…
Open the full discussion on BetterNews