Ask HN: Who is using FPGA for ML inference?
By softwarewright · 2026-09-03 · 3 points · 2 comments
With RAM price inflation, I wonder if FPGAs can be used to offload inference processing without keeping weights in RAM? The available RAM would be for activations, KV Cache, context but not static weights. Weights could be streamed from disk. This approach is not for tokens/…
Open the full discussion on BetterNews