Chunked KL loss, running Knowledge Distillation locally in less <6GB VRAM
By ikergarcia1996 · 2026-08-11 · 2 points · 1 comments
https://github.com/CompactifAI/Full-Chunked-KL-Loss
By ikergarcia1996 · 2 points · 1 comments · on Hacker News, read on BetterNews.
Open the full discussion on BetterNews