Show HN: NanoRL – RL training for LLMs in ~1,800 lines
By alex000kim · 2026-08-13 · 9 points · 0 comments
https://github.com/alex000kim/nanoRL
The smallest async RL trainer I could write: one loop that runs REINFORCE on CartPole on a laptop and async GRPO on a cluster (e.g. 8xH100 trainer, 8 vLLM workers, ran as a [SkyPilot job group]( https://docs.skypilot.ai/en/latest/examples/job-groups…
Open the full discussion on BetterNews