RL fries the models' brains on complex domains
By threevox · 2026-08-28 · 2 points · 0 comments
https://jagilley.github.io/rl-is-not-enough.html
By threevox · 2 points · 0 comments · on Hacker News, read on BetterNews.
Open the full discussion on BetterNews