An basic idea in scaling RL: Can we allocate more compute to the harder problems?
We did this: If your GRPO group has all wrong completions, sample more with probability P (~0.9) -- in search of more GRPO batches with nonzero gradient. It works!
Called "Never Give Up"
Is RL actually making your LLM better? Gains from RL are mostly on easy questions🤯 We're calling this the Matthew Effect for RL on LLMs. We then leverage async...