1 link tagged with all of: reinforcement-learning + curriculum-learning + post-training + frogs-game
Click any tag below to further narrow down your results
Links
This article updates FrogsGame results using Claude Fable 5, Opus 4.8, and GPT-5.5 to see how AI agents have improved at post-training a base model. Fable 5 solved the key failure of low-quality SFT traces by programmatically generating correct reasoning with a backtracking algorithm, then fine-tuning and RL within the time budget, boosting pass@4 and calibration. It also improved time use, data diversity, curriculum strategies, and error recovery.