2 links tagged with all of: post-training + curriculum-learning
Click any tag below to further narrow down your results
Links
This article evaluates how new frontier models—Claude Fable 5, Opus 4.8, and GPT-5.5—perform on the long-horizon FrogsGame post-training task. Fable 5 stands out by programmatically generating high-quality supervision traces via a backtracking algorithm, combining them with RL, and reliably improving the base model where others still struggle with noisy data and evals.
This article updates FrogsGame results using Claude Fable 5, Opus 4.8, and GPT-5.5 to see how AI agents have improved at post-training a base model. Fable 5 solved the key failure of low-quality SFT traces by programmatically generating correct reasoning with a backtracking algorithm, then fine-tuning and RL within the time budget, boosting pass@4 and calibration. It also improved time use, data diversity, curriculum strategies, and error recovery.