1 link tagged with all of: post-training + frogsgame + curriculum-learning + reward-design + model-crafting
Links
This article evaluates how new frontier models—Claude Fable 5, Opus 4.8, and GPT-5.5—perform on the long-horizon FrogsGame post-training task. Fable 5 stands out by programmatically generating high-quality supervision traces via a backtracking algorithm, combining them with RL, and reliably improving the base model where others still struggle with noisy data and evals.
post-training
frogsgame
model-crafting
curriculum-learning
reward-design