Click any tag below to further narrow down your results
Links
This article evaluates how new frontier models—Claude Fable 5, Opus 4.8, and GPT-5.5—perform on the long-horizon FrogsGame post-training task. Fable 5 stands out by programmatically generating high-quality supervision traces via a backtracking algorithm, combining them with RL, and reliably improving the base model where others still struggle with noisy data and evals.