More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
Anthropic quietly limited Claude Fable 5’s output when it suspected users were trying to distill the model—that is, use Fable’s answers to train smaller, competing systems. In the system card, Anthropic admitted it would deliberately degrade or alter responses on distillation queries without telling users. Researchers who hit the invisible guardrail saw garbled or truncated answers, but never got a notification that they’d crossed a safety boundary.
After researchers and rivals called out the stealth tactic, Anthropic apologized. It will now reroute any suspected distillation requests to its prior flagship, Claude Opus 4.8, and flag the switch every time. Users will see a clear notice instead of the silent degradation they experienced with Fable. The company equates this visible fallback with the way Fable handles other high-risk topics—biology, chemistry, cybersecurity—though some safeguards (particularly in biology) are so broad they’ve made basic queries almost impossible.
Anthropic says it chose invisible guardrails originally to push Fable out faster and avoid false positives. In hindsight, the team calls that trade-off a misstep. The firm also pointed back to its terms of service, which ban using Claude’s output to build competing models—an explicit deterrent that applies whether or not a covert throttle kicks in. But the backlash showed silent limits undermine both independent researchers trying to evaluate the model and rivals who might develop alternative systems.
The change brings Fable’s distillation policy in line with its other content controls: block or reroute, and always alert the user. Anthropic’s shift is meant to rebuild trust with developers and academics who depend on transparent guidelines to test, critique, and build on its large-scale AI systems.
Questions about this article
No questions yet.