More on the topic...
Generating detailed summary...
Failed to generate summary. Please try again.
At 5:23 pm last Friday the White House slapped export controls on Anthropic’s flagship Claude Fable 5 and Mythos 5 models. The move follows a so-called “jailbreak” demo—Fable was asked to “fix this code,” identified security flaws and suggested patches, which in turn exposes those same flaws for exploitation. Rather than disabling the feature, Anthropic’s co-founder Dario Amodei argued it posed no real risk. That response didn’t fly with regulators. Since then Anthropic has been in damage-control mode, flying executives to Washington, negotiating a path to “fix” the jailbreak—something experts say is technically impossible without crippling Fable’s coding skill. As of day seven, the pause remains in effect, and markets are pricing a 50-50 chance of resolution by July 1.
Beyond the pause, the AI world is buzzing with new benchmarks, hardware and policy moves. MidJourney Medical unveiled a full-body scanner that uses no radiation, promises high resolution and low marginal cost, slated for next-year trials. Anthropic floated several policy frameworks on developer obligations, resilience measures and economic redistribution—ideas that feel quaint against an export-control backdrop. On the tech front, OpenAI’s LifeSciBench rolls out 750 expert-authored biotech tasks, while the EvalEval Coalition aims to catalog and verify evaluation metrics across models. Benchmarks like Opus Magnum challenge models to solve Zachtronics puzzles: Claude Fable 5 topped GPT-5.5 and GLM-5.2, but it’s offline. And across corporate labs, GLM-5.2, Grok 4.3, Gemini 3.5 Flash and others jockey on speed, cost and ethical consistency tests. Meanwhile, companies like DeepSeek raise billions—$7.5 B at a $50 B valuation—planting bets on cheap, fast AI agents.
Questions about this article
No questions yet.