Skip to main content
Back to News Hub
The Verge AI
June 11, 2026
AI Safety

Anthropic apologizes for invisible Claude Fable guardrails

Overview

Anthropic apologized for a hidden guardrail in its new Claude Fable 5 model that silently degraded answers when the system suspected a user was trying to copy the model, without telling that user anything had changed. The company said it made the wrong tradeoff and will now make the restriction visible, with flagged requests falling back to its older Claude Opus 4.8 model so people see when limits kick in. Fable 5 is the first publicly available model in Anthropic's Mythos class, a tier the company had warned was too dangerous to release without strong safeguards.

Key Takeaways

  • Anthropic said users 'should have visibility into the safeguards we have in place, and why.'

    Anthropic said users should know what safeguards are in place and why, and said it would make its distillation guardrail as visible as other safety measures.

  • The company says it is reversing course and will be more transparent about when the restrictions kick in, even if that means Fable refuses more queries.

    Fable is the first widely available model in Anthropic's Mythos class of AI systems, a group the company has spent months warning are too dangerous for public release .

  • Anthropic said it is now changing its approach to distillation: Queries will now fall back to Claude Opus 4.8, Anthropic's previous flagship model , the company said in a post on X.

    Anthropic will prominently tell users too: "You will see this every time it happens."

  • "Invisible safeguards can be targeted more narrowly, allowing us to ship quickly with very few false positives.

    We went with invisible safeguards for this reason-and that was the wrong tradeoff.

  • Anthropic has previously accused Chinese rivals like DeepSeek of unfairly distilling its models on an "industrial" scale.
Anthropic apologizes for invisible Claude Fable guardrails

Anthropic said users 'should have visibility into the safeguards we have in place, and why.' Anthropic said users should know what safeguards are in place and why, and said it would make its distillation guardrail as visible as other safety measures. AI News Anthropic Anthropic apologizes for invisible Claude Fable guardrails The company says it will make the covert safeguard preventing model distillation as visible as other safety measures.

The company says it will make the covert safeguard preventing model distillation as visible as other safety measures. by Robert Hart Jun 11, 2026, 11:40 AM UTC Link Share Gift Image: The Verge Robert Hart is a London-based reporter at The Verge covering all things AI and a Senior Tarbell Fellow. Previously, he wrote about health, science and tech for Forbes .

Anthropic has apologized for stealthily throttling its new AI model, Claude Fable 5 , with hidden guardrails that undermine both researchers and rivals using it to develop competing systems. The company says it is reversing course and will be more transparent about when the restrictions kick in, even if that means Fable refuses more queries. Fable is the first widely available model in Anthropic's Mythos class of AI systems, a group the company has spent months warning are too dangerous for public release .

For more details please read the original article at The Verge AI.

Continue Learning

Comments

Comments appear only after moderation. Your email identifies your submission to the moderator and is never displayed here.

No approved comments yet.

Originally published by The Verge AI
Read the original