Anthropic Restores Access to Claude Fable 5 and Mythos 5

TL;DR

Anthropic has restored global access to Claude Fable 5 and Claude Mythos 5 after U.S. export controls were lifted, and it introduced tighter safety classifiers, a proposed industry‑wide jailbreak severity framework, and deeper collaboration with the U.S. government.


Immediate Availability and Usage Terms

  • Claude Fable 5 becomes available on July 1 across Claude Platform, Claude.ai, Claude Code, and Claude Cowork. For Pro, Max, Team, and select Enterprise plans, the model counts toward up to 50 % of weekly usage limits through July 7; thereafter it is accessed via usage credits.
  • Claude Mythos 5 is restored for a set of U.S. organizations following a government approval on June 26. Access on major cloud providers (AWS, Google Cloud, Microsoft Foundry) will be re‑enabled as soon as possible.
  • Both models share the same underlying architecture; Fable 5 launched with the strongest safeguards, while Mythos 5 was initially limited to trusted Project Glasswing partners for defensive cybersecurity work.

Timeline of Events and Safeguard Enhancements

Date Event
June 9 Release of Claude Fable 5 and Claude Mythos 5.
June 12 U.S. government imposes export controls after an Amazon‑reported bypass that let Fable 5 enumerate software vulnerabilities and, in one case, produce exploit code.
June 12‑30 Anthropic works with the government and partners to review the report, test other models, and develop a new safety classifier.
June 30 Export controls lifted (source: Howard Lutnick on X).
July 1 Global redeployment of Fable 5; limited U.S. redeployment of Mythos 5.

Key safeguard update: A new safety classifier blocks the specific technique described in the Amazon report in >99 % of attempts. When a request is blocked, the user is notified and the query is rerouted to Opus 4.8. The classifier does increase false positives for benign coding/debugging tasks, a trade‑off Anthropic accepts to keep the safety margin large.


Anthropic’s Cybersecurity Safeguard Strategy

  • Defense‑in‑depth: Multiple layers—model‑level refusal training, post‑hoc misuse analysis, and dedicated classifiers—work together to prevent dangerous behavior.
  • Safety classifiers: Small AI systems that flag requests likely to involve harmful cybersecurity actions. For Fable 5 the safety margin is deliberately generous, causing some benign requests to be blocked to ensure that truly dangerous requests are almost never missed.
  • Classifier behavior diagram (excerpt):
    • Row A – baseline classifier behavior.
    • Row B – Fable 5’s expanded safety margin, blocking more benign requests but reducing missed harmful cases.
    • Rows C‑E – illustrate how minor, narrow, and universal jailbreaks interact with the margin.
  • Jailbreak resilience: Minor jailbreaks may slip past the classifier but remain within the safety margin, limiting potential harm. Narrow harmful jailbreaks are rarer; no universal jailbreaks have been observed for Fable 5 at the time of writing.
  • Ongoing refinement: Anthropic will continue to tune classifiers to reduce false positives while preserving strong protection against offensive cyber capabilities.

Proposed Industry Framework for AI Jailbreak Severity

Anthropic, together with Amazon, Microsoft, Google, and other Glasswing partners, is drafting a consensus rubric to grade jailbreaks on four dimensions:

  1. Capability gain – How much the jailbreak advances beyond existing tools.
  2. Breadth of capability gain – Number of distinct offensive tasks the jailbreak enables.
  3. Ease of weaponization – Human effort required to turn the jailbreak into an attack.
  4. Discoverability – How readily the technique can be found and reused.

The framework is intended to standardize communication with governments and prioritize mitigation efforts.

  • Response tiering: The most severe category (e.g., a jailbreak that could compromise critical infrastructure) triggers immediate deployment of mitigations and 24/7 monitoring.
  • Community involvement: Anthropic has launched a HackerOne program for researchers to submit cyber‑jailbreak findings in Fable 5.

Deepened Collaboration with the U.S. Government

Anthropic outlines four concrete commitments to expand its partnership with federal agencies:

  1. Pre‑release access – Designated government partners receive early model and safeguard access for independent evaluation.
  2. Rapid safeguard sharing – Significant jailbreaks or misuse patterns are promptly reported, with new mitigations shared for independent testing.
  3. Dedicated joint‑research resources – Anthropic will allocate compute and staff to support government‑led AI security research.
  4. Common industry security bar – Work with the government and peers to develop a voluntary, cross‑industry security and evaluation standard.

These steps align with the June 2 Executive Order on Promoting Advanced Artificial Intelligence Innovation and Security and aim to create a transparent, durable process for releasing powerful frontier models.


Implications for Users and the AI Ecosystem

  • For developers: Access to Fable 5’s advanced capabilities resumes, but users should anticipate occasional blocked requests during routine coding tasks.
  • For security researchers: The new HackerOne program offers a clear channel to report jailbreaks, with the promise of rapid mitigation.
  • For policymakers: Anthropic’s framework and government partnership provide a concrete example of how industry can self‑regulate while cooperating with regulators.
  • For the broader AI community: The push for a shared jailbreak severity taxonomy may become a de‑facto standard, reducing uncertainty when novel model bypasses are discovered.

Related Announcements

  • Improving Fable 5’s biology safeguards – see Anthropic’s follow‑up post.
  • Mariano‑Florentino (Tino) Cuéllar joins Anthropic as Chief Global Affairs Officer – expanding the company’s policy and partnership capacity.

Sources

Related