Anthropic and the Limits of Fable 5: What Cannot Be Said
Anthropic has decided that its Fable 5 model will not talk about certain topics deemed too dangerous. The company launched Fable 5, a model of the so-called "Mythos" class, promising to outperform its previous versions, but with some very clear restrictions. The model does not answer questions about cybersecurity, biology, and chemistry. The concern is that these areas could "empower" malicious actors.
Fable 5 operates on the same baseline as Mythos 5, but while the latter is available only to a select group of trusted cyberdefenders, Fable 5 is accessible to the general public. However, when someone tries to address these sensitive topics, the system redirects the query to the previous model, Claude 4.8 Opus, and warns the user. Anthropic adjusted these guardrails to be stricter than ideal, acknowledging it might frustrate everyday users. But they argue it is a small price to pay to prevent the model from assisting in harmful activities.
An interesting point is that across more than a thousand hours of testing with external red teams, no one managed to bypass these protections. Anthropic is especially concerned about Mythos 5's capability to perform "agentic hacking"—that is, complex cyberattacks—something it does better than previous models. Recent tests showed Mythos Preview performs similarly to OpenAI's GPT-5.5 on security challenges, suggesting there isn't a specific leap for a single model.
Additionally, Anthropic expanded blocks on biology- and chemistry-related queries, fearing that well-equipped malicious actors could use this information for risky biological research. The company understands that this restriction is a double-edged sword. What is useful for cybersecurity professionals and biology researchers can be dangerous in the wrong hands. This puts Anthropic in the delicate position of deciding who is trusted enough to access Mythos 5.
Anthropic plans to expand its Project Glasswing program, in consultation with the US government, to include more cybersecurity professionals and life sciences organizations, allowing them to access the model without biology and chemistry restrictions. For those wanting to use Fable 5, the cost is $10 per million input tokens and $50 per million output tokens, rates significantly higher than OpenAI's GPT-5.5. The company hopes that in the future, Fable 5 can be part of standard subscription plans, once it has sufficient capacity.





Comments (0)
Comments are moderated and if they violate our Terms and Conditions of use, the comment will be deleted. Persistence in violation will result in a ban of your account.