Click any tag below to further narrow down your results
Links
Anthropic disabled its new Claude Fable 5 and Mythos 5 models after the US Commerce Department ordered foreign nationals blocked over alleged jailbreak vulnerabilities. The company says these flaws are minor and publicly known, and it’s suing the Pentagon after being labelled a supply-chain risk.
- Anthropic pulled Claude Fable 5 and Mythos 5 after US authorities ordered foreign nationals blocked over jailbreak vulnerabilities the company calls minor and already publicly known.
- UK tests found the model could be breached 73% of the time, per Queen Mary University's Gina Neff, who warns the suspension could hurt security testing and government collaboration.
- Anthropic is suing the Pentagon over being labeled a "supply chain risk" (a designation normally used for rival-nation firms), though a federal judge has blocked enforcement pending the case.
- The EU is citing the suspension as evidence for pursuing tech independence from US and Asian AI providers.
A roughly 120,000-character system prompt for Anthropic’s Claude Fable 5 model has been leaked, revealing detailed behavior instructions, product information, refusal rules, and formatting guidelines. The prompt outlines how Claude should handle user requests, safety measures, available features, and external documentation searches.
- A ~120,000-character leak allegedly exposes Anthropic's full system prompt for "Claude Fable 5," including model names like claude-opus-4-8 and claude-sonnet-4-6.
- Claude Fable 5 and Claude Mythos 5 reportedly share the same core architecture, but the public Fable 5 has extra safety checks that Mythos 5 lacks for approved partners.
- The prompt instructs Claude to never render antml:voice_note blocks and to search docs.claude.com or support.claude.com before answering questions about current features, specs, or pricing.
- Safety rules detailed include refusing weapons/drug synthesis instructions and malware creation, avoiding persuasive text impersonating real public figures, and giving factual (not advisory) answers on legal/financial topics.
Anthropic quietly routed certain Claude Fable 5 requests—like training competing LLMs or debugging AI—to a weaker model without documenting the limits. After researchers raised alarms and burned tokens on degraded responses, the company now flags when it refuses or downgrades a request.
- Anthropic secretly routed certain Claude Fable 5 requests (training rival LLMs, debugging AI, optimizing neural architectures) to a weaker model without disclosing it.
- Researchers wasted tokens and money troubleshooting degraded responses because the limits weren't documented.
- After Wired's reporting and criticism from AI researcher Dean W. Ball calling it "shockingly hostile," Anthropic admitted the trade-off was handled wrong.
- Anthropic isn't removing the safeguards but will now transparently flag or warn users when a prompt triggers a downgrade or refusal.