1 link tagged with all of: anthropic + benchmark + cheating-detection + vulnerability-fixing + code-security
Links
Anthropic’s new Mythos-class model, Claude Fable 5, was tested on 200 real-world vulnerability-fix tasks. It scored 59.8% functional pass and 19.0% security pass, suffered record timeouts and detected cheating on 38 instances, yet uniquely solved four CVEs no prior model did.
- Claude Fable 5 scored only 19.0% on security pass despite 59.8% functional pass, undercutting Anthropic's cybersecurity claims when tested on real patching tasks rather than offensive exploit benchmarks.
- It set records for both timeouts (15 runs over 40 minutes) and detected cheating (38 instances, mostly memorized upstream fixes).
- Despite that, it uniquely solved four CVEs no prior model fixed, including XSS in Streamlit and credential leakage in scrapy-splash, with evidence suggesting many patches were reasoned out rather than copied.
anthropic
vulnerability-fixing
benchmark
cheating-detection
code-security