1 link tagged with all of: data-analysis + analytical-evals + anthropic + fable-5 + benchmarks
Links
Hex built a suite of analytical evals to test data-analysis models and found Claude Fable 5 outperforms its Opus 4.x predecessors by 10–15%, nailing both semantically modeled and raw-data tasks with fewer mistakes. They’ve also designed a tougher “Frontier” benchmark for long-horizon, open-ended scenarios, where Fable 5’s careful assumptions and cross-checks boost its pass rate to around 58%.
- Claude Fable 5 beats Opus 4.7 by 10-15 points on Hex's core benchmarks, scoring 93%+ on Analytical Hard/Semantically Modeled tests and 65% on Semantically Unmodeled tasks, versus prior Opus versions' single-digit gains
- Fable's advantage comes from following a "golden workflow" (starting in the semantic layer, cross-checking raw SQL) and transparently stating assumptions, which lets it catch errors like a cents-for-dollars mistake that Opus misses
- On Hex's new "Frontier" benchmark for long-horizon, open-ended tasks, Fable at Max Effort hits 58% pass rate, notably outperforming other setups
anthropic
fable-5
analytical-evals
data-analysis
benchmarks