1 link tagged with all of: drug-discovery + ai-benchmarking + biologics
Click any tag below to further narrow down your results
Links
Researchers at Latch Bio created a benchmark of 100 experimentally grounded tasks to test whether AI agents can make defensible decisions in antibody discovery. Claude Opus 5 performed best at 53% pass rate, but all models frequently failed by answering scientifically adjacent questions rather than the actual problem at hand.
- Opus 5 with Claude Code achieved the highest pass rate at 53%, with xAI and Google models within 3 percentage points, but no model was reliably accurate
- Model rankings shifted depending on the specific competency and decision type—higher cost and token usage didn't correlate with better performance
- Most failures came from scientific framing errors, not computational mistakes: agents performed internally consistent calculations on the wrong question
- The benchmark spans 10 discovery stages from target assessment through engineering and candidate de-risking, with diverse evidence types (binding kinetics, dose-response, sequence data, structural info)