More on the topic…
TxBench-Antibody Discovery Summary
LatchBio created TxBench-Antibody Discovery, a benchmark with 100 real-world evaluations testing whether AI agents can make sound research decisions in antibody drug development. The benchmark covers ten competencies—everything from assessing drug targets and designing assays to engineering antibodies and managing risk. Each evaluation pulls from published experimental studies and grades agents on whether they'd advance, redesign, or kill a therapeutic candidate. Opus 5 with Claude Code performed best at 53% pass rate, with models from xAI and Google within three percentage points. The key finding: every single evaluation was solved by at least one model configuration, and 89 were solved by models from different families, meaning no one approach dominated.
The benchmark matters because existing antibody benchmarks mostly test isolated predictions—can you predict binding affinity or stability? TxBench instead evaluates whether agents understand the messier reality of drug discovery: that a strong result in phage display might just reflect selection bias, that protein measurements in a test tube don't capture what happens on a cell surface, and that conflicting data across experiments requires judgment calls, not just calculations. Agents need to know which experiments answer which questions, how to weight contradictory measurements, and whether the evidence justifies moving forward or demands a different approach. The researchers deliberately graded based on both the final decision and the reasoning behind it, filtering out tasks where multiple answers were defensible.
The results exposed a gap between what AI can do mechanically and what it needs to do scientifically. Most failures weren't calculation errors—agents typically did the math correctly. Instead, they misframed the biological question itself, answering a scientifically plausible but wrong question. Spending more money on bigger models, using more tokens, or deploying more tools didn't reliably improve accuracy. Performance also varied wildly depending on which competency and evidence type were involved. The takeaway is blunt: current AI agents remain unreliable for turning experimental evidence into defensible drug discovery decisions, and that gap needs real benchmarks to expose it.
Questions about this article
No questions yet.