Click any tag below to further narrow down your results
Links
Researchers at Latch Bio created a benchmark of 100 experimentally grounded tasks to test whether AI agents can make defensible decisions in antibody discovery. Claude Opus 5 performed best at 53% pass rate, but all models frequently failed by answering scientifically adjacent questions rather than the actual problem at hand.
- Opus 5 with Claude Code achieved the highest pass rate at 53%, with xAI and Google models within 3 percentage points, but no model was reliably accurate
- Model rankings shifted depending on the specific competency and decision type—higher cost and token usage didn't correlate with better performance
- Most failures came from scientific framing errors, not computational mistakes: agents performed internally consistent calculations on the wrong question
- The benchmark spans 10 discovery stages from target assessment through engineering and candidate de-risking, with diverse evidence types (binding kinetics, dose-response, sequence data, structural info)
OpenAI released Rosalind Workbench, a unified environment that consolidates fragmented scientific tools and data into one workspace for life science researchers. It's built on GPT-Rosalind, a specialized model for biology, and available through ChatGPT with guided workflows for tasks like protein design, genomics analysis, and molecular docking.
- Solves a real workflow problem: researchers currently juggle data in one system, analysis in another, and experimental records elsewhere, losing context between steps
- Includes specialized viewers and guided tasks (protein design, small-molecule design, genomics, experimental validation) that keep biological questions connected across different tools
- Offers two access modes—Explore mode for general questions and Research mode for complex analysis, with verified organizations able to request access now and individual access coming later
The article argues that AI’s success in molecular design won’t cure most diseases without a deeper understanding of human biology and disease mechanisms. It calls for large-scale, causal biological measurements and better links between cellular data and clinical outcomes before AI can deliver truly transformative medicines.
- AI drug discovery has focused on molecule design, but the real bottleneck is upstream: most diseases lack a validated biological target to hit.
- 90% of clinical trials fail because they target the wrong biological mechanism, not because molecule design is too slow.
- Genuinely new drug targets have dropped from ~100 in 2015 to ~30 in 2024, while 38 targets each now have 50+ drug programs piling onto the same well-known locks.
- Fixing this requires massive causal/perturbational biology data (billions more measured cell states), not just better AI models, since current cell atlases barely scratch the surface.
Anthropic has kicked off an internal drug discovery effort focused on neglected diseases to sharpen its AI tools for biopharma clients. By running its own research alongside partners, the company aims to gather feedback and demonstrate Claude Science’s capabilities.
- Anthropic is running its own internal drug discovery program targeting neglected diseases to battle-test and improve Claude Science before selling it to pharma partners.
- As a public benefit company, Anthropic claims it can choose projects based on patient need rather than commercial potential, unlike typical biotechs.
- The company hasn't said what happens if it finds a promising drug candidate, leaving unclear how it would handle clinical trials.
- This follows a mixed track record for big tech in healthcare, including Alphabet's life sciences unit, Apple's health features, and Amazon's One Medical/PillPack acquisitions.
The article argues that AI will revolutionize drug discovery long before it can streamline clinical development, creating an abundance of candidate molecules but leaving patient trials as the main constraint. As discovery becomes commoditized and more assets target the same biology, real value will hinge on predictive toxicity, clinical efficacy, and strategic trial design.
- Drug candidate pipelines have doubled in the past decade but novel FDA approvals stayed flat at ~50/year, proving clinical development—not discovery—is the real bottleneck.
- Preclinical assets license for tens of millions, but value jumps to hundreds of millions or low-billions post-Phase 2 proof of concept—a premium set to shrink as AI floods the pipeline with candidates.
- Competition per target is already intense (100+ programs on targets like PD-1/GLP-1) and could double or triple by 2030, making individual molecules less rare and pushing investors to demand better translational data and trial design.
- AI excels at data-rich, fast-feedback problems (virtual screening, protein folding) but struggles with messy, high-variability clinical questions (endpoint selection, immune response prediction, adaptive trials)—so real value will shift to whoever masters those still-slow areas.
Anthropic released Claude Fable 5, a Mythos-class model with conservative safeguards that excels at long-context reasoning, software engineering, vision, knowledge work, and life-science research. A restricted-lifted variant, Claude Mythos 5, is available to vetted cyberdefense partners under Project Glasswing. Both are priced at $10 per million input tokens and $50 per million output tokens and lead benchmarks across multiple domains.
- Stripe used Fable 5 to migrate a 50-million-line Ruby codebase in one day, a job that would've taken a team two months
- Fable 5 and Mythos 5 cost $10/million input and $50/million output tokens, under half the previous Mythos Preview price
- Mythos 5 matched or beat expert scientists across all protein-engineering steps in a blind test, yielding nine drug candidates now under investigation
- Cybersecurity safeguards on Fable 5 silently reroute sensitive queries to Opus 4.8 in under 5% of sessions, while Mythos 5 drops those restrictions for vetted cyberdefense partners under Project Glasswing
This post lists nine key quotes from a San Francisco talk by Demis Hassabis and Sebastian Mallaby covering everything from OpenAI’s 50% bankruptcy risk to the need for new “AlphaFold” moments in drug discovery. They debate AGI probabilities, frontier cyber defense access, global AI optimism, and the economic and philosophical challenges of a post-scarcity future.
- Mallaby puts a 50% chance OpenAI goes bankrupt within 18 months
- Hassabis says six more AlphaFold-level breakthroughs are needed to cut drug-delivery timelines from 10 years to months
- Hassabis calls the idea that p(doom) is zero "crazy," acknowledging real existential risk from AI
- Hassabis argues economists, social scientists, and philosophers must join scientists in shaping decisions about how AI's value gets shared
OpenAI trained a new LLM, GPT-Rosalind, on 50 common biological workflows and major public databases to help researchers navigate massive genomic and protein datasets. The model links genotype to phenotype, suggests biological pathways, and prioritizes potential drug targets by leveraging mechanistic understanding.
- OpenAI's GPT-Rosalind is fine-tuned on the 50 most common biological workflows plus major public gene/protein/pathway databases
- It's designed to bridge jargon silos between biology subfields (e.g., helping a plant geneticist interpret neurobiology findings)
- Beyond summarizing data, it suggests biological pathways, links genotype to phenotype, and ranks drug targets to speed up hypothesis generation in drug discovery and synthetic biology
AWS introduced Amazon Bio Discovery, an AI-driven platform that lets researchers run complex drug-design workflows without coding. It provides a library of biological foundation models, an AI agent for workflow setup and analysis, and links to lab partners for synthesis and testing, cutting months of work down to weeks.
- AWS launched Amazon Bio Discovery, a no-code AI platform letting scientists run drug-design workflows using foundation models plus an AI agent, cutting months of research into weeks.
- In a Memorial Sloan Kettering/Twist Bioscience collaboration, nearly 300,000 AI-designed antibodies were narrowed to 100,000 for physical lab testing.
- Bayer, Broad Institute and Voyager Therapeutics are early adopters, and 19 of the top 20 global pharma companies already use AWS cloud services.
- AWS is separately partnering with Boston Consulting Group and Merck on an AI tool to improve clinical trial site selection.
AI model costs and rapid obsolescence are eating into margins—each new generation demands more compute, serves less time, and recoups less revenue. The only way to capture lasting value is by using closed-door models to drive high-value discoveries (like drug design) and own the instruments that generate proprietary data.
- GPT-4.5 only recouped 0.7× its $5.3B cost in five months before being surpassed, versus GPT-3's 3.6× return over 30 months—AI model economics are getting worse, not better, as inference costs now exceed training costs
- Labs are shifting from selling API tokens to using models for proprietary discovery (drugs, materials) whose value outlasts the model itself
- AI drug discovery cuts R&D costs 25-40% and timelines 30-40%, potentially saving $1B per drug against a backdrop of $6.7B average lifetime drug revenue
- Pairing closed models with owned physical instruments (like DeepMind's materials lab or OpenAI/Ginkgo's protein synthesis work) creates proprietary data flywheels competitors can't replicate