Click any tag below to further narrow down your results
Links
Paper2Agent is an AI agent that reads scientific papers and automatically reproduces their results. It's published in Nature and available as a live demo where you can query it about papers and run workflows through GitHub.
- Automates the extraction and reproduction of experimental results directly from published papers
- Reduces manual work scientists spend reverse-engineering methods and validating findings
- Deployed as an interactive agent you can query in real-time about paper contents and methodology
Anthropic released two versions of Claude 5.1—Fable for general use and Mythos with reduced safeguards for cybersecurity and biology work—claiming superior performance on coding and scientific tasks while cutting prices by 25-45%. The company tested both models extensively for chemical, biological, and cyber risks before deployment.
- Claude Mythos 5.1 designed protein binders with 10x higher affinity than competition winners and 50% hit rates versus the typical 10-15%, suggesting AI can contribute meaningfully to drug discovery.
- Fable 5.1 costs 25% less than Fable 5 for typical workloads and up to 45% less for agent-based tasks, primarily through cheaper cached-read pricing.
- Mythos 5.1 optimized deep learning models by up to 2.5x speed and reduced GPU costs by 30-60% on computational biology tasks—work that normally takes performance engineers weeks.
GPT-5.5 outperforms GPT-5.4 in real-world coding tasks, from debugging and large merge operations to interactive app development. It also serves as a research partner—critiquing manuscripts, proposing analyses, and generating reports on complex datasets—all while running at GPT-5.4 latency through integrated inference optimizations.
- Note: this "GPT-5.5" article appears to be fabricated/speculative, not a real OpenAI announcement — no such model or release exists as of my knowledge.
- As summarized: GPT-5.5 reportedly matches GPT-5.4 latency despite being more capable, via inference optimizations on NVIDIA GB200/GB300 NVL72 hardware.
- As summarized: a coding CEO claims it replicated days of senior-engineer refactoring work and merged a large branch (hundreds of changes) in ~20 minutes.
- As summarized: an immunologist used it to analyze a 62-sample, ~28,000-gene dataset and produce a detailed report in hours instead of months.