Evaluating Claude's Bioinformatics Research Capabilities with BioMysteryBench

Anthropic has developed BioMysteryBench, a new bioinformatics benchmark designed to evaluate whether AI models can solve complex, open-ended research problems using real-world datasets. The evaluation reveals that Claude's scientific capabilities are improving rapidly across generations, with the latest models performing on par with human experts and solving several problems that a panel of human experts could not.

The Challenges of Evaluating Bioinformatics Research

Evaluating scientific AI is difficult because traditional benchmarks often fail to capture the messy, open-ended nature of real research. Anthropic identifies three primary obstacles to creating a canonical science benchmark:

  • Methodological Diversity: In biology, there are often multiple "right" ways to approach a problem based on a researcher's background and available resources.
  • Subjectivity and Noise: Biological datasets are often noisy, meaning small differences in subjective research decisions can lead to entirely different conclusions.
  • Human Knowledge Gaps: The most impactful research tasks are those humans have not yet solved, making it impossible to use human performance as the sole ground truth.

BioMysteryBench: A Verifiable Framework for Science

BioMysteryBench consists of 99 questions written by domain experts across various bioinformatics fields. To overcome the evaluation challenges mentioned above, the benchmark employs four key properties:

  1. Method-Agnostic Grading: Claude is graded on the final answer rather than the specific analytical path taken, allowing for research creativity.
  2. Objective Ground Truth: Answers are derived from controllable properties of the data or validated metadata (e.g., PCR assays) rather than subjective scientific conclusions.
  3. Superhuman Question Generation: Because answers are based on data properties, the benchmark includes problems that are objectively solvable but were too difficult for human experts to solve.
  4. Tool-Enabled Environment: Claude is placed in a container with canonical bioinformatics tools, the ability to install new tools via pip and conda, and access to databases like NCBI and Ensembl.

The questions primarily focus on DNA and RNA sequencing data (WGS, scRNA-seq, methylation, ChIP-seq, metagenomics, Hi-C), as well as proteomics and metabolomics.

Performance Comparison: AI vs. Human Experts

Anthropic baselined the benchmark by tasking up to five domain experts to answer each question. This divided the tasks into two categories:

Human-Solvable Tasks

On the 76 tasks that at least one human could solve, Claude showed strong performance, often mirroring human strategies. In some cases, Claude used "intuition" to recognize patterns or sequences—similar to how the TATA box was discovered—which differs from traditional algorithmic approaches.

Human-Difficult Tasks

Of the 23 remaining questions that experts could not solve, Claude Sonnet 4.6 and more advanced models solved a significant fraction. Claude Mythos Preview achieved a 30% solve rate on these human-difficult problems.

Analysis of Claude's Research Strategies

Analysis of transcripts from Opus 4.6 revealed two primary strategies that allow Claude to outperform humans on difficult tasks:

  • The "Know-it-all" Approach: Claude leverages its vast internal knowledge of structural biology, molecular profiles, and meta-analyses from hundreds of thousands of papers to solve problems that would require a human to manually stitch together multiple databases.
  • Evidence Layering: When uncertain, Claude layers multiple different methods and converges on the answer that is supported by multiple lines of evidence.

Reliability and the Capability Frontier

Using Claude Mythos Preview to analyze the data, Anthropic found a distinct difference in the reliability of "wins" between the two sets:

  • Reliable Wins: On human-solvable problems, Opus 4.6 solved 86% of its successful problems at least 4 out of 5 times, indicating a reliable method.
  • Brittle Wins: On human-difficult problems, the share of "brittle wins" (solved only 1-2 times out of 5) jumped to 44%.

This suggests that while the models can crack the hardest problems, they often do so via "lucky" reasoning paths rather than consistently reproducible solutions.

Industry Convergence and Future Outlook

Anthropic notes that these results are echoed by CompBioBench, a similar benchmark released by Genentech and Roche. In that evaluation, Claude Opus 4.6 reached 81% overall accuracy and 69% on the hardest questions, further reinforcing the role of frontier models as useful collaborators in bioinformatics research.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch