Skip to content
Part of From DENDRAL to AlphaFold: what let AI into science was the assessment

What CASP14 actually measured

The overview says AlphaFold2 was accepted because of CASP. This is the measurement underneath that sentence: what r.m.s.d.95 is, the two numbers the paper leads with, how far ahead of the field they were, and the one thing the confidence intervals will not tell you.

4 min readFiled underScienceBenchmarksReproducibility

"Competitive with experimental structures" is the phrase everyone quotes1. It is doing a lot of work, and it is worth knowing exactly what was put on the scale to earn it.

CASP hides the answer: structures already solved but not yet released, sequences published, predictions taken by a deadline, answers revealed afterwards2. So the comparison below is not a benchmark a lab chose for itself. It is 87 protein domains nobody in the contest had seen, scored by someone else.

The unit: r.m.s.d.95

Root-mean-square deviation is the average distance between your atoms and the real ones once the two structures are superimposed, in ångströms. Lower is better. For scale the paper offers its own yardstick: the width of a carbon atom is approximately 1.4 Å1. So a median error near 1 Å is not "close on the whole" — it is smaller than an atom.

The 95 matters. r.m.s.d.95 is the deviation computed at 95% residue coverage — the worst 5% of residues are dropped before averaging. Proteins have floppy tails, and a single disordered loop can dominate a plain r.m.s.d. and make an otherwise correct fold look wrong. Trimming the tail measures the model rather than the disorder.

It also means a number like 0.96 Å is not "every atom is within 1 Å". It is "once the worst twentieth is set aside, the typical atom is."

The two numbers the paper leads with

Both are medians over the CASP14 domain set, with 95% confidence intervals, against the next best method entered1.

Backbone

AlphaFold20.960.851.16Next best method2.82.74.0

All-atom

AlphaFold21.51.21.6Next best method3.53.14.201.22.43.64.8

Å · lower is better

AlphaFold2 against the next best CASP14 entrant, on the two accuracy measures the paper reports. Whiskers are the 95% confidence interval of the median. Redrawn from the values in the text; the paper is CC BY 4.0, so its own figures could have been reproduced instead — a plot rebuilt from stated numbers is checkable against them, and a copied image is not. Jumper et al. 2021, Nature 596, 583–589

Two things are visible in that shape that the quoted phrase does not carry.

The gap is not incremental. On backbone accuracy the median is 0.96 Å against 2.8 Å. That is not a better model in the same class; that is a different regime, which is why the field's reaction was a change of practice rather than a citation.

The intervals do not touch. AlphaFold's backbone interval runs 0.85–1.16 Å and the next best method's runs 2.7–4.0 Å. Whatever the sampling noise across 87 domains, the two do not meet. A result whose error bars overlap its comparison is an argument; this one is not.

What the interval will not tell you

A confidence interval describes sampling variation within the set that was measured. It says nothing about whether the set was the right one.

That is why CASP14, and not the number, is the load-bearing part. The same 0.96 Å computed on structures the model had seen during training would be meaningless, and would look identical on this chart. The taxonomy of ways that goes wrong in ML-based science is exactly the survey Kapoor and Narayanan ran across 17 fields and 329 affected papers5: nothing crashes, the number just improves.

CASP is the guard against this and it is a social one, not a statistical one — a deadline, an unreleased answer, and a scorer who is not you.

The interval measures the model. The venue measures whether the measurement means anything.

what the chart above is actually evidence of

What this deep dive leaves out

The paper reports far more than these two medians. The per-target scatter against the other 145 entries, the ablations that show which architectural pieces carry the result, the recycling behaviour, and the accuracy breakdown by MSA depth are all in the paper and none of them are redrawn here — this page takes the two headline numbers and asks what they mean, which is a narrower job than reviewing the work.

One number in wide circulation is deliberately absent. A CASP14 median GDT_TS of 92.4 is quoted almost everywhere this result is described, including in the first draft of the overview this page hangs off. It is not the figure the paper leads with, and it could not be traced to this paper's own text, so it is not here.

Also absent: any claim about how AlphaFold performs outside CASP14. The paper addresses recent PDB structures separately, and the database that followed covers over 214 million sequences4 at confidence levels that vary enormously per residue. The predecessor's CASP13 result is a useful baseline for how fast this moved3, and it is a different measurement on a different set.

Back to the overview

The overview's claim was that a field moves when it has prior data, an assessment that hides the answer, and a validation path outside the model. The chart above is the second condition doing its job once, on 87 domains, in 2020 — and the reason the first sentence of this series is about the assessment rather than the architecture.

References

  1. [1]Jumper, J., Evans, R., Pritzel, A. et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589. doi:10.1038/s41586-021-03819-2
  2. [2]Moult, J., Pedersen, J. T., Judson, R. & Fidelis, K. (1995). A large-scale experiment to assess protein structure prediction methods. Proteins: Structure, Function, and Bioinformatics 23(3). doi:10.1002/prot.340230303
  3. [3]Senior, A. W., Evans, R., Jumper, J. et al. (2020). Improved protein structure prediction using potentials from deep learning. Nature 577, 706–710. doi:10.1038/s41586-019-1923-7
  4. [4]Varadi, M., Bertoni, D., Magana, P. et al. (2024). AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Research 52(D1), D368–D375. doi:10.1093/nar/gkad1011
  5. [5]Kapoor, S. & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4(9), 100804. doi:10.1016/j.patter.2023.100804

Read next

Have something in mind?

Tell us what you are building and we will tell you honestly whether we are the right studio for it.