What CASP14 actually measured
The overview says AlphaFold2 was accepted because of CASP. This is the measurement underneath that sentence: what r.m.s.d.95 is, the two numbers the paper leads with, how far ahead of the field they were, and the one thing the confidence intervals will not tell you.
"Competitive with experimental structures" is the phrase everyone quotes1. It is doing a lot of work, and it is worth knowing exactly what was put on the scale to earn it.
CASP hides the answer: structures already solved but not yet released, sequences published, predictions taken by a deadline, answers revealed afterwards2. So the comparison below is not a benchmark a lab chose for itself. It is 87 protein domains nobody in the contest had seen, scored by someone else.
The unit: r.m.s.d.95
Root-mean-square deviation is the average distance between your atoms and the real ones once the two structures are superimposed, in ångströms. Lower is better. For scale the paper offers its own yardstick: the width of a carbon atom is approximately 1.4 Å1. So a median error near 1 Å is not "close on the whole" — it is smaller than an atom.
The 95 matters. r.m.s.d.95 is the deviation computed at 95% residue coverage — the worst 5% of residues are dropped before averaging. Proteins have floppy tails, and a single disordered loop can dominate a plain r.m.s.d. and make an otherwise correct fold look wrong. Trimming the tail measures the model rather than the disorder.
It also means a number like 0.96 Å is not "every atom is within 1 Å". It is "once the worst twentieth is set aside, the typical atom is."
The two numbers the paper leads with
Both are medians over the CASP14 domain set, with 95% confidence intervals, against the next best method entered1.
Backbone
AlphaFold20.960.85–1.16Next best method2.82.7–4.0All-atom
AlphaFold21.51.2–1.6Next best method3.53.1–4.201.22.43.64.8Å · lower is better
Two things are visible in that shape that the quoted phrase does not carry.
The gap is not incremental. On backbone accuracy the median is 0.96 Å against 2.8 Å. That is not a better model in the same class; that is a different regime, which is why the field's reaction was a change of practice rather than a citation.
The intervals do not touch. AlphaFold's backbone interval runs 0.85–1.16 Å and the next best method's runs 2.7–4.0 Å. Whatever the sampling noise across 87 domains, the two do not meet. A result whose error bars overlap its comparison is an argument; this one is not.
What the interval will not tell you
A confidence interval describes sampling variation within the set that was measured. It says nothing about whether the set was the right one.
That is why CASP14, and not the number, is the load-bearing part. The same 0.96 Å computed on structures the model had seen during training would be meaningless, and would look identical on this chart. The taxonomy of ways that goes wrong in ML-based science is exactly the survey Kapoor and Narayanan ran across 17 fields and 329 affected papers5: nothing crashes, the number just improves.
CASP is the guard against this and it is a social one, not a statistical one — a deadline, an unreleased answer, and a scorer who is not you.
The interval measures the model. The venue measures whether the measurement means anything.
What this deep dive leaves out
The paper reports far more than these two medians. The per-target scatter against the other 145 entries, the ablations that show which architectural pieces carry the result, the recycling behaviour, and the accuracy breakdown by MSA depth are all in the paper and none of them are redrawn here — this page takes the two headline numbers and asks what they mean, which is a narrower job than reviewing the work.
One number in wide circulation is deliberately absent. A CASP14 median GDT_TS of 92.4 is quoted almost everywhere this result is described, including in the first draft of the overview this page hangs off. It is not the figure the paper leads with, and it could not be traced to this paper's own text, so it is not here.
Also absent: any claim about how AlphaFold performs outside CASP14. The paper addresses recent PDB structures separately, and the database that followed covers over 214 million sequences4 at confidence levels that vary enormously per residue. The predecessor's CASP13 result is a useful baseline for how fast this moved3, and it is a different measurement on a different set.
Back to the overview
The overview's claim was that a field moves when it has prior data, an assessment that hides the answer, and a validation path outside the model. The chart above is the second condition doing its job once, on 87 domains, in 2020 — and the reason the first sentence of this series is about the assessment rather than the architecture.
References
- [1]Jumper, J., Evans, R., Pritzel, A. et al. (2021). Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589. doi:10.1038/s41586-021-03819-2
- [2]Moult, J., Pedersen, J. T., Judson, R. & Fidelis, K. (1995). A large-scale experiment to assess protein structure prediction methods. Proteins: Structure, Function, and Bioinformatics 23(3). doi:10.1002/prot.340230303
- [3]Senior, A. W., Evans, R., Jumper, J. et al. (2020). Improved protein structure prediction using potentials from deep learning. Nature 577, 706–710. doi:10.1038/s41586-019-1923-7
- [4]Varadi, M., Bertoni, D., Magana, P. et al. (2024). AlphaFold Protein Structure Database in 2024: providing structure coverage for over 214 million protein sequences. Nucleic Acids Research 52(D1), D368–D375. doi:10.1093/nar/gkad1011
- [5]Kapoor, S. & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4(9), 100804. doi:10.1016/j.patter.2023.100804