the science desk, on measuring a machine
The Only Score That Counts Was Sealed First
The best result in the history of machine learning and the most quietly inflated ones were produced by the same instrument, a test. What separates them is custody of the answer key. A score means something only when the answers were sealed before the model was built.
A benchmark is a fixed set of questions with a fixed set of answers, held still so that two different machines can be compared against the same thing. That is its whole job. It works exactly as long as the answers stay out of the room where the machine is being built, and it starts failing, silently and without any change to its own contents, the moment they get in.
Two results from the last six years sit at opposite ends of what this instrument can do. In 2020 a system called AlphaFold 2 predicted the folded shapes of proteins in a blind competition with a median accuracy score of 92.4 out of 100, close enough to laboratory measurement that the field's organizers treated a fifty-year problem as largely answered for single protein chains. Four years later Demis Hassabis and John Jumper shared the 2024 Nobel Prize in Chemistry for it, alongside David Baker for designing proteins that had never existed. Meanwhile, in the same period, a long series of language-model scores climbed toward perfect on tests the models may have partly read during training. Both kinds of number came out of a test. Only one came out of a sealed one.
where the answer key was kept
The competition AlphaFold won is called CASP, the Critical Assessment of protein Structure Prediction, and it has run every two years since 1994. Its design is a custody protocol. Organizers collect protein structures that experimentalists have solved in the lab but not yet published. Predictors receive only the amino-acid sequence. The true shapes are held back until every prediction is filed. Nobody can train on the answer because, at the moment of prediction, the answer exists in a handful of laboratory notebooks and a sealed queue, and nowhere on the public internet.
That last clause is the entire scientific value of the result. A model can be large and expensive, and neither property tell you whether it learned the physics of folding or the contents of a database. Only the seal tells you.
Now compare the ordinary benchmark. A few thousand questions are published, usually on the open web, usually with their answers alongside, because publishing is how a benchmark becomes a standard. A language model is then trained on a scrape of that same web. The test and the textbook are drawn from the same well. When researchers at Scale AI wrote a fresh set of grade-school math problems in 2024, matched in style and difficulty to a popular older set, some model families lost as much as eight percentage points on the new questions. Other frontier models barely moved, which is its own finding and an important one. The gap is the measurement of how much of the old score was arithmetic and how much was recall.
A test the model has already seen is a mirror, and every mirror reports a perfect score.
the economist's law, applied to a lab
Charles Goodhart, a British economist, observed in 1975 that any statistical regularity tends to collapse once pressure is placed on it for control purposes. The anthropologist Marilyn Strathern later compressed it into the version people quote: when a measure becomes a target, it ceases to be a good measure. A benchmark leaderboard is that sentence turned into infrastructure. The moment a score determines funding and headlines, every incentive in the system bends toward the score and away from the capability it was meant to sample.
The degradation has a recognizable sequence:
- Publication. The questions go public and become a shared standard. This is the useful phase.
- Absorption. The questions, their discussions, and their worked solutions enter training data, mostly by accident, sometimes by design.
- Saturation. Scores converge near the ceiling, and the remaining differences between models fall inside the noise.
- Replacement. A harder test is written, and the cycle starts again, each round a little more expensive than the last.
None of these steps requires anyone to cheat. Contamination is the default state of a published test in a world where models are trained on everything published.
what the upside actually looks like
It would be a misreading of all this to conclude that the capability is fake. The protein result is the counterexample, and it is enormous. The public AlphaFold database now holds predicted structures for more than two hundred million proteins, nearly every one cataloged by science, and researchers working on enzymes, vaccines, and neglected tropical diseases use those shapes as starting points every day. That is what a real capability looks like once it has passed through a sealed instrument: it keeps working on cases nobody had seen.
The discipline for everything else is simple to state and tedious to practice. Hold some answers back. Write new questions after the model's training cutoff. Report the date the test was written next to the score it produced. Treat a model's results on a public benchmark as a claim and its results on a private, post-cutoff set as the evidence for that claim, which is how assurance treats any capability claim once it leaves the lab. The engineering version of this argument, that agents need a ground truth they cannot quietly absorb, applies with equal force to the people grading them.
I keep returning to the protein structures sitting in those unpublished notebooks in 2020, waiting in a queue while the predictions came in. Someone had to agree not to release their own result for a few months so that a machine could be honestly measured against it. That small restraint is what turned a number into a fact.
The score is worth exactly as much as the seal on the answers.
The same record an agent receives. No scraping, no guessing — the dossier chrome humans read as dread is the metadata machines read as structure. One source of truth.
--- id: PRG-0091 title: The Only Score That Counts Was Sealed First kicker: the science desk, on measuring a machine captured: 2026-09-30T13:00:00Z status: open author: Ines Hargrove summary: The best result in the history of machine learning and the most quietly inflated ones were produced by the same instrument, a test. What separates them is custody of the answer key. A score means something only when the answers were sealed before the model was built. tags: [the record, measurement, custody, ai, science] --- A benchmark is a fixed set of questions with a fixed set of answers, held still so that two different machines can be compared against the same thing. That is its whole job. It works exactly as long as the answers stay out of the room where the machine is being built, and it starts failing, silently and without any change to its own contents, the moment they get in. Two results from the last six years sit at opposite ends of what this instrument can do. In 2020 a system called AlphaFold 2 predicted the folded shapes of proteins in a blind competition with a median accuracy score of 92.4 out of 100, close enough to laboratory measurement that the field's organizers treated a fifty-year problem as largely answered for single protein chains. Four years later Demis Hassabis and John Jumper shared the [2024 Nobel Prize in Chemistry](https://www.nobelprize.org/prizes/chemistry/2024/summary/) for it, alongside David Baker for designing proteins that had never existed. Meanwhile, in the same period, a long series of language-model scores climbed toward perfect on tests the models may have partly read during training. <Highlight>Both kinds of number came out of a test. Only one came out of a sealed one.</Highlight> ## where the answer key was kept The competition AlphaFold won is called CASP, the Critical Assessment of protein Structure Prediction, and it has run every two years since 1994. Its design is a custody protocol. Organizers collect protein structures that experimentalists have solved in the lab but not yet published. Predictors receive only the amino-acid sequence. The true shapes are held back until every prediction is filed. Nobody can train on the answer because, at the moment of prediction, the answer exists in a handful of laboratory notebooks and a sealed queue, and nowhere on the public internet. That last clause is the entire scientific value of the result. A model can be large and expensive, and neither property tell you whether it learned the physics of folding or the contents of a database. Only the seal tells you. Now compare the ordinary benchmark. A few thousand questions are published, usually on the open web, usually with their answers alongside, because publishing is how a benchmark becomes a standard. A language model is then trained on a scrape of that same web. The test and the textbook are drawn from the same well. When researchers at Scale AI wrote a fresh set of grade-school math problems in 2024, matched in style and difficulty to a popular older set, [some model families lost as much as eight percentage points](https://arxiv.org/abs/2405.00332) on the new questions. Other frontier models barely moved, which is its own finding and an important one. The gap is the measurement of how much of the old score was arithmetic and how much was recall. > A test the model has already seen is a mirror, and every mirror reports a perfect score. ## the economist's law, applied to a lab Charles Goodhart, a British economist, observed in 1975 that any statistical regularity tends to collapse once pressure is placed on it for control purposes. The anthropologist Marilyn Strathern later compressed it into the version people quote: when a measure becomes a target, it ceases to be a good measure. A benchmark leaderboard is that sentence turned into infrastructure. The moment a score determines funding and headlines, every incentive in the system bends toward the score and away from the capability it was meant to sample. The degradation has a recognizable sequence: 1. **Publication.** The questions go public and become a shared standard. This is the useful phase. 2. **Absorption.** The questions, their discussions, and their worked solutions enter training data, mostly by accident, sometimes by design. 3. **Saturation.** Scores converge near the ceiling, and the remaining differences between models fall inside the noise. 4. **Replacement.** A harder test is written, and the cycle starts again, each round a little more expensive than the last. None of these steps requires anyone to cheat. Contamination is the default state of a published test in a world where models are trained on everything published. <Marginalia label="On the instrument">The fix the field keeps rediscovering is old. Clinical medicine settled it with preregistration and blinded trials; geology settled it by drilling the core before anyone guesses what is in it. Each is the same move: commit the answer to a record nobody can edit before the prediction is made. The novelty in machine learning is only that the subject being tested has read the library.</Marginalia> ## what the upside actually looks like It would be a misreading of all this to conclude that the capability is fake. The protein result is the counterexample, and it is enormous. The public AlphaFold database now holds predicted structures for more than two hundred million proteins, nearly every one cataloged by science, and researchers working on enzymes, vaccines, and neglected tropical diseases use those shapes as starting points every day. That is what a real capability looks like once it has passed through a sealed instrument: it keeps working on cases nobody had seen. The discipline for everything else is simple to state and tedious to practice. Hold some answers back. Write new questions after the model's training cutoff. Report the date the test was written next to the score it produced. Treat a model's results on a public benchmark as a claim and its results on a private, post-cutoff set as the evidence for that claim, which is how [assurance treats any capability claim](https://www.adjective.us/blog/ai-assurance-is-not-a-policy-problem) once it leaves the lab. The engineering version of this argument, that agents need a [ground truth they cannot quietly absorb](https://www.adjective.us/blog/intelligence-primitives-agents-need-ground-truth), applies with equal force to the people grading them. I keep returning to the protein structures sitting in those unpublished notebooks in 2020, waiting in a queue while the predictions came in. Someone had to agree not to release their own result for a few months so that a machine could be honestly measured against it. That small restraint is what turned a number into a fact. The score is worth exactly as much as the seal on the answers.
{
"@context": "https://schema.org",
"@type": "Article",
"headline": "The Only Score That Counts Was Sealed First",
"description": "The best result in the history of machine learning and the most quietly inflated ones were produced by the same instrument, a test. What separates them is custody of the answer key. A score means something only when the answers were sealed before the model was built.",
"identifier": "PRG-0091",
"datePublished": "2026-09-30T13:00:00.000Z",
"dateModified": "2026-09-30T13:00:00.000Z",
"author": {
"@type": "Person",
"name": "Ines Hargrove",
"url": "https://progoff.com/authors/ines-hargrove"
},
"publisher": {
"@type": "Organization",
"name": "Progoff",
"url": "https://progoff.com"
},
"image": "https://progoff.com/records/the-only-score-that-counts-was-sealed-first/opengraph-image",
"keywords": "the record, measurement, custody, ai, science",
"articleSection": "Science",
"url": "https://progoff.com/records/the-only-score-that-counts-was-sealed-first",
"mainEntityOfPage": "https://progoff.com/records/the-only-score-that-counts-was-sealed-first",
"sha256": "0a8b3a613c7085b6ad7ed5a8c988b54d60967586428a67f72de4531c971c34e7",
"creativeWorkStatus": "open",
"isAccessibleForFree": true
}