How this museum is made
Almost everything in this museum was produced by a machine: the exhibits were written by one, and the measurements are retaken by another every day with no one watching. This page says which is which, how each works, and what happens when one goes wrong.
A museum that publishes automatically has a problem an ordinary one does not. Its authority rests on the reader believing the numbers, and a number produced by an unattended program is exactly the kind a careful reader should distrust. The usual answer is to keep quiet about it. That trades the credibility of the whole building against the convenience of not explaining, and it only works until somebody notices.
The answer here is the opposite. The automation is stated at the door, documented on this page, and built so that the failure mode is silence rather than invention. What follows is the whole mechanism.
All 45 exhibits in this museum were drafted by a language model, not typed by a person. What a person did was decide which 34 moments belong in a museum of this subject at all, and read what came back. The measurements are automated on a different schedule and by a different program. Nothing here is published without it being clear which is which.
| Surface | Made by | Cadence |
|---|---|---|
| The instrument panel at the entrance | Counted by a scheduled program from a public dataset | Recounted daily |
| Who still tells you how big their model is | Counted by the same program, same dataset | Recounted daily |
| The 34 works in the gallery | Drafted by a language model from a curated list, sources attached | Regenerated or corrected by hand |
| The 6 portraits and 5 artifacts | Drafted by a language model the same way | Regenerated or corrected by hand |
| Which works belong here at all | Chosen by Alan Salomon | Edited by hand |
| This page | Written by an agent session, counts read from the live data | Edited when the system changes |
Once a day a scheduled job on the server that hosts this site fetches one file: the notable-models dataset published by Epoch AI, a research organisation that maintains it publicly. The job recounts every panel from that file, writes the result as a dated snapshot, and points the site at the newest one. No page is rebuilt and no code is deployed; the numbers are data, and data can change underneath a running site.
Every snapshot is kept. The sources that publish this kind of data overwrite themselves, which makes the question “what did this say in March” unanswerable; keeping the series means this museum can eventually answer it.
The job validates before it writes. A truncated download, a renamed column in the source, or a file that is not the dataset at all leaves the previous snapshot untouched and stops. The site then keeps showing the last good reading, with the date it was taken printed beside it.
That is the deliberate trade. A page showing last week’s figure with last week’s date is honest and slightly stale. A page showing a number derived from half a file looks identical to a correct one and is worth nothing. Given the choice, this museum goes stale rather than wrong.
The site validates a second time when it reads, because a file written by an unattended job is not to be trusted simply because it exists. A snapshot missing its method, or carrying an impossibly small corpus, is rejected in favour of the copy built into the site.
Right now the site is serving a snapshot written after the current build was deployed, which means the scheduled job is running and its output is being accepted.
One dataset, unmodified except by the counting described on each exhibit. It is published under a licence that permits reproduction with credit, which is why it can be used here and why the credit appears wherever the numbers do.
Epoch AI, 'Data on AI Models'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/data/ai-models-documentation'
https://epoch.ai/data/ai-models-documentation · CC BY 4.0
The two kinds of automation here are not the same and should not be trusted the same way. The panels count, and a count can be checked against the dataset by anyone who downloads it. The exhibits were written by a language model, and prose is not checkable that way: what can be checked is whether its sources say what it says they say, which is why every exhibit carries them.
The judgement that is not automated is which moments belong in this museum at all. That list is chosen, argued over and edited by hand, and it is the part that decides what the building is about.
An absent confident figure is not proof a developer said nothing. It is proof that Epoch, who maintain this dataset, could not trace a figure to a source they were willing to stand behind, and that judgement has false negatives: Kimi K3 is marked speculative at 2.8T while Moonshot's own model card states the figure.