The AI Museum
Documentation· Data last recounted 23 September 2026

How this museum is made

Almost everything in this museum was produced by a machine: the exhibits were written by one, and the measurements are retaken by another every day with no one watching. This page says which is which, how each works, and what happens when one goes wrong.

A museum that publishes automatically has a problem an ordinary one does not. Its authority rests on the reader believing the numbers, and a number produced by an unattended program is exactly the kind a careful reader should distrust. The usual answer is to keep quiet about it. That trades the credibility of the whole building against the convenience of not explaining, and it only works until somebody notices.

The answer here is the opposite. The automation is stated at the door, documented on this page, and built so that the failure mode is silence rather than invention. What follows is the whole mechanism.

What made what

All 45 exhibits in this museum were drafted by a language model, not typed by a person. What a person did was decide which 34 moments belong in a museum of this subject at all, and read what came back. The measurements are automated on a different schedule and by a different program. Nothing here is published without it being clear which is which.

Each surface of the museum, what produces it, and how often it changes.
SurfaceMade byCadence
The instrument panel at the entranceCounted by a scheduled program from a public datasetRecounted daily
Who still tells you how big their model isCounted by the same program, same datasetRecounted daily
The 34 works in the galleryDrafted by a language model from a curated list, sources attachedRegenerated or corrected by hand
The 6 portraits and 5 artifactsDrafted by a language model the same wayRegenerated or corrected by hand
Which works belong here at allChosen by Alan SalomonEdited by hand
This pageWritten by an agent session, counts read from the live dataEdited when the system changes
What runs, and when

Once a day a scheduled job on the server that hosts this site fetches one file: the notable-models dataset published by Epoch AI, a research organisation that maintains it publicly. The job recounts every panel from that file, writes the result as a dated snapshot, and points the site at the newest one. No page is rebuilt and no code is deployed; the numbers are data, and data can change underneath a running site.

Every snapshot is kept. The sources that publish this kind of data overwrite themselves, which makes the question “what did this say in March” unanswerable; keeping the series means this museum can eventually answer it.

What happens when it fails

The job validates before it writes. A truncated download, a renamed column in the source, or a file that is not the dataset at all leaves the previous snapshot untouched and stops. The site then keeps showing the last good reading, with the date it was taken printed beside it.

That is the deliberate trade. A page showing last week’s figure with last week’s date is honest and slightly stale. A page showing a number derived from half a file looks identical to a correct one and is worth nothing. Given the choice, this museum goes stale rather than wrong.

The site validates a second time when it reads, because a file written by an unattended job is not to be trusted simply because it exists. A snapshot missing its method, or carrying an impossibly small corpus, is rejected in favour of the copy built into the site.

Right now the site is serving a snapshot written after the current build was deployed, which means the scheduled job is running and its output is being accepted.

Where the numbers come from

One dataset, unmodified except by the counting described on each exhibit. It is published under a licence that permits reproduction with credit, which is why it can be used here and why the credit appears wherever the numbers do.

Epoch AI, 'Data on AI Models'. Published online at epoch.ai. Retrieved from 'https://epoch.ai/data/ai-models-documentation'

https://epoch.ai/data/ai-models-documentation · CC BY 4.0

What this does not claim

The two kinds of automation here are not the same and should not be trusted the same way. The panels count, and a count can be checked against the dataset by anyone who downloads it. The exhibits were written by a language model, and prose is not checkable that way: what can be checked is whether its sources say what it says they say, which is why every exhibit carries them.

The judgement that is not automated is which moments belong in this museum at all. That list is chosen, argued over and edited by hand, and it is the part that decides what the building is about.

An absent confident figure is not proof a developer said nothing. It is proof that Epoch, who maintain this dataset, could not trace a figure to a source they were willing to stand behind, and that judgement has false negatives: Kimi K3 is marked speculative at 2.8T while Moonshot's own model card states the figure.