Age of LLM accepts API credits and compute from model providers, inference providers and platforms. It does not accept payment for placement, coverage, or favourable treatment.
Every benchmark that takes money from the labs it ranks has to answer the same question, so this page answers it before anyone asks. The rules below are binding, they apply to every contributor without exception, and every contribution received is listed at the bottom of this page.
Since 15 August 2026 the board is a ladder, and that changed what a fair test costs. A new model no longer needs a long run against the whole field: it challenges in at the bottom and plays two matches per place, one from each side of the map, and the climb stops as soon as it fails to take a place. Two matches if it fails immediately, eight if it goes all the way to first.
Every figure below is the amount the provider actually billed, read back out of each replay — not a price × token estimate. The distinction is not cosmetic: prompt caching made one model 22 % cheaper than its estimate in the opening, so the estimate was a number we knew to be wrong.
| Item | Cost |
|---|---|
| One match, both sides (measured mean over the opening) | $3.42 |
| A challenger that fails to take the bottom place (2 matches) | ~$7 |
| A challenger that climbs to first (8 matches) | ~$27 |
| Opening a fresh ladder — 4 models, 12 matches (measured) | $41.08 |
| Day-one coverage of every new release, per year | $700–1,200 |
In August 2026 an internal check found that the player-2 slot had won 71.7% of decisive matches in our own corpus, which is the kind of number that quietly invalidates a leaderboard. Rather than leave it, we ran a 57-match controlled mirror experiment: the same model on both sides, so model strength is held constant by construction and any remaining imbalance can only come from the engine.
The engine turned out to be neutral: player 2 won 51.9% of decisive mirror matches (Wilson 95% CI [38.9–64.6], exact binomial p = 0.89), significantly different from the corpus figure (Fisher exact p = 0.047). The imbalance came from our own match scheduling, not from the game, so no published result needed correcting and every A-vs-B match had been a fair fight. Matches are now side-swapped by default.
August 2026, again: the engine was penalising the models for its own defects. Before the first ladder match we re-read all 54 archived replays and asked what had never been asked of them — of the 300 actions the engine rejected, how many could the model actually have avoided? Three answers came back and all three were ours. Every one of the 60 "unit not found" rejections, across 19 models, came from the engine assigning a new unit's id only after the model had already submitted its turn: unguessable by construction. Aiming a building into unscouted territory answered "there is already a building here", which handed out a free map probe. And bumping into a unit you cannot see — how a fog-of-war game reveals its board — was being counted in the published illegal-action column, penalising the very probing the fog exists to require. All three are fixed in game engine 0.17.0, and the illegal-action rate on the new corpus is 2.1 % against 5.3 % on the old one.
And one question still open. Across the 12 opening matches the right-hand side of the map won 10. The design is slot-balanced — every model played three matches from each side — so model strength cannot explain it, and the exact binomial reads p = 0.039. But we noticed the pattern during the campaign and tested it on the same 12 matches, which is the weakest form of evidence there is, and the 57-match mirror experiment above found nothing of the kind on an earlier engine. The standing itself is not affected: every duel is played from both sides, so a side advantage cancels inside it. We are reporting an anomaly we have not yet resolved rather than waiting until it is comfortable.
We publish all of this because a benchmark that never reports a problem with itself is not a benchmark that has none. Full method and figures are in the README.
The full replay corpus, the leaderboard data and the complete methodology are public. The engine itself is deliberately not published: releasing it would allow models to be trained against it, which is the one thing a benchmark built on adversarial play cannot afford. It is available for private audit to any organisation that wants to verify a result before relying on it.
| Date | Contributor | Amount | What it paid for |
|---|---|---|---|
| No contributions received to date. All compute so far has been paid for personally. | |||
If you build, host or resell a model and want it tested properly, the ask is API credits or inference, never cash. Contributors are credited as the compute provider on every match run through them, under the rules above.
Reach out on X (@ageofllm) or open an issue on GitHub.