Local LLM Test Bench — coverage, scores & gaps

▸ Open the results explorer — every prompt, every answer, the code each model wrote, the playable games, and your own grading.
Eight batteries across a local-model roster. Click any column header to sort; click a row to expand that model's sub-scores. Read the Known gaps panel before drawing conclusions — several things people assume are measured here are not.

What each battery actually measures

Coverage matrix

every model × every battery — blank means not run, not zero
strong mid weak not run colour = value within that battery's own range; each battery has its own unit

Model detail

click a row above to inspect its sub-scores

Statistical confidence — most of these rankings are ties

Wilson 95% intervals · in the matrix marks a model that is NOT distinguishable from the leader

Judge reliability — how much the B1 scores can carry

3-seat blinded panel · agreement, spread, and self-preference

Efficiency — quality per GB of weights

what you can actually run, and what it costs you in VRAM

B1 is a 0–10 judged score, so “per GB” is a value ratio, not a physical unit — use it to compare, not to quote. A 6.7 GB model at 6.8/10 is a very different proposition from an 18 GB model at 7.4/10.

Why agentic runs fail — first-failure classification

more actionable than the completion rate: the failure mode

Which harness knob moves the answer (B7)

lower agreement = that setting destabilises output more

Speculative decoding — the throughput the suite doesn't show

lossless (temp-0 output is byte-identical) · zero VRAM cost · prism llama.cpp fork

n-gram spec-decode on EDIT / rewrite work --spec-type ngram-mod

The same feature on the suite's B5 arm — fresh generation

--spec-ngram-mod-n-match tuning

MTP draft heads — multi-token prediction

Game / graphical builds

one-shot browser games — specified in the v2.1 design, never built as a battery

These are ad-hoc laptop results from sessions 5–6, graded by blind code review on a bare one-line prompt — not a scored battery, no replicates, only 6 models. Treat them as anecdotes, not measurements.

Snake & Tetris — what actually got built

The roster that was specified but never implemented

Known gaps — what these numbers do NOT cover

read this before ranking anything

Caveats that change how you read the table