Local LLM Test Bench — coverage, scores & gaps
▸ Open the results explorer — every prompt, every answer, the code each model wrote, the playable games, and your own grading.
Eight batteries across a local-model roster. Click any column header to sort;
click a row to expand that model's sub-scores. Read the Known gaps panel before drawing
conclusions — several things people assume are measured here are not.
What each battery actually measures
Coverage matrix
every model × every battery — blank means not run, not zero
strong
mid
weak
not run
colour = value within that battery's own range; each battery has its own unit
Model detail
click a row above to inspect its sub-scores
Statistical confidence — most of these rankings are ties
Wilson 95% intervals · ≈ in the matrix marks a model that is NOT distinguishable from the leader
Judge reliability — how much the B1 scores can carry
3-seat blinded panel · agreement, spread, and self-preference
Efficiency — quality per GB of weights
what you can actually run, and what it costs you in VRAM
B1 is a 0–10 judged score, so “per GB” is a value ratio, not a
physical unit — use it to compare, not to quote. A 6.7 GB model at 6.8/10 is a very different
proposition from an 18 GB model at 7.4/10.
Why agentic runs fail — first-failure classification
more actionable than the completion rate: the failure mode
Which harness knob moves the answer (B7)
lower agreement = that setting destabilises output more
Speculative decoding — the throughput the suite doesn't show
lossless (temp-0 output is byte-identical) · zero VRAM cost · prism llama.cpp fork
n-gram spec-decode on EDIT / rewrite work — --spec-type ngram-mod
The same feature on the suite's B5 arm — fresh generation
--spec-ngram-mod-n-match tuning
MTP draft heads — multi-token prediction
Game / graphical builds
one-shot browser games — specified in the v2.1 design, never built as a battery
These are ad-hoc laptop results from sessions 5–6, graded by
blind code review on a bare one-line prompt — not a scored battery, no replicates, only 6 models.
Treat them as anecdotes, not measurements.
Snake & Tetris — what actually got built
The roster that was specified but never implemented
Known gaps — what these numbers do NOT cover
read this before ranking anything
Caveats that change how you read the table