From docs/research/commander-ratings-v3.md.

Commander residual ratings v3: all battles, with modelled strengths

Run 2026-09-25 under the reviewed grade E design, which extends the commander-rating design on the owner's decision that battles should not be omitted for want of a reported strength and that the all-battles test is the single verdict. This is an exploratory diagnostic, not a ranking of record, a measure of skill or a causal estimate. The baseline is unchanged. Status: separately reviewed (review); its five required corrections are applied.

Authorization

The owner authorized the run for the exact files named in the question: "You're good to run it" (record). The run refuses to start unless the design (15bdb6a2…), the grade E output (f798faf3…), both v2 ledgers (7855d4d3…, 495cef30…), the registry (91fc2526…) and the campaign table (93813220…) match. Code 82901a9; make commander-ratings-v3; outputs artifacts/commander-ratings-v3.{json,md}.

Grade E

artifacts/strength-imputation-v1.json fills the 262 grade D sides with a typical-size model (command echelon, side, period, theater) fitted on 328 sides with reported counts. Calibration widened σ by k = 1.08 (leave-one-out 80% coverage 0.78 before, 0.80 after; 0.80 leave one campaign out); the median absolute error is 0.67 on the log scale, about a factor of two. 70 sides are clipped by an applicable lower bound and 11 by an upper bound; 63 had a bound set aside under the rule 7 test; 15 are naval and 15 joint-command sides use an unknown echelon. (Design §2 counts 3 flotilla training rows by ledger echelon; the output records 2, because one is a joint-command side coded unknown. Both merge into unknown, so the fit is unaffected.)

The verdict: lower held-out log loss on all battles

301 battles in 108 campaigns; 162 have a modelled side and 17 a post-start side. Leave one campaign out, probabilities averaged over 20 imputations.

ModelLog loss, battle-weightedLog loss, campaign-weighted
Commander model0.64760.6586
Strength only0.65970.6651

What the estimates look like

94 commanders have two or more modelled battles (52 US, 42 Confederate).

What this means

How sure is the verdict? (2026-10-06)

The design reads the verdict only as "held-out log loss was lower on these rows", with no significance test. A descriptive uncertainty analysis (make rating-uncertainty) asks how firm that is. It uses the committed held-out predictions; nothing is refitted, and the verdict above stands as recorded.

Without both, the stored predictions give 0.0088 and 0.0021.

So the battle-weighted improvement is marginal and the campaign-weighted one cannot be told apart from zero. Most of either rests on a few well-known commanders and on rows that may carry outcome information. The successor design for run 4 sets an uncertainty rule and the leakage handling in advance.

Coverage correction (2026-10-06)

The per-commander coverage counted only the 301 modelled rows, so battles dropped as nested (GA012, LA009, VA033, VA034) vanished from it:

battles_attributed now counts every in-scope ledger attribution. Each commander lists their dropped_as_nested battles, and the three commanders with no modelled battle appear as coverage_only. The run was re-emitted with the corrected code. Every estimate, interval, rank, view and held-out prediction is bit-identical, and the Markdown report is unchanged. Only these coverage fields differ (primary-verified by a structural diff).

Limits