From artifacts/commander-ratings-v3-uncertainty.md.

How sure is rating run 3's verdict?

Descriptive, from the committed held-out predictions of run 3 (SHA-256 24f1040ec9a7…). Nothing is refitted and the run's own verdict stands as recorded. Differences are commander-model log loss minus strength-only log loss, so negative favours commanders.

301 held-out rows in 108 campaigns.

WeightingDifference95% interval (campaign bootstrap)80% intervalResamples ≥ 0
Battle-weighted-0.0121-0.0241 to -0.0005-0.0199 to -0.00452.1%
Campaign-weighted-0.0065-0.0226 to +0.0102-0.0170 to +0.004421.7%

Campaigns: commander model better in 61, strength only better in 47, ties 0; exact two-sided sign-test p = 0.211.

Where the difference sits (no refit)

Rows keptRowsBattle-weightedCampaign-weighted
All301-0.0121-0.0065
Without rows with a post-start strength side284-0.0109-0.0060
Without rows with a joint-command side283-0.0097-0.0021
Without either267-0.0088-0.0021
Without the rows of Nathan Bedford Forrest and Ulysses S. Grant278-0.0046-0.0024

Net log-loss gain over all rows: 3.643. The commanders whose rows carry the most and the least of it (a row counts for both of its commanders):

CommanderRowsNet gain
Nathan Bedford Forrest12+1.389
Ulysses S. Grant11+0.982
P.G.T. Beauregard7+0.913
Benjamin Franklin Butler6+0.716
Thomas J. Jackson8+0.694
John Bell Hood9+0.601
John McNeil3+0.340
David G. Farragut4+0.332
Horatio G. Wright2-0.313
Abel Streight1-0.319
Ambrose E. Burnside7-0.364
James H. Wilson4-0.589
Jubal A. Early8-0.623

Limits