From docs/research/commander-ratings-v4.md.

Commander residual ratings v4: context comparator, estimated τ and leakage handling

Run on 2026-10-06 under the pre-registered run-4 design. This is an exploratory diagnostic, not a ranking of record, a measure of skill or a causal estimate. The baseline is unchanged.

Status: primary-verified. No separate review was run, under the owner's policy of 2026-10-06.

Authorization and pre-registration

Rows

The verdict: no improvement under the pre-registered rule

The pre-registered comparison is the commander model (context terms, with τ estimated in each training fold) minus the context model. Its held-out log loss, by weighting:

WeightingDifference95th percentile (the rule)95% interval
Battle-weighted−0.0174−0.0054−0.0320 to −0.0030
Campaign-weighted−0.0086+0.0092−0.0289 to +0.0130

Where the difference comes from (descriptive)

All rows below are held-out log-loss differences, with campaign-bootstrap 95% intervals.

ComparisonBattle-weightedCampaign-weighted
Run 3 as recorded: commander − strength only (v2 ledger)−0.0121 (−0.0241 to −0.0005)−0.0065 (−0.0226 to +0.0102)
Bridge: run 3's configuration on the v3 ledger−0.0134 (−0.0252 to −0.0017)−0.0073 (−0.0233 to +0.0096)
Run 4's rows, leakage handled, no context (τ = 0.5) − strength only−0.0118 (−0.0232 to −0.0005)−0.0049 (−0.0203 to +0.0114)
Context only − strength only+0.0111 (−0.0082 to +0.0325)+0.0093 (−0.0127 to +0.0312)
Commander + context (τ-hat) − strength only−0.0063 (−0.0272 to +0.0161)+0.0007 (−0.0245 to +0.0263)
The verdict: commander + context (τ-hat) − context−0.0174 (−0.0320 to −0.0030)−0.0086 (−0.0289 to +0.0130)
  1. Ledger v3 left run 3's result about where it was. The bridge reproduces run 3's configuration on the new ledger and gives slightly larger gains.
  2. Leakage handling reduced the gain a little. The reduction is mostly campaign-weighted: from −0.0073 to −0.0049.
  3. The context comparator was weaker than intended. Its held-out log loss was higher than force size alone, by 0.0111 and 0.0093.
    • Out of sample, the 19 theater, period and theater-by-period terms (prior SD 0.5) added noise rather than signal.
    • Beating the context model is therefore a lower bar than beating strength only. With context terms in both, the commander model does not beat strength only (−0.0063 and +0.0007).
    • This was not known before the run. Any change to the comparator, such as a tighter context prior, would have to be fixed before the next run's data are seen. It cannot rescue this run.
  4. The best held-out model was the commander model without context terms. It was still not better than force size alone beyond campaign-resampling noise under the campaign weighting.
  5. Campaigns. The commander model did better in 65 of 108 campaigns (exact sign test p = 0.043).
    • The battle-weighted gain survives resampling, but the campaign-weighted one does not. The gain therefore sits in campaigns with several battles.
  6. Concentration (stored predictions, no refit).
    • Forrest's 12 rows carry 46% of the net gain.
    • Beauregard's 7 rows carry 19% and Grant's 10 rows 15%. A row counts for both of its commanders, so shares overlap.
    • Without Forrest's rows, the verdict difference is −0.0098 battle-weighted and −0.0054 campaign-weighted. The battle-weighted 95th percentile is then −0.0004, only just below zero.
  7. Leakage rows (no refit). Without the 34 rows that had a post-start side in the ledger or a joint-command side, the difference is −0.0159 and −0.0055. The campaign-weighted 95th percentile is +0.0143.
  8. Monte Carlo check. Imputations 1–10 give −0.0179 and −0.0090; imputations 11–20 give −0.0168 and −0.0082.
  9. Temporal split (1861–1863 → 1864–1865, no verdict). The battle-weighted log losses are:
ModelLog loss
Commander + context0.6321
Context only0.6457
Strength only0.6505

Do commanders differ? (τ, descriptive)

All rows. The Laplace marginal likelihood peaks at τ = 0.6, with an approximate 95% profile interval of 0.2 to 1.0. That interval excludes zero.

Calibration.

Within folds and splits.

Reading. In sample, results differ by credited commander more than force size and the coarse context explain, and more than chance would produce under the context model. This is consistent with commander effects. It is equally consistent with army, subordinate, opponent or theater effects that commander identity carries and four theaters and three periods do not. It did not become a held-out improvement under both weightings.

What the estimates look like (no ordered ranking)

Who is rated.

Estimates. At τ = 0.6:

Labels.

The report lists commanders alphabetically, as the rule requires.

What changed from run 3, and what did not

Changed:

Unchanged:

What this means

Limits