From docs/commander-ratings-v4.md.

Commander residual ratings, run 4: context comparator, estimated τ and leakage handling

Status: primary-verified design, fixed before run 4 computed any held-out result, 2026-10-06. It extends the commander-rating design and the grade E design. Runs 1–3, their records and the frozen baseline are unchanged.

On 2026-10-06 the author listed eight improvements. Item 1 was:

Run 3's verdict is fragile, and its uncertainty isn't reported.

and item 8 was:

Before starting the Napoleonic Wars: make the code work for any war.

The owner replied "Sounds good, go ahead and fix all 8" (authorization record). Under the owner's self-verification policy of the same day, the primary designed, implemented and verified this run with no separate review. The record binds the hashes of this design, every input and the code the run executes; the run refuses to start if any differs.

Order of work. The code was tested on synthetic data and smoke-tested end to end on the real rows with outcomes shuffled, so that no real held-out result was seen. This design, the code and the authorization were then committed before the run. Run 3's results and their uncertainty analysis were known when this design was written; it responds to them, and its own rule is fixed before its own numbers exist.

1. Why

Run 3 found lower held-out log loss for the commander model than for force size alone, under both weightings (by 0.0121 and 0.0065). Four weaknesses limit what that shows:

  1. Uncertainty. A campaign bootstrap puts the campaign-weighted gain's 95% interval at −0.0102 to +0.0226, which includes zero. Forrest's 12 rows carry 38% of the net gain.
  2. Leakage-prone rows. Two kinds of row may carry outcome information:
    • 17 rows have a side whose strength uses post-start information;
    • 18 rows have a side whose commander rule 5 chose as "the force compelling the result".

Without both, the stored predictions give 0.0088 and 0.0021.

  1. No context comparator. Commanders are tied to theaters and periods, so a commander term can absorb when and where a battle was fought. Run 3 compared commanders only with force size.
  2. τ was a convention. The commander spread was fixed at 0.5, never estimated. Yet whether commanders differ at all is the question τ answers.

2. Inputs and the pre-start view

Inputs. Run 4 uses:

The rows are those of run 3 (§3). The war is described by a profile (§10).

The pre-start view. The proposal said to "exclude those rows from the test". Run 4 instead removes the outcome-dependent information and keeps the battles. This follows the owner's decision of 2026-09-25 that battles should not be omitted for want of a strength (record), and it scales to sparser wars. The exclusion is still reported as a check (§7).

For each side whose ledger estimate carries post_start_information (20 sides in 17 rows), the frozen engine (estimate_side, strength design §4) re-estimates the side after removing every input in a dependence group that matches §3 row 3.

Grade E v2 (artifacts/strength-imputation-v2.json, generalship/imputation_v2.py).

3. Rows and attribution

4. Models

For battle i, with force ratio x_i = (A − B)/(A + B):

logit P(side A wins) = α + β·x_i + γ[theater_i] + δ[period_i] + η[theater_i × period_i] + θ[a_i] − θ[b_i]
α, β ~ Normal(0, 1)
γ, δ, η ~ Normal(0, κ²), κ = 0.5          (the context effects)
θ_k ~ Normal(0, τ²), independently          (the commanders)

5. Multiple imputation

There are M = 20 imputations, drawn as in run 3 (grade E design §5) with seed 20261006. Every fold, view and model uses the same imputed data sets.

6. The held-out test: the single verdict, pre-registered

Leave one campaign out over the 301 rows. In each fold, for each imputation:

The difference. For each row, d is the commander model's log loss minus the context model's. This is the clipped log loss of every scored run. Its mean is taken battle-weighted and campaign-weighted.

The rule.

  1. Resample the 108 campaigns with replacement 20,000 times (seed 20261008), using the floor-index percentile convention of uncertainty.py.
  2. The commander model improves only if the bootstrap's 95th percentile of the mean difference is below zero under both weightings. That is a one-sided 95% bound on each.

Reading the result.

7. Reported beside the verdict (descriptive, not part of it)

Other held-out models. Each comes from the same folds and imputations, and each comparison gets the same bootstrap.

ModelDescription
Commander + context at τ = 0.5Run 3's convention with context
Commander, no context, τ = 0.5Run 3's model on run 4's rows
Strength onlyForce size alone

Comparisons. Every model is compared with the context model or with strength only.

Bridge. This is run 3's configuration on the v3 ledger:

The commander model minus strength only, with its bootstrap, separates the ledger change from the method changes.

Leakage rows. The verdict difference is recomputed without the rows that had a post-start side in the ledger or a joint-command side, using the stored predictions (no refit).

Concentration (leave one commander out, stored predictions, no refit):

Other diagnostics.

τ on all rows.

Temporal split. Training on 1861–1863 and testing on 1864–1865, every model is fitted with τ-hat chosen on the training years. It has no verdict.

8. Ratings

Pooled estimates follow run 3's method (grade E design §5), with these changes:

Ranks, rank sets (two or more modelled battles), connectivity and prior dominance follow design §4. Each commander also reports:

9. Views

Base. Views run at the grade E medians, at the ratings' τ, unless stated. The single fit at the medians is the reference view median_imputation.

Robustness views. These set view_sensitive and unranked_in_view by the rules of design §6:

Other views:

Storage. Run 3 copied every view into every commander, which made its output 40 MB. Run 4 stores views compactly. For each commander with two or more modelled battles, each view keeps only these numbers in a single table:

10. A profile for each war

generalship/frame.py describes a war:

imputation_v2.py and ratings_v4.py read the profile instead of Civil War constants. The commander sign comes from each row, so a commander may command either side, as coalition wars need.

Still Civil War–specific:

Runs 1–3 keep their own frozen code.

11. Outputs, checks and gate

Outputs:

Checks before the run. The run refuses unless:

Tests. tests/test_ratings_v4.py and tests/test_imputation_v2.py cover:

Model runs are not part of the tests.

Outputs never enter artifacts/baseline.json. The run admits no feature.

12. Limits