384 engagements in 119 campaign groups · 384 dossiers · 3,577 claims · 16,090 citations · 1,849 registered sources · 305 engagements in the strength and command ledgers · 313 commanders credited
Latest rating run (4): no improvement over the context model under the pre-registered rule; no ordered ranking is given. Report · commander-ratings-v4.json

Generalship

An evidence-first research project for measuring historical command performance.

The long-term aim is to compare what commanders accomplished with the opportunities and constraints they faced, including the advantages they created before battle. Every result should be traceable from commander to campaign to engagement to disputed assumptions and supporting passages.

Explore the evidence: every commander, campaign, engagement, claim, quoted passage and pinned source, rebuilt from each commit.

Inspired by Ethan Arsht's military rankings, this project starts with a reproducible battle-level baseline, then builds toward campaign-level contribution. A battle residual is not an estimate of how many wins a commander caused. No validated commander ranking exists here yet.

Status

As of 2026-10-06:

How the research is produced

AI agents do the research, under rules meant to make every claim checkable. The project owner (the maintainer) sets the scope and authorizes each rating run.

  1. A frozen frame. The engagement list comes from a pinned, checksum-verified source table and is fixed before modelling, with defeats and inconclusive results included.
  2. Draft dossiers. Agents research by complete campaign. Each dossier covers seven dimensions: strength, terrain, logistics, information, objectives, responsibility and outcome. Every claim quotes an exact passage with its source ID and locator, and every source snapshot is stored with its SHA-256. Gaps and disagreements become explicit unknowns or disputed claims, never model recollection. A first pass inspects up to three source families per battle, plus one targeted follow-up.
  3. Verification. Every Civil War dossier batch had a separate, fresh-context AI review (GPT-6 Astra at xhigh reasoning effort, then Claude Opus 5.5 at high from 2026-09-24). The authoring agent checked each proposed correction against the source before applying it. Since 2026-10-06, by owner policy, the authoring agent verifies its own work against the sources and checks before committing, and labels it primary-verified. A separate review runs only when the owner asks for one.
  4. Gated model inputs. Draft dossiers never change model inputs. The frozen baseline uses only the pinned source tables. Rating runs use versioned strength and command ledgers (separately reviewed through v2, primary-verified since), and each run is bound to the SHA-256 of its exact inputs, and from run 4 its code, by a recorded owner authorization.

AI reviews and AI verification are not human historical adjudication or proof that sources are independent; a passage-backed claim can still be historically wrong. The code itself never calls a model, and only the explicit fetch command uses the network. See AI's role and validation and the evidence contract.

Results so far

Every result below is scored on held-out campaigns: each campaign in turn is left out of the fit and then predicted. Lower Brier score and log loss are better. None of these is a ranking of skill or a causal estimate, and the later runs leave the frozen baseline unchanged.

RunBattlesQuestionResult
Pilot baseline23 of the 127 pilot engagementsDoes relative force size beat equal odds?No: Brier 0.277 against 0.250
Strength evaluation21–37 pilot battles with graded strengthsDo reviewed strength estimates predict results?Weakly: worse than equal odds on the 21 grade A rows; on the 37 A–C rows, better than equal odds but not than the training prior (campaign-weighted)
Rating run 137 pilot battlesDoes adding commanders lower held-out log loss?No detectable commander signal, so no ordered list
Rating run 2126 full-war battles with graded strengthsThe same test on the full warLower under both weightings (0.6540 vs 0.6633 by battle, 0.6795 vs 0.6844 by campaign); every commander's interval includes zero
Rating run 3301 in-scope battles, 162 with a modelled sideThe same test with missing strengths modelled (grade E)Lower again (0.6476 vs 0.6597; 0.6586 vs 0.6651), but fragile: the campaign-weighted gain's 95% bootstrap interval includes zero, and Forrest's and Grant's rows carry much of the gain (uncertainty); only Forrest's 80% interval excludes zero
Rating run 4301 battles on ledger v3, 120 with a modelled sideDo commanders beat a theater-and-period comparator, beyond campaign-resampling noise under both weightings, with τ estimated and leakage removed?No: −0.0174 battle-weighted (bound −0.0054, cleared) but −0.0086 campaign-weighted (bound +0.0092, not cleared); the comparator itself did worse than force size alone, and Forrest's rows carry 46% of the gain. In sample τ ≈ 0.6 (95% profile interval 0.2–1.0, calibrated p = 0.01)

Runs 2 and 3 passed the test fixed in advance for them, which is read only as lower held-out log loss on those rows and had no uncertainty threshold. Run 4's pre-registered rule added that threshold, and the result did not meet it. The residual still mixes command with army quality, subordinates, theater, opponents, coding choices and modelling error, so the ordered lists in the run 2 and 3 reports are point summaries, not rankings. Each record names the make target that reproduces it.

Planned work

Beyond the Napoleonic work, the roadmap plans a campaign-level estimand defined before any enriched model (battle execution and campaign contribution are separate and never added together), opponent and army context, and graded outcomes rather than win or loss only. None of this is implemented yet. The explorer is a first, static version of the planned research interface.

Run it

Requires Python 3.11 or newer and nothing else. From the repository root:

make check                              # tests plus source, evidence and pipeline validation
make reproduce                          # regenerate the committed artifacts offline
python3 -m generalship inspect TN003    # print one engagement's record (TN003 is Shiloh)
python3 -m generalship packet TN003     # prepare a research assignment; makes no AI call
python3 -m generalship admission-check  # offline proposal audit; promotes nothing
python3 -m generalship site             # write the static explorer to _site/

python3 -m generalship fetch restores missing pinned upstream CSVs; it is the only command that uses the network. Checked-in historical snapshots are restored from Git, not silently refreshed. To run from another directory, use python3 -m generalship --root /path/to/generalship check with the package on the Python path, or install it with python3 -m pip install -e ..

Repository layout

PathContents
generalship/Python package and command-line interface
data/Source registry (sources.json) and raw snapshots, cohorts, dossiers (evidence/), strength and command ledgers, and authorization records
artifacts/Reports, predictions, research packets, review records and the input/output hash receipt
docs/Methodology, evidence contract, designs and roadmap; docs/research/ holds pass, ledger and run records
tests/Offline tests
design/, reviews/, prompts/, scripts/Supporting material for the Shiloh packets and review, and the researcher prompt

Documentation

Terms

TermMeaning
Engagement, campaign groupOne record in the CWSAC battle list, and the list's grouping of records into campaigns. Research, review and evaluation keep whole campaigns together.
CohortA fixed list of engagements chosen before modelling: v1 is the 127-engagement 1862–1863 pilot, v2 the 384-engagement full war.
DossierAn engagement's research record in data/evidence/: source-cited claims across the seven dimensions.
Explicit unknownA claim recording that the inspected sources do not establish a dimension.
Source familySources that share an origin; copies and reprints count once.
First passThe bounded research protocol: up to three source families per battle and one targeted follow-up.
PrimaryThe main AI agent that prepares and maintains the evidence, as distinct from the separate reviewer.
OwnerThe project maintainer. Owner decisions and run authorizations are recorded with dates.
Strength gradesA–C: reported figures, from A (directly applicable) to C (for example an opponent's estimate); D: no usable figure; E: a modelled typical size for a D side, used in runs 3 and 4.
AdmissionThe reviewed contract an enriched predictor must pass before it can join the frozen baseline. The rating runs are exploratory diagnostics outside it.

License and attribution

Project code is MIT-licensed. Third-party data keeps its own terms and attribution; see NOTICE. The engagement frame and baseline strengths come from four pinned, checksum-verified CWSAC tables in Jeffrey B. Arnold's American Civil War Battle Data. Nothing needs a paid service. GitHub Actions runs make check on Python 3.11 and 3.14 and publishes the explorer.