From docs/extraction-audit.md.

Extraction audit v1 (2026-10-06)

Every dossier had a separate AI review, and those reviews logged hundreds of applied corrections, but no audit had estimated how many errors remain after review. This audit draws a seeded random sample of cited dossier claims and checks each one against its passages. The record is data/audit/extraction-audit-v1.json. Its scored summary is artifacts/extraction-audit-v1.md, which make check validates and make reproduce regenerates.

Sample

Checks and verdicts

Each cited quote was read in its surrounding text. Where an element of the claim lay outside that window, the full cited section was searched. Four checks were applied:

Each claim then gets one verdict:

Result

VerdictClaims
Correct142
Minor8
Material0

The other three are:

All eight are corrected in dossier revisions named extraction-audit-correction-2026-10-06. Each audited version is archived in data/evidence/history/ and linked by supersedes, so frozen ledgers still replay. The audit record keeps the audited hashes.

Independence and limits

This is a primary self-audit. The auditor is the same model that drafted and verified the dossiers and that ran most of their separate reviews, so it cannot see errors that this model makes systematically. The owner set a self-verification policy on 2026-10-06. Roadmap Milestone 2 still calls for an independent extraction sample, by a human or by a model from a different family; the frozen sample and protocol here can be reused for that.

The audit tests extraction against the cited passages. It does not test:

A claim can match its sources and still be historically wrong.

Review records

The same module normalizes every historical separate-review record into one vocabulary: applied, partly_applied, primary_found_applied, not_applied, deferred and decided_by_primary. The 17 disposition strings and two record shapes become one index, artifacts/review-index.json. A record with a disposition outside the vocabulary fails make check.