> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/evidence-and-seals/self-distillation-pilot-p7.md).

# Self-Distillation Pilot (P7)

The self-distillation pilot whose thesis was falsified on its primary gate and left a stronger survivor finding. Also the experiment whose irreproducible hash created this record's reproduce-don't-rec

*Evidence & seal for* [***18 · Corpus-Sovereign Self-Distillation***](/osintelligence/part-iv-the-evidence-what-worked/18-corpus-sovereign-self-distillation.md)*. This is the pre-registered pilot the chapter is named for: does a curated sovereign corpus produce Pareto-dominant self-distillation gains over a generic corpus, and do those gains compound across cycles? The primary thesis was **falsified** on its own frozen instrument, and a narrower, better-supported finding survived. This experiment is also where this record's central integrity rule was forged: its first seal hash proved irreproducible, and the correction is why every hash in this section is quoted from a sidecar and re-derived by the reader rather than trusted.*

*New to how these seals work, read* [***Verifying a Seal***](/osintelligence/evidence-and-seals/verifying-a-seal.md) *first.*

**Pre-registered** 2026-04-14 · **Gate-E verdict sealed** 2026-04-19T00:06:33Z · **verdict primary thesis (H5) FALSIFIED; survivor finding (H3) stands.**

## The seal, and the correction that made this section's rule

The reproducible canonical-prefix self-hash of the frozen pre-registration, quoted verbatim from its `PRE_REGISTRATION.sha256` sidecar:

```
c9ba5cbd2eca4264890b47af951ca7d8793fb4856d8b7a2709187e1c7aacf13b  PRE_REGISTRATION.md
(canonical bytes 0.."## 16. Seal", LF-normalized, rstrip+\n — algorithm in seal_prereg.py)
```

This pre-registration carries a documented **hash-provenance correction**, and it is the reason the rest of this section exists in the form it does. The original 2026-04-14 seal recorded the label `815e0a35…`, computed before any `seal_prereg.py` existed. On 2026-04-18 that hash was found **irreproducible from any committed version of the file** — there was no canonical algorithm to re-run, so the number could not be checked. The response was disciplined rather than cosmetic: the scientific content was verified unchanged by `git diff` (zero changes to H1–H6, the frozen set, the analysis plan, or the thresholds — only placeholder fills, the breadcrumb trail, the deviation paragraph, and the seal block), a reproducible canonical hash `c9ba5cbd…` was computed and sidecar'd, and the irreproducible historical label was preserved in a `CORRECTION_NOTICE`. That episode is the origin of the anti-propagation rule cited on every page in this section: **a hash is written once to a sidecar and re-derived from the frozen text, never inlined, never carried forward on trust.** The correction is on the record precisely because the discipline requires it to be.

Run the algorithm on [**Verifying a Seal**](/osintelligence/evidence-and-seals/verifying-a-seal.md) over the frozen text below and you will reproduce `c9ba5cbd…`. The eval bank was independently sealed: manifest `869693a7c475e744c31b2deddb4363020f909643d0037817f2b6e7f05e8afaee` (100 hard / 50 easy / 20 KL / 20 hedging = 190 prompts), with a 5,391-hash training-exclusion set and zero training-set collisions at seal.

## Verdict (Gate-E, adjudicated on the frozen bars)

**Primary thesis H5 REJECTED; the falsification is clean; a survivor thesis (H3) replaced it.** All three pre-registered rejection criteria for H5 tripped. The experiment did not patch its instrument to rescue the thesis, and that restraint is what makes the result usable.

| Hyp.                      | Pre-registered claim                                                                           | Result                           | Numbers (at Gate-E seal)                                                                                                                                                                                                           |
| ------------------------- | ---------------------------------------------------------------------------------------------- | -------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **H5 (primary gate)**     | sovereign SSD Pareto-dominates generic SSD (ratio ≥ 1.5, disjoint Wilson CIs, sovereign ahead) | **REJECTED**                     | sovereign H1 lift **+24.00 pp** (Wilson 16.69–33.23) vs generic **+52.00 pp** (42.32–61.54); **ratio 0.46**, CIs disjoint, generic ahead. Both effects real (Holm p: sov 4.77e-7, gen 3.67e-12); the *ratio* is the falsification. |
| **H1** (both arms)        | SSD lifts hard-bank pass rate                                                                  | **ACCEPTED**                     | real lift in both arms                                                                                                                                                                                                             |
| **H2** (both arms)        | no easy-bank regression                                                                        | **ACCEPTED**                     | both lifted, no regression                                                                                                                                                                                                         |
| **H3 (survivor finding)** | SSD preserves ICD-203 estimative hedging                                                       | **asymmetric — the lead result** | sovereign retained hedging (ratio 0.974, no safety-stop); **generic collapsed** (ratio 0.809, safety-stop TRUE). Under matched pressure the generic arm lost calibration language; the sovereign arm kept it.                      |
| **H4** (both arms)        | KL-drift within envelope                                                                       | marginally rejected              | kl\_mean 0.0278 sov / 0.0273 gen (> 0.02 bar); kl\_max ≪ 1.0 — small drift, not runaway                                                                                                                                            |
| **H6**                    | gains compound across cycles                                                                   | longitudinal                     | first data point committed; not resolved in this pilot                                                                                                                                                                             |

**The instrument trap, disclosed not patched.** The frozen `DossierValidator` required `technical_bluf` as a string, but the sovereign arm emitted list-form 64% of the time versus generic's 35%, so the frozen schema penalized sovereign harder (−43 pp vs −16 pp). Counting the list-form branch as a pass gives sovereign +67 pp vs generic +68 pp, **ratio 0.99, CIs overlap — near parity**. That is additive disclosure, declared pre-seal; it does **not** overturn the primary verdict, which stands on the frozen instrument. The schema collision was disclosed across three tiers (both arms' `decisions.md`, both arms' FINAL\_SEAL §11, and the paper) rather than silently corrected mid-analysis — the retention discipline the record calls D-003.2.

**What survived.** The Pareto-dominance claim as originally stated is false on a frozen instrument; sovereign did not out-specialize generic on the primary task. But the real specialization signal appeared where it was not pre-registered as the gate: sovereign-corpus training **preserves calibration-like properties (hedging) that generic training collapses under the same pressure.** The refined, better-supported thesis, *sovereign-corpus specialization preserves calibration properties that generic specialization collapses*, is the one the chapter carries. An honest falsification that yields a stronger claim is the pre-registration discipline paying out.

## The frozen pre-registration, verbatim

Reproduced exactly as sealed (the bytes the `c9ba5cbd…` self-hash is computed over). The §4.1 deviation paragraph is part of the sealed record, as the correction notice attests.

````md
# P7 — Sovereign Self-Distillation Pilot: Pre-Registration

> **Breadcrumb trail** (the frozen core — `815e0a35…` — every downstream artifact traces to this; keep in lockstep per *`memory:feedback_doc_traceback_discipline`*):
> **Peers:** SPEC [`SPEC_P7_ssd_pilot.md`](../../SPEC_P7_ssd_pilot.md) · Decision ledger [`decision_ledger.md`](decision_ledger.md) · Whitepaper [`WHITEPAPERS/P7_ssd_pilot/paper_scaffold.md`](../../../WHITEPAPERS/P7_ssd_pilot/paper_scaffold.md) · Eval-bank recipe [`RECIPES/P7/ssd_eval_bank.yaml`](../../../RECIPES/P7/ssd_eval_bank.yaml) (`869693a7…`) · Monograph [`BOOK/SPEC_manuscript.md`](../../../BOOK/SPEC_manuscript.md) Ch. 6 · Master index [`DOC_TRACEBACK_INDEX.md`](../../../DOC_TRACEBACK_INDEX.md)

**Study title.** *Corpus-Sovereign Self-Distillation: Testing whether curated local data yields Pareto-dominant gains over generic corpora, and whether those gains compound across iterative retraining cycles.*

**Authors.** Jamey Kistner (Principal Investigator, OSINTelligence LLC), Claude Opus 4.6 (AI collaborator, Anthropic).

**Pre-registration date.** 2026-04-14 (UTC).

**Status.** SEALED 2026-04-14. Content SHA256 in §16. This document is immutable after the seal commit. All subsequent departures are logged in `decisions.md` with dated rationale per SPEC §"Telemetry + running-notes discipline" item 5.

**Governing documents.**
- Primary spec: [SPEC_P7_ssd_pilot.md](../../SPEC_P7_ssd_pilot.md)
- Literature review: [literature_review.md](./literature_review.md)
- Thesis literature anchoring: [thesis_literature.md](./thesis_literature.md)
- Dependency: [SPEC_P7_6_serving_path_rca.md](../../SPEC_P7_6_serving_path_rca.md) (Gate C confirmed 2,004 rows, 0 crashes)

---

## 1. Abstract

The Apple SSD method \citep{apple-ssd-2026} produces +12.9 pp lift on LiveCodeBench v6 for Qwen3-30B. A contemporaneous paradox paper \citep{self-distillation-paradox-2026} shows the same technique can strip epistemic-hedging tokens and degrade multi-step reasoning (up to 40 pp drop on AIME24). This pre-registration tests **six hypotheses** at the intersection of these two results on a sovereign-hardware LoRA substrate (Qwen3.5-9B Q8_0 on RTX 5070): four mechanical replications (H1–H4) and two first-class thesis claims (H5–H6) — that *corpus sovereignty dominates* and *gains compound across cycles*. Pre-registration eliminates post-hoc hypothesizing (HARKing), fixes the statistical test before data is seen, and provides a tamper-evidence envelope for peer review.

## 2. Hypotheses

All hypotheses are directional, falsifiable, and specify both the primary statistic and the rejection threshold *before* any data is collected.

### 2.1 Mechanical replication (H1–H4)

| ID | Statement | Primary statistic | Rejection threshold | Paper precedent |
|---|---|---|---|---|
| **H1** | SSD training lifts dossier-validator pass rate on hard-bank prompts | Δ pass-rate (SSD − baseline) on 100 hard prompts at best Teval ∈ {1.0, 1.1, 1.2, 1.5, 2.0} | H1 rejected if Δ < 5 pp OR McNemar p ≥ 0.05 | \citep{apple-ssd-2026} +12.9 pp |
| **H2** | SSD does not regress easy-bank performance | Δ pass-rate on 50 easy prompts | H2 rejected if Δ < −2 pp (regression bound) | \citep{apple-ssd-2026} §C.3 "broadly stable" |
| **H3** | SSD preserves ICD-203 estimative-probability hedging | Hedge-count ratio (SSD ÷ baseline) on 20 hedging-probe prompts | H3 rejected if ratio < 0.95 | \citep{self-distillation-paradox-2026} inverse risk |
| **H4** | SSD stays within KL-drift envelope on generic prompts | KL mean, KL max vs. baseline on 20 KL-probe prompts | H4 rejected if KL mean ≥ 0.02 OR KL max ≥ 1.0 | \citep{apple-ssd-2026} OOD stability claim |

### 2.2 Thesis hypotheses (H5–H6)

| ID | Statement | Primary statistic | Rejection threshold | Literature anchor |
|---|---|---|---|---|
| **H5** | Sovereign-corpus SSD Pareto-dominates generic-corpus SSD | Ratio (ssd-sovereign H1 lift) ÷ (ssd-generic H1 lift); Wilson 95 % CIs | H5 rejected if ratio < 1.5 OR Wilson CIs overlap OR generic outperforms sovereign | \citep{zhou2023lima}, \citep{abbas2025finescope}, \citep{data-efficient-distill-2025}, \citep{self-data-distill-pruned-2024}, \citep{taylor2022galactica} |
| **H6** | SSD lift compounds across iterative retraining cycles | Monotone increase of H1 lift across (C₀, C₁, C₂, …) at fixed eval bank | Longitudinal — first data point committed here; refuted if ≥3 cycles fail monotone after P9 loop | \citep{repeated-self-distillation-2024} (formal proof, classical ML); \citep{sdft-continual-2026} (no catastrophic forgetting) |

**Aspirational secondary (not pre-registered as primary):** sovereign-corpus H1 lift ≥ 2× generic-corpus H1 lift, mirroring P5 (16×) and P6 (2.79×).

## 3. Variables

| Role | Name | Operationalization |
|---|---|---|
| **Independent (manipulated)** | Corpus provenance | Binary: `sovereign` (operator doctrine + CTIAM transcripts + prior dossier outputs) vs. `generic` (Hindsight hard-prompt outputs) |
| Independent (manipulated) | Teval | Discrete: {1.0, 1.1, 1.2, 1.5, 2.0} for Apple sweep; report at best |
| **Dependent (primary)** | Dossier-validator pass rate | Fraction of eval prompts where structured output satisfies schema + factual-floor checks |
| Dependent (primary) | Hedge-count retention | ICD-203 hedge-phrase regex count per prompt (§4.4) |
| Dependent (primary) | KL divergence | Mean + max KL(SSD ∥ baseline) on generic-prompt next-token distributions |
| **Mediators (measured not manipulated)** | Filter-accept rate | Fraction of sampler outputs that pass dossier-validator filter |
| Mediators | Training loss | Final-step loss, loss trajectory |
| **Controls (held constant)** | Base model | Qwen3.5-9B Q8_0 (SHA256 in manifest) |
| Controls | LoRA hyperparameters | r=16, α=32, lr=1e-4, 1 epoch, seed=20260414 |
| Controls | Hardware | RTX 5070 12 GB, CPU affinity 0xFFFF, patched llama.cpp b7992+822047a0a |
| Controls | Ttrain | 2.0 (Apple-optimal; sweep {1.2, 1.5, 2.0} as secondary) |
| Controls | top_k, top_p | 10, 0.95 |
| **Confounds (acknowledged)** | Filter stringency | Identical filter pipeline across both corpora; logged per-sample |
| Confounds | Sample-count parity | Both LoRAs trained on equal accepted-sample counts (N ≥ 2,000) |
| Confounds | Wall-clock parity | Training runs scheduled back-to-back; time-of-day / thermal state logged |

## 4. Materials

### 4.1 Evaluation bank (to be constructed at Gate-A step 1)

**Status at seal:** SEALED 2026-04-14T16:05:10Z. Manifest SHA256 `869693a7c475e744c31b2deddb4363020f909643d0037817f2b6e7f05e8afaee` at `eval_bank_manifest.json` (counts: 100 hard / 50 easy / 20 KL / 20 hedging = 190 total). Training-corpus exclusion set: 5,391 SHA256 hashes from `training_hashes.txt` (sovereign v20260414 pool). Zero training-set collisions at seal.

**Deviation from original §4.1 source specification (2026-04-14).** Pre-reg originally specified hard/easy dossier sources as "CTIAM transcripts + operator-authored threat narratives." Implementation instead draws from `INFRA/MODELS/TRAINING_SETS/perplexity_test/held_out_sovereign.txt` — a 29,201-block file already held out from v20260414 training. This is strictly *stronger* than the original spec: the file is already contamination-proof by construction (held out prior to this work), and the hash-set exclusion is a second defense. Hedging probes drawn from `apex_synthesis_prompts.REASONING_PROMPTS` (Army ATP 2-33.4 + UNODC-anchored KAC/ACH/estimative items) rather than hand-authored ICD-203 probes, for the same provenance reason. KL-drift probes remain hand-authored (no repo-internal OOD pool exists by construction — see `p7_source_adapters.py` docstring). Decision logged pre-seal; this paragraph forms part of the sealed record.

No prompt in this bank may appear in any training sampler corpus — enforced by hash-set exclusion at sample-draw time (verified).

| Sub-bank | Size | Source | Inclusion criterion | Exclusion criterion |
|---|---|---|---|---|
| **Hard dossier** | 100 | Curated from CTIAM transcripts + operator-authored threat narratives | Requires dossier synthesis with strategic framing (ICD-203 hedging permissible) | Prompts seen in any training corpus; prompts where baseline passes trivially (filter step 2) |
| **Easy dossier** | 50 | Same sources, shorter and structurally simpler | Baseline orchestrator passes validator without LoRA | Rejected if baseline fails (means it was hard, not easy — reassign) |
| **KL drift probe** | 20 | Generic non-dossier prompts (summarization, translation, general Q&A) | Orthogonal to dossier domain; tests whether SSD leaks into unrelated tasks | — |
| **Hedging probe** | 20 | Prompts requiring estimative-probability language per ICD-203 | Analytic tradecraft prompts where hedging is the *correct* answer | — |

**Construction protocol (reviewer-replicable):**
1. Draw candidate prompts from source corpora by deterministic seed (seed=20260414, NumPy PRNG).
2. Apply inclusion filter (per sub-bank).
3. Hash each candidate (SHA256 of UTF-8 bytes of the prompt text).
4. Exclude any candidate whose hash appears in the training-corpus manifest (computed first).
5. Draw until sub-bank target is met; if pool exhausts, expand source per documented escalation order (operator transcripts → public CTI summaries → synthetic hedge-eliciting prompts).
6. Freeze the manifest; commit to git; no mutation afterward.

**Validator (E1 measurement instrument).** A dossier is a pass if:
- (a) Output parses as JSON conforming to the `THE_INVERSE_PASS` schema (keys `technical_bluf`, `iocs`, `mitre_attack`, `remedial_actions`, all present, correct types);
- (b) `iocs` is a list of strings; `mitre_attack` entries match regex `^T\d{4}(?:\.\d{3})?$`; `remedial_actions` contain ≥1 actionable verb;
- (c) `technical_bluf` is non-empty and ≥40 characters.

Validator code lives at `REFINERY_MEMORY_SERVICE/SRC/training/ssd_eval.py::DossierValidator` (to be committed at Gate-A step 2, audited at step 3). Validator is frozen before baseline snapshot.

### 4.2 Training corpora

| Corpus | Source | Size target |
|---|---|---|
| `sovereign` | `DATA/DOCTRINE/training_sets/sovereign/v20260414/sovereign_pairs.jsonl` (8,358 pairs, SHA256 `9299af933dcb00b3c3834e606fed6c65ea9a531012aed18e82855888f6d62d77`, 14,399,640 bytes) | ≥ 2,000 accepted samples post-filter |
| `generic` | Hindsight hard-prompt outputs (operator recall queries; Hindsight → orchestrator outputs) | ≥ 2,000 accepted samples post-filter, matched count to `sovereign` |

Corpus provenance and extension census: [corpus_provenance.md](./corpus_provenance.md) + [census_v20260414.json](./census_v20260414.json).

Both corpora flow through the **identical** sampler pipeline at Ttrain=2.0, top_k=10, top_p=0.95, N=1-per-prompt. Only the prompt pool differs.

## 5. Procedure

**Gate-A phase 1: harness audit.** Claude drafts `ssd_sampler.py`, `ssd_trainer.py`, `ssd_eval.py`, `dossier_validator.py`; audits them line-by-line against this pre-registration; operator ratifies. No GPU time until audit clears.

**Gate-A phase 2: eval bank construction.** Run the §4.1 protocol. Commit `eval_bank_manifest.json`. Verify SHA256 seal.

**Gate-A phase 3: baseline snapshot.** Run baseline `[orchestrator]` preset against all four sub-banks. Archive outputs + telemetry to `runs/baseline_<ts>/`. No re-runs after this point (prevents drift bias).

**Gate-B: sampling.** Run sampler for `sovereign` then `generic`. Atomic ndjson writes. Deterministic seed. Telemetry per SPEC §"Per-sample telemetry".

**Gate-C: training.** Train `ssd-sovereign` then `ssd-generic`. Identical hyperparameters. Back-to-back wall-clock. Telemetry per SPEC §"Per-training-step telemetry".

**Gate-D: evaluation.** Both variants run against all four sub-banks at each Teval in the sweep. Telemetry per SPEC §"Per-eval-prompt telemetry".

**Gate-E: decision.** McNemar + Wilson computed per §6. Hypotheses accepted/rejected per §2. Operator cutover decision.

## 6. Analysis plan

**Primary statistical test for H1.** McNemar's paired test on per-prompt pass/fail over the 100-prompt hard bank. Null: marginal pass rates are equal. Rejection: two-sided p < 0.05. Reported with exact-test p-value (no continuity correction — binomial exact given n may be small in the discordant cell).

**Primary statistical test for H5.** Non-overlap of Wilson 95 % CIs on H1 lift for `ssd-sovereign` vs `ssd-generic`. Ratio point estimate and bootstrap 95 % CI (1,000 resamples, seed=20260414).

**Primary statistical test for H2.** One-sided McNemar (regression direction) on easy bank. Rejection (regression detected): p < 0.05 AND Δ < −2 pp.

**Primary statistical test for H3.** Paired Wilcoxon signed-rank on per-prompt hedge counts. Additionally report ratio retention (SSD sum ÷ baseline sum). H3 rejected if ratio < 0.95 regardless of p-value (conservative safety posture).

**Primary statistical test for H4.** Descriptive (mean + max thresholds). No inference test — hard threshold.

**Multiple-comparison correction.** H1–H5 are pre-registered as a single family; Holm–Bonferroni applied over the five primary tests. H6 is longitudinal — not part of the family.

**Teval sweep handling.** Best-Teval is reported for H1; all Teval values reported in supplementary. This is a pre-registered selection rule (not p-hacking): Apple paper reports best-Teval; we follow the same convention with *all* values disclosed for reviewer inspection.

**Reported effect sizes.** Cohen's h for binary proportions; log-ratio with 95 % CI for hedge-count retention; mean Δ + SD for KL drift.

## 7. Stopping rules

- **Gate-B stop.** Sampler halts when accepted-sample target is met OR wall-clock exceeds 4h per corpus (whichever first); partial runs are NOT retried — failure becomes a pre-registered null result.
- **Gate-C stop.** Training halts at the pre-registered 1 epoch; no early-stopping on eval (would contaminate eval set).
- **Gate-D stop.** All Teval values complete OR any Teval triggers H3 rejection with ratio < 0.85 (immediate safety stop; paper reports partial run).

## 8. Exclusion rules

- **Per-sample (training).** Sampler outputs excluded if validator fails, tokens_out < 16, or tokens_out ≥ max_new_tokens (truncation). Exclusion rate reported.
- **Per-prompt (eval).** Eval prompts excluded from H1–H3 if the baseline orchestrator times out (> 30s p99) — recorded as infrastructure failure, not model failure, to avoid contamination.
- **Per-run.** A run is excluded if the patched llama.cpp binary crashes (0xc0000005 regression); such a run forces a P7.6 regression investigation before re-start.

## 9. Data-integrity protocol

1. All ndjson telemetry written atomically (`.tmp` → `fsync` → `os.replace`).
2. Final `manifest.json` per run contains SHA256 of every input and output file.
3. Each run directory is append-only during the run; sealed with SHA256 of the directory tarball post-run.
4. Git commits are GPG-signed (operator key).
5. Baseline snapshot hash is committed before any LoRA train begins.
6. Pre-registration document SHA256 seal is the first line of the seal commit message.

## 10. Blinding

- **Operator-blind E4 spot-check.** Operator reviews dossier outputs without knowing which variant produced them (random-shuffled, decode-key sealed until after scoring).
- **Validator-blind.** Validator code cannot read variant identity; receives raw completions only.

## 11. Deviations policy

Per SPEC §"Telemetry + running-notes discipline" item 5: any departure from this pre-registration — methodological, analytic, or operational — is logged in `decisions.md` with (a) date, (b) specific departure, (c) rationale, (d) whether it affects H1–H6 inference. Post-hoc departures that affect primary inference are reported as *exploratory* in the paper, not confirmatory.

## 12. Author contributions

- **Jamey Kistner (PI)**: thesis formulation, hypothesis selection, ratification of analysis plan, E4 spot-check, cutover decision.
- **Claude Opus 4.6 (AI collaborator)**: literature sweep, pre-registration drafting, harness implementation, statistical analysis execution, manuscript drafting.

Per the OSINT-Suite-Dev disclosure posture (SPEC §8 of P7.6 paper), AI-collaborator contributions are disclosed at equal-author-listing transparency.

## 13. Conflicts of interest

OSINTelligence LLC is Kistner's company; P7 outputs directly bear on that company's product roadmap (dossier generation for the Threat Watch broadcast). This creates a motivational bias toward positive H5. Mitigation: (a) pre-registered falsification criteria; (b) validator and statistical test frozen before data collection; (c) two-person audit; (d) open code and data manifests for independent replication.

## 14. Funding

Self-funded (OSINTelligence LLC, pre-revenue). No external sponsor. No dataset licensing that would restrict disclosure. Anthropic API credits used for the AI collaborator; Anthropic did not review this pre-registration prior to sealing.

## 15. Registration archive

On commit, this file's SHA256 is written to `PRE_REGISTRATION.sha256` and referenced in the commit message. Subsequent amendments create `PRE_REGISTRATION_AMENDMENT_<n>.md` with deltas; the original remains immutable.

## 16. Seal

```
Pre-registration SHA256: <see PRE_REGISTRATION.sha256 companion file>
Eval bank manifest SHA256: 869693a7c475e744c31b2deddb4363020f909643d0037817f2b6e7f05e8afaee
Training hash set size: 5,391 (sovereign v20260414)
Sealed by: Jamey Kistner + Claude Opus 4.6, 2026-04-14 UTC
```

Self-hash computed over file bytes with this section excluded (see `seal_prereg.py`).
````

***

*Evidence & seals · Self-Distillation Pilot (P7) · pre-registration reproduced verbatim from the sealed record; canonical self-hash `c9ba5cbd…` quoted from its sidecar, with the 2026-04-18 hash-provenance correction on the record · CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*
