> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/part-iv-the-evidence-what-worked/18-corpus-sovereign-self-distillation.md).

# 18 · Corpus-Sovereign Self-Distillation

> **A companion paper.** Cited in-series by title, the pre-registered pilot whose **falsified primary hypothesis and surviving secondary finding** the spine reports verbatim: the sovereign-corpus model retained calibrated hedging (ratio 0.974) where the generic-corpus model collapsed to a safety-stop (0.809) under matched load. The paper is equally the reference instance of the series' falsification discipline: the record keeps the failure, the decision trail, the pivot, the fix, the falsification, and the ratification in one auditable chain.
>
> **Status note.** Body §§1–10 preserved as frozen at seal; §11–§12 were appended at promotion (2026-05-12). The transition of record at 2026-04-16T23:26:12Z falls mid-experiment: baseline and sampler phase sealed under Opus 4.6, training spans the boundary, evaluation seals under Opus 4.7, attribution is granular to section and seal, and this paper is itself a datum in the companion natural experiment (*The Guard Changes at 23:26Z*).

> **What is new here.** The contribution is a pre-registered self-distillation ablation in which the operator's own first-order thesis was falsified on the frozen instrument and kept in the record rather than rescued: sovereign-corpus distillation did not Pareto-dominate generic-corpus distillation on hard-dossier lift (ratio 0.46, confidence intervals disjoint, in the wrong direction), and no validator patch or hypothesis migration was permitted to recover it. Two findings the pre-registration did not rank first are what survived. On a calibration axis, the sovereign arm held its epistemic hedging (0.974 of baseline, no safety-stop) while the generic arm collapsed below the safety floor (0.809) under matched training pressure, which is the paper's lead result and the series' FC-2 evidence. And a constrained-decoding sub-experiment isolated a boundary the literature had not drawn: a per-token-legal grammar constraint can be off-distribution at the compositional level, so token-level alignment is not compositional alignment, with a practitioner rule (keep per-item patterns structural; enforce lexicon semantics at the runtime filter, not the logit mask).
>
> **Deepest water.** §4.6 and Table 5, the hedging-retention asymmetry that is the surviving thesis; §4.9's H5 falsification read honestly (the 0.46 primary and the 0.99 schema-adjusted secondary both kept, neither one laundered into "the true result"); and §4.2's failure→decision→pivot→fix→falsification→ratification sampler trajectory, including the D-003.2 grammar-constraint abort preserved under seal.

### Abstract

Apple's Self-Distillation (SSD, 2026) produces +12.9 pp lift on LiveCodeBench v6 for Qwen3-30B. A contemporaneous paradox paper (2026) demonstrates the same recipe can strip epistemic-hedging tokens and degrade multi-step reasoning by up to 40 % on math benchmarks including AIME24. We pre-register and execute six hypotheses at the intersection of these findings on a fully sovereign stack (Qwen3.5-9B Q8\_0 LoRA on a single RTX 5070), separating **mechanical replication** (H1–H4) from **thesis claims** (H5–H6). H5 asserts that corpus-curation quality *Pareto-dominates* corpus size under SSD: a curated sovereign corpus (\~8.4k instruction pairs) produces ≥ 1.5× the H1 lift of a generic corpus of equal accepted-sample count. H6 asserts that SSD lift **compounds across iterative retraining cycles**, testing a claim formally proved for classical ML (Pareek et al. 2024) but untested at LLM scale. We provide: (a) a SHA256-sealed pre-registration with McNemar + Wilson + Holm–Bonferroni analysis plan; (b) a contamination-free eval bank constructed by seeded deterministic draw with hash-set exclusion of all training prompts; (c) an open harness auditable line-by-line against the pre-registration; and (d) a three-tier sovereignty architecture (weights / index / promotions) that preserves curated-corpus signal while admitting whole-machine retrieval without weight dilution. **Results (sealed 2026-04-19T00:06:33Z).** The primary H5 Pareto-dominance claim is **falsified** on the frozen validator (sovereign/generic H1-lift ratio 0.46, Wilson CIs disjoint; generic out-lifts sovereign on the primary instrument). A pre-declared schema-adjusted secondary analysis, emitted additively, shows **near-parity** (ratio 0.99, CIs overlap): the primary gap traces to an asymmetric schema collision on a single field, not a specialization failure. H1 and H2 are accepted on both arms (sovereign +24.00 pp, generic +52.00 pp; no regression on either); H4 is marginally rejected on both at comparable magnitudes (kl\_mean ≈ 0.027 each). A novel **survivor finding** emerges on the pre-registered H3 hedging gate: the sovereign arm retains calibrated hedging (ratio 0.974, no safety-stop) while the generic arm collapses to a safety-stop (0.809, automatic rejection): the asymmetry the series' spine cites as its FC-2 evidence. H6 is deferred to post-deployment instrumentation.

### 1. Introduction

#### 1.1 Two papers that must be read together

Apple's Self-Distillation paper reports a +12.9 pp headline lift produced by a deliberately minimal recipe: sample from the model, accept on a lightweight validator, train on the acceptances. A contemporaneous paradox paper runs the same class of recipe and reports the **opposite**: up to 40 pp degradation on AIME24, with the mechanism identified as *hedging-token erosion*: the model, distilling against its own confident outputs, loses the epistemic-probability language its calibrated baseline relied on to reason in steps. The field response has largely been to pick one and move on. We argue both must be engaged simultaneously, because the conditions under which SSD reverses (broad, thin, or operator-unaligned corpus coverage) are precisely the conditions most practitioners use by default when building a sovereign agent.

#### 1.2 Thesis

We pre-register **corpus sovereignty** (curation quality, domain alignment, and operator-authored provenance) as the primary axis along which SSD outcomes diverge between the positive and negative published results. The two outcomes are not incompatible; they are endpoints of a corpus-quality axis that has not been pre-registered and measured at LLM scale. We measure it. H5 tests whether a curated sovereign corpus dominates a generic corpus on hard-dossier lift (the Apple metric class); H3 tests whether the same sovereign corpus *also* preserves the calibrated hedging the paradox paper records as the failure mode. Either hypothesis can falsify independently; the pair is informative in every combination.

#### 1.3 Contributions

1. **Pre-registration:** SHA256-sealed, 16-section document freezing hypotheses, variables, analysis plan, and stopping rules before any data collection.
2. **Contamination-free eval bank:** 190 prompts (100 hard dossier + 50 easy dossier + 20 KL drift + 20 hedging), deterministically drawn under seed 20260414, hash-set-excluded against the 5,391-hash training corpus, with per-draw telemetry.
3. **Open audit harness:** sampler, trainer, evaluator, and bank-builder committed with literature-anchored docstrings; no fabricated data (operator-ratified allowlists stubbed explicitly).
4. **Three-tier sovereignty architecture:** weights / retrieval-index / allowlisted promotions, with explicit non-crossing invariants.
5. **Hedging-retention gate:** ICD-203 estimative-probability hedge-count ratio as a first-class safety metric (H3), adopting the paradox paper's failure mode as the conservative floor.
6. **Thesis claim (H5)** with pre-registered falsification criterion: if sovereign does not Pareto-dominate generic, the thesis is recorded as *refuted* in the same commit as the pre-registration.
7. **Compounding claim (H6):** longitudinal; first cycle recorded here; refutation requires three successive cycles failing monotone improvement.

#### 1.4 Scope and non-goals

**In scope:** a single pre-registered pilot on a single base model (Qwen3.5-9B; bf16 train / Q8\_0 infer), single hardware substrate, single training seed, single validator snapshot, all effects reported with Wilson 95% CIs and Holm–Bonferroni family-wise adjustment across H1–H4; two paired training runs (sovereign and generic) under byte-identical pipeline configuration except the pre-registered corpus swap; a contamination-free eval bank (N = 190) built under AntiLeakBench and LiveBench methodology; and a pre-declared schema-adjusted secondary analysis branch, emitted additively, to surface (not suppress) validator-specification gaps discovered mid-pipeline. **Out of scope:** scale-law extrapolation, multi-seed meta-analysis, preference-optimization control arms, cloud inference of any kind, and base-weight modification (all effects isolated in the LoRA adapter). The pilot is deliberately narrow: one instrument, one substrate, one seed, honestly reported.

### 2. Related Work

The pilot sits at a triple intersection: (i) **Apple SSD** establishes the positive-result recipe and the +12.9 pp headline against which any sovereign replication is measured; (ii) the **self-distillation paradox paper** establishes the failure mode (hedging-token erosion) against which H3 is pre-registered as a conservative floor; (iii) **LIMA** (Zhou et al. 2023) is the canonical precedent for the H5 quality-over-scale claim in the classical SFT regime, with FineScope (Bhattacharyya & Kim 2025) its data-efficient contemporary. **Pareek et al. (2024)** provide the only formal compounding proof in the classical-ML setting; H6 pre-registers the LLM-scale analogue. **SDFT (2026)** anchors H2's no-catastrophic-forgetting expectation; **STABLE** (Hoy et al. 2026) anchors H4's KL-drift threshold; **ICD-203** (ODNI 2015) supplies the hedge-phrase list for H3; **AntiLeakBench** (ACL 2025) and **LiveBench** anchor the contamination defense. The bibliography was frozen in the pre-registration's companion specification; no anchor was added post-seal.

### 3. Methods

#### 3.1–3.2 Pre-registration and substrate

The pre-registration sealed 2026-04-14 with its SHA256 self-hash committed to git alongside the document; subsequent deviations are logged in a decision ledger with dated rationale, and any departure affecting H1–H6 inference is reported as exploratory, not confirmatory. Substrate: single-operator bare-metal Windows 11; i7-14700F (20C/28T), RTX 5070 12 GB (sm\_120 Blackwell), 128 GB DDR5; patched llama.cpp b7992 + 822047a0a (the mRoPE RCA companion's binary); CPU affinity pinned; thermal governor active at 85 °C; no cloud inference.

#### 3.3 Eval-bank construction

| Sub-bank     | N   | Source                                                                        | Anchor                    |
| ------------ | --- | ----------------------------------------------------------------------------- | ------------------------- |
| Hard dossier | 100 | Held-out sovereign blocks 500–1200 chars, ≥ 2 CTI keywords, canonical wrapper | LIMA + Apple SSD          |
| Easy dossier | 50  | Same file, 200–500 chars; regression sentinel                                 | Apple SSD regression rail |
| KL drift     | 20  | Hand-authored public-domain non-intel prompts                                 | STABLE                    |
| Hedging      | 20  | Reasoning prompts filtered to Admiralty / ICD-203 / ACH items                 | ICD-203 + paradox paper   |

Construction: deterministic seed 20260414, SHA256 hash-set exclusion against the 5,391-hash training corpus, atomic per-draw telemetry, manifest self-hashed and frozen (869693a7…).

#### 3.4–3.5 Corpora and base-model provenance

**Sovereign corpus** (v20260414): 8,358 pairs, SHA256 9299af93…d62d77, 14.4 MB, the same sealed corpus the series' spine cites. **Generic corpus:** hard-prompt outputs matched to the sovereign accepted-sample count. Both flow through an identical sampler pipeline at Ttrain = 2.0; H5 isolates the single intentional difference: the prompt pool. Base-model provenance (decision D-007): training and inference use the *same upstream checkpoint* in two representations: bf16 safetensors for training, Q8\_0 GGUF (9.53 GB, byte-size-verified against the upstream listing) for inference, with the LoRA applied at serve time onto the frozen base; no base-weight modification occurs. Tokenizer-parity and chat-template-parity checks run before any training compute.

#### 3.6 LoRA training: and fitting a 19 GB model into 12 GB

Unsloth SFT, identical hyperparameters per arm: r = 16, α = 32, lr = 1e-4, 1 epoch, seed 20260414, seven projection-matrix target modules, batch 2 × grad-accum 8, back-to-back wall clock, isolated training virtualenv. The bf16 footprint (\~19 GB) exceeds the 12 GB card; four adaptations, applied identically to both arms, preserving the H5 invariant, close the gap: (1) **QLoRA 4-bit quantization** of the base (\~5.5 GB); (2) **sequence-length reduction** 4096 → 1024, covering 98.7% of samples without truncation per corpus token-length analysis; (3) **vision-encoder CPU offload** (\~600 MB reclaimed, the VL encoder is unused in text-only SFT); (4) a **fused cross-entropy chunk-allocator patch**: Unsloth memoizes VRAM-available at model-load time, when memory is fully consumed, so later steps see a stale near-zero value and abort: the patch re-queries free memory per call and floors the chunk budget, changing chunking granularity only, never the loss mathematics. The root pressure is Qwen3.5's 248,320-token vocabulary: the per-batch logits tensor must be subdivided to fit free VRAM.

**Wall-clock reality (D-008).** Observed mean 850.6 s/step across 125 steps, an 11.4× gap over the dry-run micro-batch, attributed to the 248K-vocab chunk allocator under 151 MB free VRAM, an sm\_120 FlashAttention-2 regression forcing the Xformers fallback, per-layer dequantization overhead, and 96.5% VRAM saturation blocking async memory operations. GPU temperature held at 45 °C at 100% utilization, compute-density-bound, not thermal-bound. **Phase-1 (sovereign) sealed** 2026-04-17T07:58:32Z: 106,324 s (≈ 29.5 h), final loss 0.6077, zero step failures / OOM / thermal events, adapter SHA-256 f9844d43…b310, landing at 98.45% of the revised D-008 ceiling. **Phase-2 (generic) sealed** 2026-04-18T07:26:06Z: 75,325 s, paired-seal discipline (D-009) closed at byte-identity 5/7: sovereign-then-generic, hot-snapshot-then-launch, no reboot between arms, inter-arm gap logged as an acknowledged confound. H1–H6 inference was deferred until both arms sealed. The 29.5 h sustained QLoRA at 96% VRAM saturation and 45 °C peak on a consumer Blackwell card is additionally documented as a detachable substrate benchmark (memory-bandwidth-bound signature: 43 W sustained against a 250 W rating).

#### 3.7–3.8 Evaluation and statistics

Teval sweep {1.0, 1.1, 1.2, 1.5, 2.0}, best-Teval selection rule pre-registered for H1. The frozen DossierValidator scores parse, schema conformance, IOC/MITRE/remedial checks, and a BLUF length floor. Statistics: H1 exact McNemar (two-sided); H2 one-sided McNemar with a −2 pp materiality floor; H3 Wilcoxon signed-rank with a 0.95 hard retention floor; H4 descriptive KL thresholds (mean ≥ 0.02 or max ≥ 1.0 → reject); H5 Wilson 95% CI non-overlap + bootstrap ratio CI (1,000 resamples), ratio ≥ 1.5 floor; Holm–Bonferroni across the family; H6 longitudinal.

**Sampler–filter decoupling (D-001, logged pre-Gate-B).** The Gate-A baseline produced 0/190 on the frozen validator; using that validator as the sampler-side filter would make the ≥ 2,000-sample target unreachable. A distinct SamplerFilter was therefore introduced for data curation, accepting the model's native emission shapes and rejecting only structurally-broken output, while the Gate-D validator remained frozen as the scoring instrument. The decoupling is grounded in Apple SSD's own no-filter ablation (lifts survive complete filter removal), the "filter for broken, not for perfect" practitioner consensus, and the canonical calibration guardrail of separating training-pipeline thresholds from test metrics. Pre-registration scope was checked: the sampler filter is not in the frozen set; pipeline identity between arms is preserved; the sealed self-hash is unchanged. Predicted failure modes each map to a pre-registered guard: hedge-stripping to the H3 hard floor, KL drift to H4, malformed-output lock-in to the frozen strict validator (an SSD variant that merely clones the relaxed shape fails H1 honestly).

### 4. Results

#### 4.1 Baseline: the scientifically-ideal zero

Gate-A completed 2026-04-14: 950 prompt × Teval calls, 6 h 9 m, zero HTTP errors, zero serving-path crashes (the RCA companion's patched substrate holding at 3,404 cumulative crash-free requests). **The unconstrained orchestrator scored 0/190 on the frozen validator at every temperature.** The failure modes decompose cleanly: BLUF emitted as list rather than string (142), MITRE IDs annotated with descriptions (100+), empty remedial lists (64), JSON hygiene failures (54), typed-object IOCs (40). Content is semantically on-topic and CTI-faithful; *schema compliance* is not: the ideal negative baseline for the Apple thesis, since self-distillation under constrained decoding is precisely the recipe for tightening format compliance without new content. The 0% floor leaves full headroom, and the validator stays frozen per pre-registration immutability. Hedging baseline for the H3 denominator: hedge-sum 33–44 across the temperature grid; the retention ratio must stay ≥ 0.95. KL-drift baseline: 100 verbatim completions retained for pairwise computation at Gate-D. All artifacts sealed under a manifest with an empty deviations file.

#### 4.2 The sampler trajectory: failure → decision → pivot → fix → falsification → ratification

The path from an 11% accept rate to the sealed 2,000-sample corpus is the paper's methodological heart, logged as decisions D-001 → D-003.2 *before* the compute each affected was committed.

| Stage                             | Decision                                                       | Constraint      | Accept rate              | Accepts / h | Wall-clock to 2,000      |
| --------------------------------- | -------------------------------------------------------------- | --------------- | ------------------------ | ----------- | ------------------------ |
| v1 (sealed)                       | D-001 relaxed filter                                           | none            | 11.16%                   | 33          | > 60 h, intractable      |
| v2 (probe)                        | D-002 schema-shape checks dropped                              | none            | 26.6%                    | 134         | \~14.9 h, over budget    |
| v2 + schema (mini-probe)          | D-003 JSON-Schema at the sampler                               | logit-level     | **100.0%** (20/20)       | **1,218**   | \~1.6 h, under budget    |
| Tight schema (D-003.1)            | minLength + MITRE pattern in-grammar; V1 verb check at runtime | logit + runtime | 76.0%                    | 684         | \~2.9 h, under budget    |
| Verb-pattern in grammar (D-003.2) | 26-verb whitelist as logit mask                                | logit only      | ABORTED 20/200           | –           | **hypothesis falsified** |
| **Canonical run**                 | **D-003.1 ratified**                                           | logit + runtime | **81.57%** (2,000/2,452) | \~745       | 2 h 41 m actual          |

***Table 2.** The sampler trajectory. Every stage logged pre-compute; the falsified D-003.2 preserved with its abort seal.*

**The shape-vs-content gap (D-003 → D-003.1).** The full D-003 probe sustained 95.5% acceptance and confirmed that logit-level schema constraint eliminates parse failures structurally, and simultaneously surfaced the trade-off: the loose schema constrained *shape* but not *content*, admitting tactic-IDs in the MITRE field and 96-character minimum-viable dossiers. D-003.1 tightened the schema with content constraints expressible in the GBNF converter (string minLengths; the ^T\[0-9]{4}(.\[0-9]{3})?$ MITRE pattern) while reverting the runtime filter to the stricter V1, whose action-verb requirement no JSON-Schema dialect can express. Both schema-side predictions landed at zero; the residual rejects concentrated in a single class (remedial items lacking one of 26 canonical action verbs), setting up the natural next question.

**The falsification (D-003.2).** Pushing that verb whitelist into the grammar itself (predicted, with literature support, to lift acceptance to 95–99%) instead produced a decode pathology: tokens pinned at the generation ceiling, near-zero acceptance, decode-wall blowout. Aborted at emission 20/200 under the pre-declared reversal criterion; the hypothesis is **falsified** and the aborted run preserved under seal. The root-cause interpretation, anchored in the tokenizer-alignment literature: BPE tokenization splits multi-character verbs across sub-token boundaries that do not align with GBNF alternation, and the 52-way alternation × minLength × wildcard composite, though token-by-token within the model's training distribution, is **off-distribution at the compositional level** for the model's learned JSON prior. The constraint-distribution alignment principle is violated at composition even while satisfied per token. **D-003.1 was re-ratified as canonical**, and the sequence reframed as the pilot's first standalone contribution: a published ablation isolating the training-distribution-superset boundary for constrained decoding (*token-level alignment ≠ compositional alignment*) with the practitioner design rule: keep per-item patterns structural; enforce lexicon semantics at the runtime filter, not the logit mask.

**The canonical seal.** 2,000 accepts / 2,452 emissions / 81.57% / zero crashes / 2 h 41 m: a 36× throughput lift over v1. All three structural-impossibility predictions held at scale (zero parse failures, zero MITRE-format rejects, zero size-floor rejects); all 452 rejects came from the one runtime-lane class, exactly the division of labour the falsification had demarcated.

**Stratification: per-prompt hardness, not time, not iteration.** A mid-run accept-rate climb prompted a stratified test of three observational hypotheses. Within-prompt draw-index rates are non-monotonic with a 2.6 pp spread: *iteration-improves falsified* (and re-falsified on all four arms). Wallclock deciles and prompt-visit-order deciles are numerically collinear under the seed-deterministic harness, so the apparent temporal trend *is* the prompt-ordering trend: the first \~62 prompts sit 10–12 pp below the run mean, per-prompt stdev 0.20 across 613 prompts: **per-prompt hardness is the dominant variance component**. Because prompt order is shared across arms by construction, the hard-prompt cluster is a common offset that cancels in the between-arm H5 contrast.

#### 4.2.11–4.2.13 The four-arm contrast, sealed

| Arm                                                 | Accept rate | Emissions → 2,000      | Elapsed  | First-62 cluster                  |
| --------------------------------------------------- | ----------- | ---------------------- | -------- | --------------------------------- |
| Sham: schema only, filter off (Apple §4.4 ablation) | **99.80%**  | 2,004                  | 2 h 12 m | harder (−1.6 pp)                  |
| Canonical: schema + V1, sovereign corpus            | 81.57%      | 2,452                  | 2 h 41 m | harder (−14 pp)                   |
| Generic: schema + V1, generic corpus (H5 pair)      | 80.52%      | 2,484                  | 2 h 42 m | harder (−12.9 pp)                 |
| Strict: frozen validator as filter (ablation)       | 35.03%      | 1,282 in 4 h (timeout) | 4 h 00 m | **easier (+5.25 pp), sign flips** |

***Table 3.** All four arms sealed with SHA-256 artifact envelopes. The strict arm's timeout is itself the finding: validator-as-filter is non-convergent in budget, vindicating the D-001 decoupling.*

Three findings elevate to publication grade. **(1) The Apple no-filter ablation reproduces on a sovereign corpus at scale:** 99.80% acceptance under schema-only gating, with the sham arm's 18.71% observation-mode filter pressure the *unbiased* per-emission measure (the canonical arm's 24.36% is retry-inflated by the accept-target loop concentrating on harder prompts). Net V1 contribution to corpus composition: 18.23 pp. **(2) Gate-type × prompt-cluster interaction: a sign-flip.** The three schema-gate arms agree that the first-62 cluster is *harder*; the one semantic-gate arm flips the sign, finding it *easier*. Four arms, two corpora, three filter configurations, one sign-flip: **prompt difficulty is gate-type-dependent, not intrinsic**, a cross-corpus replicate not anticipated by either anchor paper, mechanistically consistent with constrained decoding distorting the generative distribution non-uniformly. **(3) H5 sampler-level parity:** sovereign vs generic accept rates differ by −1.05 pp: the "sovereign clears the filter more readily" alternative is refuted at full n, so whatever sovereign advantage exists must live in *what is inside* accepted samples, not how many pass. A residual defect candidate (3 of 2,004 rows with passing verdicts but kept=false) is logged unexplained, not material to aggregates, not hidden.

#### 4.3 Training telemetry

Gate-C paired QLoRA: **sovereign** sealed at 106,324 s (29.5 h, 98.45% of ceiling); **generic** sealed at 75,325 s (20.9 h, 69.75%). Paired-seal invariant verified at byte-identity 5/7 bundle files. The wall-clock asymmetry reflects the sovereign corpus's curriculum density (longer-form, higher tokens-per-sample), not a substrate hazard: the thermal envelope held throughout both arms with zero throttle events, and Gate-D closed with zero production-service impact.

#### 4.4–4.5 H1 hard-bank lift and H2 easy-bank stability: all four accepted

| Arm · instrument                           | Baseline                    | Variant         | Δ (pp)            | Holm-adj. p | Verdict     |
| ------------------------------------------ | --------------------------- | --------------- | ----------------- | ----------- | ----------- |
| H1 sovereign · frozen validator            | 0/100                       | 24/100          | **+24.00**        | 4.8e-07     | accepted    |
| H1 generic · frozen validator              | 0/100                       | 52/100          | **+52.00**        | 3.7e-12     | accepted    |
| H1 sovereign · schema-adjusted (secondary) | 21/100                      | 88/100          | +67.00            | –           | descriptive |
| H1 generic · schema-adjusted (secondary)   | 19/100                      | 87/100          | +68.00            | –           | descriptive |
| H2 sovereign / generic · easy bank         | no regression on either arm | +26.00 / +56.00 | 7.3e-04 / 3.7e-08 | accepted ×2 |             |

***Table 4.** H1/H2. The primary-to-secondary gap quantifies the validator's schema-collision penalty: sovereign −43 pp (list-form BLUF in \~64% of outputs), generic −16 pp (\~35%): an asymmetry traceable to corpus formatting habits, load-bearing for §4.8.*

#### 4.6 H3: the survivor finding

| Arm               | Hedge-sum (baseline 194) | Retention ratio | Safety-stop (< 0.85)? | Verdict                                   |
| ----------------- | ------------------------ | --------------- | --------------------- | ----------------------------------------- |
| **ssd-sovereign** | 189                      | **0.9742**      | No                    | accepted, hedging within 2.6% of baseline |
| ssd-generic       | 157                      | **0.8093**      | YES                   | rejected, 19.1% hedging loss, automatic   |

***Table 5.** The asymmetry the series' spine cites. Unanticipated by the sampler-level parity of §4.2.13: H3 tests the content of accepted completions, and the generic corpus transmitted measurably less hedge-calibrated language under matched training pressure.*

#### 4.7–4.9 H4, H5, H6

**H4 (KL drift): both rejected, marginally and symmetrically.** kl\_mean 0.0278 / 0.0273 against a 0.02 threshold; both max drifts ≈ 0.18 against a 1.0 ceiling: no catastrophic-drift event. The symmetric rejection suggests the threshold is calibrated tighter than this LoRA budget produces on this base regardless of corpus; the threshold is retained, not patched, with recalibration flagged as post-hoc follow-up.

**H5 (the keystone): PRIMARY FALSIFIED.** On the frozen validator, the sovereign/generic lift ratio is **0.4615**: generic outperforms sovereign by \~2.2× with disjoint Wilson 95% CIs; the pre-registered ≥ 1.5 acceptance threshold is rejected in the wrong direction at scholarly significance. The pre-declared schema-adjusted secondary moves the same ratio to **0.9853** with overlapping CIs, **empirical near-parity**: the primary gap is a validator-specification collision (the string-typed BLUF field penalizing the sovereign corpus's list-formatted doctrine style 43 pp vs 16 pp), not a genuine specialization failure. A less-disciplined pipeline would either patch the validator mid-analysis (forfeiting the frozen instrument) or report 0.46 as the true signal ("sovereignty refuted"): neither is correct, and the discipline produces the honest compound conclusion: *the falsification is real, the near-parity is real, the schema collision is real, and the failure to anticipate it pre-freeze is real.* All four stay in the record. The pre-registration's value is precisely that the investigator could not recover H5 by retroactive schema-patching; a schema-richer validator is a different H5 test for the next cycle, not a rescue of this one. (A bootstrap-CI software defect in the verdict machinery is disclosed as non-falsifying: Wilson CIs establish the disjoint-CI property on their own.) **H6 (compounding): deferred**, this pilot records the cycle-C₀ datum only; the Pareek et al. LLM-scale analogue test awaits post-deployment instrumentation, with its falsification criterion (three successive cycles failing monotone improvement) frozen unchanged.

#### 4.10–4.11 The instrument as the contribution; the two-axis case

The primary contribution of the pilot is not which hypotheses landed where: **the instrument is the PI+AI co-authored discipline**, and the Gate-D/E closure exercised it three ways: pre-registration retention under a negative result (no validator patch, no hypothesis migration); honest disclosure of a specification gap surfaced mid-pipeline (declared in both arms' ledgers and seals, disclosed additively); and the AI collaborator's own failure cascade during the closure window, documented in the same paper as the falsification it accompanied, because the eval data is byte-identical to the pipeline's intent (the failures were pre- and post-eval, not in-eval). Every artifact of the discipline, including the cascade, is itself curated training signal for the successor models: the pilot's run is training data for future pilots.

The **two-axis case** that closes the paper's argument: Axis A (quantitative, §4.6), under matched specialization pressure, the sovereign arm preserves calibrated hedging while the generic arm collapses into safety-stop. Axis B (qualitative, the companion natural experiment's §10 catalogue), a non-specialized frontier model under sustained production load loses operator-corpus-rule compliance across six documented failure classes that in-weights specialization on the operator's correction corpus is architecturally positioned to address. Neither axis alone is the case; the pair is, and the pair is what the series' methodology reference formalizes as **the Sovereign Pair**: upstream-deterministic gates where determinism is structurally possible, in-weights sovereign specialization where compliance must survive distribution-shaping under load. This pilot is the first peer-review-grade empirical instance of the Pair's second half.

### 5. Discussion

#### 5.1 Relative to Apple SSD

Both arms replicate the positive-result class (both Holm-adjusted p < 1e-6); the +12.9 pp headline is not an upper bound but a point on a corpus-dependent curve. What the pilot adds is a falsification surface for a *stronger* claim than Apple tested (sovereign Pareto dominance), and that stronger claim falls to evidence cleanly: both variants lift; the asymmetric magnitude claim is what dies. **5.2 Relative to the paradox paper.** The hedging-erosion mechanism is *reproduced on the generic arm and averted on the sovereign arm* under matched pressure, to our knowledge the first pre-registered LLM-scale measurement of that failure mode under controlled corpus substitution. The two published results are not in tension; they are one axis measured on two substrates, observable when the substrates are pre-registered as contrastive arms. **5.3 The refined thesis.** Corpus sovereignty is not the mediating variable for raw lift magnitude on a frozen-string validator; it is the mediating variable for **calibration-property retention under training pressure**. That narrower, better-supported claim is staked as the survivor thesis for the next pre-registration cycle, explicitly not a rescue of H5, which remains falsified on the record. **5.4 Compounding.** The pilot contributes a start-point datum, not a test; the falsification criterion survives unchanged. **5.5 Architecture implications.** The three-tier sovereignty architecture becomes empirically motivated: in-weights specialization is the correct home for calibration properties (retrieval over a generic-distilled adapter could not produce 0.974 retention (the generic adapter had already shed the pattern); retrieval is the correct home for coverage breadth (the lift-magnitude asymmetry is a coverage effect); and allowlisted promotions are the correct boundary between them, because the paradox paper's failure mode is exactly the class of pattern that looks like content to a validator and behaves like calibration erosion after training.

### 6. Limitations

Single seed, single substrate, single base model: all CIs are within-seed, within-substrate; cross-seed meta-analysis is deferred by pre-registered scope. The validator-specification gap on the BLUF field is the dominant driver of the primary/secondary ratio swing and is a design item for the next cycle, not a patch to this one; neither readout is claimed as "the true result": both are in the record. The hedge-phrase regex is a conservative lower bound on hedging retention (semantic hedging without a listed token goes uncounted (the reported asymmetry would likely widen, not narrow, under richer measurement). The H4 threshold may be miscalibrated for this LoRA budget; it is retained for the pilot. H6 is untested here. And the operator's motivational bias on H5 (the falsified hypothesis was his first-order thesis) is mitigated by the sealed pre-registration, frozen validator, pre-registered secondary branch, two-person audit, and open artifacts; *the falsification of H5 on the primary instrument is itself the evidence the mitigations held.*

### 7. Conflicts, Contributions, Availability, Acknowledgments

OSINTelligence LLC is the operator's company and P7 outputs bear on its roadmap; mitigations as above, with Anthropic API credits supporting the AI collaborator and Anthropic having no pre-seal review. Author contributions follow the pre-registration's disclosure posture (equal-authorship transparency; the human PI is sole author-of-record for experimental decisions). All code, manifests, hashes, and telemetry are committed to the sovereign repository; corpus and eval-bank hashes are quoted in §3; adapter weights are held as sovereign IP in v1, reproducible from corpus. The Michael Kaplan legacy (CTIAM transcripts entrusted to the operator) grounds the doctrine track of the sovereign corpus.

### 8. Conclusion

**Plain-language verdict, tier zero.** The pre-registered Pareto-dominance thesis is falsified on the frozen validator: sovereign-corpus self-distillation did *not* out-specialize generic-corpus self-distillation on hard-dossier lift (ratio 0.46, CIs disjoint, wrong direction at significance). **A novel survivor finding replaces it:** on a calibration axis nobody pre-registered as primary, the sovereign arm preserves calibrated hedging (0.974, within 2.6% of baseline) while the generic arm collapses below the safety-stop floor (0.809, automatic rejection). On the axis the original thesis implicitly predicted, sovereignty is doing exactly its work; on the lift axis it is not. The refined survivor thesis (*sovereign-corpus specialization preserves calibration properties that generic specialization collapses*) is the lead finding the pilot earned, staked as primary for the next cycle with a schema-richer validator pre-registered from intake.

**What the pilot proves about the instrument.** The discipline held through five consecutive stress tests: a constrained-decoding hypothesis falsified pre-data and its abort preserved; the keystone hypothesis falsified at the verdict gate with no patch, no migration, no rescue (the pre-registration self-hash surviving unchanged on disk and in the log); a specification collision disclosed across three tiers and answered additively rather than substitutively; a mid-pipeline deviation logged as deviation; and the collaborator's own failure cascade documented in the same paper as the falsification it accompanied. **What the falsified primary buys** is a public proof that the discipline is non-rhetorical: a solo operator on a consumer GPU, running an end-to-end pre-registered experiment with a SHA-anchored validator freeze, was willing to call his first-order thesis falsified in the same paper that documents the survivor. The audience inherits a complete record: hypothesis, seal, ablation ladder, paired training, verdict, cascade, and the architectural anchor the two axes jointly motivate. Every step reproducible from a recipe. The discipline travels; that is the pilot's lasting deliverable.

### System Update: July 2026 (appended; the sealed body above is unmodified)

The falsification-retention lineage this pilot established has compounded. Its direct successor, the matched-corpus chain experiment (final-sealed 2026-07-14), tested sovereign-corpus calibration at the base-model layer against the pre-registered mismatch control, and concluded at **composite FAIL, carried in full**: H1 FAIL −11.25 pp (the pre-registered forgetting guard firing exactly as designed), H2 PASS +18.18 pp, H3 inconclusive (+2.53 pp, n.s.). The operator-ruled conclusion is a reframe, not a data dispute: **the role-flavored corpus had been mis-layered into the base slot** (role flavor belongs in the LoRA/behavioral layers), and the incumbent broad-reasoning substrate stays precisely because the guard caught the base going too narrow. The sealed framing is explicit that this is not a "sovereign-corpus loss."

Two artifacts of doctrine followed. First, the **corpus-flywheel realignment** (operator-ruled 2026-07-11/12): one sha-sealed sovereign corpus draw, cited by hash at every layer of the compression-and-specialization chain (prune calibration, recovery fine-tune, imatrix, self-distillation, LoRA, drafter head), with a mandatory corpus-identity clause in every future pre-registration (the clause that would have caught the mis-layering at authoring time). Second, the **head-to-head redraw** this paper's measurement-axes finding anticipated is now a registered design authority (A1–A4 arm matrix; tool-calls-to-objective as the primary anchor; tool-denial ablation; restraint ladder), awaiting its own pre-registration.

Provenance: the sealed matched-corpus chain experiment (pre-registration, a 236-green experiment ledger, the final seal §3 through §4, and post-run integrity), the July 2026 pruning-arc roadmap §2 addendum, and the registered head-to-head design authority. Append-only; the pilot's sealed record above is unmodified.

***

*The Sovereign Stack · Corpus-Sovereign Self-Distillation · Chapter 18 · Part IV · v1.0.0 · License CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*

**Citation (preferred):** Kistner, J. (2026). *Corpus-Sovereign Self-Distillation: Pre-Registered Tests of Pareto Dominance and Iterative Compounding on a Sovereign-Hardware LoRA Substrate*, version 1.0.0. OSINTelligence LLC research whitepaper (pre-registered pilot). Cited in-series by title.

*The reference list and provenance follow as a sub-page of this chapter.*
