> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/part-iv-the-evidence-what-worked/16-sovereign-domain-pruning/the-quick-version.md).

# The quick version

The chapter in minutes: the video explainer, the audio deep dive, the one-view infographic, chapter notes, and a self-test quiz on the Sovereign Domain Pruning.

**Continue the tour →** [Next: 17 · Sovereign Big-Model Compression, the quick version](/osintelligence/part-iv-the-evidence-what-worked/17-sovereign-big-model-compression/the-quick-version.md)

The short version of Chapter 16, three ways: the video walks the argument in a few minutes, the deep dive talks it through at a listening pace, and the infographic holds the whole chapter in one view. The full result, with the four-file integrity contract, the mismatch control, and the two ancillary failures kept as the honesty architecture, lives in the chapter itself: [16 · Sovereign Domain Pruning](/osintelligence/part-iv-the-evidence-what-worked/16-sovereign-domain-pruning.md).

{% embed url="<https://youtu.be/ziAUJcnTYDM>" %}

**The deep dive.** A podcast-style audio conversation about this chapter: two AI hosts walk through the argument, the incidents behind it, and what it means, at a listening pace. Generated in Google's Gemini LM (formerly NotebookLM) from the chapter itself; the link opens the audio on Google's site.

{% embed url="<https://notebook.google.com/notebook/28f1f15f-0f00-4353-92ca-6f6746faeb19/artifact/cf3a4b28-8e8f-4b25-be03-6cc8888c1852?utm_source=nlm_web_share&utm_medium=google_oo&utm_campaign=art_share_1&utm_content=&utm_smc=nlm_web_share_google_oo_art_share_1>\_" %}

*The conversation is AI-generated: an interpretation of the chapter, not the chapter. It can compress, paraphrase, or get details wrong. The written chapter is the authoritative, canonical source:* [*16 · Sovereign Domain Pruning*](/osintelligence/part-iv-the-evidence-what-worked/16-sovereign-domain-pruning.md)*.*

***

![The Sovereign Domain Pruning, the chapter in one view.](https://137900913-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2Fxx2bv6VR9dSJDJ9HsDER%2Fuploads%2FL8tXWVBvrMfR5YmIUCru%2Fpruning-infographic.png?alt=media)

***

### Chapter notes

Section-by-section notes in two registers: the technical note on the left, the same idea in plain language on the right. Every row is one idea, so you can read straight across from one register to the other. The technical terms stay visible in the plain column on purpose; they are the vocabulary worth keeping.

#### 1. Abstract

**The point:** cut 197 experts out of a 35B model against the operator's own corpus, under a sealed test designed to kill the result, and the model got slightly better.

| The technical note                                                                                                                                                                                                                                                                                                                         | In plain language                                                                                                                                                                                                                                                                                 |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Destructive expert-pruning of 197 of 10,240 experts (\~1.9%) at blocks {20, 21, 25} in a 35B-A3B MoE against a sovereign CTI corpus, with a cryptographically pre-registered Gate-B verdict GREEN; sovereign held-out perplexity Δ −0.0519 (−0.41% relative, 2.4× under threshold; pruned slightly better than baseline).                  | A mixture-of-experts model is a committee of tiny specialists (**experts**); most never speak on this operator's workload. This chapter permanently deletes 197 of them (**destructive pruning**), chosen against the operator's own data, and the model's quality on that data went slightly UP. |
| The mandatory long-form gate (H6, anchored against the two published results most likely to sink it) lands at pass-rate ratio 0.9826, decisive GREEN at the ≥0.90 floor; combined Phase-2 wall-clock ≈68 h 20 m, 0 errors, on the RTX 5070 substrate.                                                                                      | The test most likely to fail was made mandatory on purpose (**the long-form gate**): published research says pruned models lose their long chain-of-thought reasoning. Here that reasoning held at 98.26% of baseline, a clear pass, across a three-day error-free run on the desk PC.            |
| Integrity is a four-file γ.1 + γ.2 re-seal with a 92-file cascade manifest; the Phase-B companion validation closed all three audit surfaces: two ancillary hypotheses failed as pre-registered and are reported as falsifications, one passed with margin, and the Phase-2 GREEN emerged unfalsified across all three adversarial lenses. | The paperwork is cryptographic (**the four-file contract**), and the result was then attacked from three angles by a follow-up audit (**companion validation**). Two of the audit's own side-hypotheses failed and are printed as failures; the main result survived all three attacks.           |

#### 2. Introduction

**The point:** why prune at all, and the two published landmines the experiment was aimed straight at.

| The technical note                                                                                                                                                                                                                                                                                                                                                                                                      | In plain language                                                                                                                                                                                                                                                                                                             |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| The 9B orchestrator is a 12 GB-VRAM compromise; the 35B architect has 4× the parameters and stronger reasoning but does not fit as primary; pruning the experts inactive on the sovereign workload produces a model that started with 35B knowledge, trimmed to what matters, fitting consumer VRAM: the imatrix philosophy extended from quantization into architecture.                                               | The motive (**breaking the VRAM ceiling**): the best model does not fit the card. Instead of settling for the smaller one, delete the parts of the big one this workload never uses. The previous chapter tuned compression with home data; this one re-architects with it.                                                   |
| The threat model is explicit: Wang et al. establish that layer-pruned models can lose test-time-scaling chain-of-thought even when perplexity preserves; Su et al. establish that removing 3 of 6,144 super-experts causes catastrophic repetition collapse; H6 is MANDATORY-falsifiable against the first, and a pre-registered super-expert cross-reference guards the second: 0 of 15 candidates in the removal set. | The two known ways this dies were built into the test (**the threat model**): pruned models that look fine but lose deep reasoning, and a handful of irreplaceable **super-experts** whose removal breaks everything. The gate targets the first; a pre-registered check confirmed none of the second were touched (0 of 15). |
| Contributions include the mixed-backend eval-bank under shared pre-registration, the hash-first integrity discipline, the document-as-found scope reduction (AMENDMENT\_4), and the four-point substrate envelope (0.35 / 2.56 / 8.91 s-per-row / 26.14 tok/s), offered as peer-review-novel consumer-hardware measurements.                                                                                            | The side contributions are method, not scores: a two-engine evaluation design declared in advance, a hash-before-fix discipline, a mid-run scope cut handled by the book, and a four-point speed map of what this class of hardware can actually do (**the substrate envelope**).                                             |

#### 3. Methods

**The point:** the screen that chose the 197, and the sealed paperwork that froze the rules before the data existed.

| The technical note                                                                                                                                                                                                                                                                                                | In plain language                                                                                                                                                                                                                                                        |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| The pre-registered conjoint screen: Block-Importance p05 < 0.05 + expert frequency < 0.5% + Gini > 0.55, with a max-activation ≥ 0.4 exclusion; blocks {20, 21, 25}, 72 + 68 + 57 = 197 removed; mechanism: router rows masked to −10000.0, shared-expert hashes byte-identical pre/post.                         | Four criteria had to agree before an expert was cut (**the conjoint screen**): the block barely matters, the expert rarely fires, its usage is lopsided, and it never spikes. The cut is a routing mask, and the untouchable shared experts are hash-verified untouched. |
| The eval-bank spans five probes under one pre-registration (sovereign PPL at 0.20 weight, MMLU 0.10, GSM8K 0.15, HumanEval 0.10, long-form TTS 0.45), split across two backends by feasibility: HF-bf16 for forward-pass probes, production-matched Q6\_K llama-server for generation, with the rationale sealed. | Five tests, weighted in advance, with the heaviest weight on the one most likely to fail (**the eval-bank**). Two engines share the work because the slow engine would take 340 hours for one test; the split was declared before any data, not improvised after.        |
| The four-file integrity contract: PRE\_REGISTRATION + AMENDMENT\_2 + AMENDMENT\_3 sealed at γ.1 (2026-04-23, unchanged at γ.2), AMENDMENT\_4 sealed at the γ.2 re-seal (2026-04-25), with a 92-file cascade manifest.                                                                                             | The rulebook is four fingerprinted files (**γ.1 and γ.2**): the original rules plus three amendments, each hashed at its moment, none editable afterward, with all 92 result files indexed under one manifest.                                                           |

#### 4–5. Results, Phase-1 and Phase-2

**The point:** the sealed scoreboard: profiling clean, perplexity better, long-form reasoning held, one honest regression flagged and routed to audit.

| The technical note                                                                                                                                                                                                                                                                                                          | In plain language                                                                                                                                                                                                                                                                               |
| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Phase-1 activation profiling: 26.54 h, 8,358/8,358 rows over 2.59M tokens, all gates green, 10.20 GiB VRAM peak, zero NaN/Inf; Phase-2 verdict GREEN at the γ.2 re-seal.                                                                                                                                                    | The measuring pass that chose the targets ran a day, clean (**Phase-1**), and the destructive experiment it fed returned green (**Phase-2**).                                                                                                                                                   |
| H5 primary: sovereign PPL 12.6418 → 12.5899 (−0.41% relative, pass by 2.4×, pruned better); H6 mandatory: AIME24 0.9453, MATH-500 1.0000, aggregate 0.9826 GREEN outside the CI-straddle band, with parity-or-better at middle budgets; Wang's degradation signature does not transfer to expert-only pruning at this rate. | The two headline gates (**H5 and H6**): quality on home data improved slightly with a fifth fewer experts in three blocks, and the deep-reasoning test the literature predicted would fail instead held at 98–100% across tasks. The published failure mode belongs to a different kind of cut. |
| Secondaries: MMLU −1.25 pp (within the power floor), GSM8K +1.74 pp (pruned BETTER), HumanEval −5.49 pp (the largest regression, marginal at N=164 with ±\~7 pp CI, flagged as the Phase-B primary audit case); weighted Δ −1.20 pp.                                                                                        | The full scoreboard is printed with its one sore spot (**the HumanEval flag**): general knowledge flat, math better, code completion down by an amount inside the test's own noise band, explicitly routed to the follow-up audit rather than explained away.                                   |
| The substrate envelope: HF-bf16 sampling <0.35 tok/s, greedy 2.56, forward-pass 8.91 s/row, production Q6\_K 26.14 tok/s; combined 68 h 20 m at 0 errors, 28.5% of the operator's 240 h ceiling.                                                                                                                            | The speed map nobody publishes for consumer cards (**the four-point envelope**): the same model spans a 75× speed range depending on engine and mode, which is why the two-engine design was the only one that fit the time budget.                                                             |

#### 6. Discussion

**The point:** the disciplines the run minted: hash before fixing, document the deviation as found, and match the eval to the production artifact.

| The technical note                                                                                                                                                                                                                                                                                                                                                                                        | In plain language                                                                                                                                                                                                                                                                         |
| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| The Gate-B aggregator emitted a spurious RED from a schema-drift bug, discovered in a window where a code edit was about to touch a verdict-informing file; the operator's directive, "Hash first, don't contaminate the experiment," sealed the three governing files \~30 s BEFORE the fix landed, eliminating the post-hoc-amendment attack surface; all buggy versions preserved as forensic anchors. | The best story in the chapter (**hash first**): a scoring bug appeared at the worst possible moment, and before fixing it, the rules were fingerprinted so nobody could ever claim the rules were bent to fit the fix. The buggy outputs were kept, labeled, as evidence of the sequence. |
| A gated-dataset authentication failure mid-run triggered the pre-existing symmetric fail-soft: both arms reduced identically to 30 cells / 1,100 generations, three pre-specified equalities verified, AMENDMENT\_4 ratifying the reduction at the re-seal, with the Wilson-CI widening (\~1.22×) still leaving the verdict decisive.                                                                     | A mid-run failure shrank the test (**document-as-found**): one dataset would not authenticate, both arms were cut back by exactly the same amount, the cut was ratified as a sealed amendment, and the math confirms the verdict survives the smaller sample.                             |
| The Q6\_K configuration used for generation probes is byte-identical to the production architect preset (backend-faithful: what is evaluated is what is deployed); the bf16 paths are evaluation-only, with the precision-gap question explicitly deferred to Phase-B Component A.                                                                                                                        | The evaluation deliberately tests the exact artifact production runs (**backend-faithful**), and the one gap that leaves, whether the compressed and full-precision versions agree, is named and handed to the follow-up audit instead of assumed away.                                   |

#### 7–8. Limitations and Reproducibility

**The point:** the honesty ledger, updated by the audit: each open item marked CLOSED with its verdict, including the two that closed as failures.

| The technical note                                                                                                                                                                                                                                                                                                                                              | In plain language                                                                                                                                                                                                        |
| --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| The limitations section carries the Phase-B closures in place: the random-ablation control CLOSED as a mechanical FAIL revealing selection-invariance; the bf16 cross-validation CLOSED as a conjoint FAIL (Spearman 0.204) with the Bland-Altman axis passing (LoA 0.030 vs ≤0.10); GPQA recovery CLOSED as PASS at 0.9866 inside the pre-registered interval. | The limitations list is not static (**closed with verdicts**): each open question was later answered by the audit, and the answers are written into the list, including the two that came back "your hypothesis failed." |
| Reproducibility: hash-pinned eval sources, a local 2B judge (sovereign local-first), pinned artifact hashes for both GGUFs and the stratified subset, a reference rebuild seal, and the full decision ledger D-P8-001 through D-P8-027 archived with per-file SHA-256 sidecars.                                                                                 | Everything needed to redo it is fingerprinted (**the reproducibility envelope**), down to the grading model, which is itself a local model rather than a cloud judge.                                                    |

#### 9. Conclusions

**The point:** the verdict, the deployment, and the methodology extracted: companion-validation as inferential strategy.

| The technical note                                                                                                                                                                                                                                                                                       | In plain language                                                                                                                                                                                                               |
| -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| The pipeline validates at the pre-registered Gate-B: H5 at 2.4× margin, H6 decisive GREEN, weighted Δ −1.20 pp modest, HumanEval routed to audit; the pruned model is deployable as the Architect at production Q6\_K with sovereign CTI specialization.                                                 | The close is operational (**deployable**): the pruned model is not a demo, it is the production architect candidate, with its one soft spot under audit rather than under a rug.                                                |
| Seven peer-review-grade methodology contributions are extracted and companion-published, exercised simultaneously in this running experiment, with the full decisional provenance trail (including the spurious-RED and degenerate-merge outputs, preserved as forensic anchors) auditable at the forge. | The lasting export is method (**the seven contributions**): the disciplines this run minted under fire are published for reuse, and the embarrassing intermediate outputs are kept on purpose as proof the sequence was honest. |

#### Appendix §10. Phase-B Companion Validation

**The point:** the deepest water: the audit's own hypotheses failed, the failures were kept, and the failures are what prove the main result robust.

| The technical note                                                                                                                                                                                                                                                                                                                                                                                         | In plain language                                                                                                                                                                                                                                                                                                                                                                                                                      |
| ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Phase-B audits three surfaces around the primary claim under falsification-bar preservation contracts (no rescue by redefinition, no threshold relaxation post-data, no hypothesis migration); Phase-2 sealed bytes never change as a function of Phase-B outcome; all 7 cryptographic anchors re-verified MATCH at every sweep.                                                                           | The follow-up audit ran under handcuffs of its own (**the falsification-bar contracts**): it could not soften a failing test, move a goalpost, or touch the sealed result. Its job was to attack the main claim from three sides and report whatever happened.                                                                                                                                                                         |
| Component B, the random-197 ablation: the discriminative-screen hypothesis mechanically FAILED (random worse on only 1/4 probes), with the striking detail that HumanEval pass-count is byte-equal (115/164) across two disjoint 197-expert sets at 25.4% overlap; the screen's one material win is GSM8K (+4.09 pp); implication: the GREEN verdict is robust to selection method at this prune fraction. | The audit's centerpiece backfired beautifully (**the random-ablation control**): deleting 197 RANDOM experts worked almost as well as the carefully screened cut, and on the code test the two produced literally identical scores (**byte-equal, selection-invariant**). The clever screen only demonstrably earns its keep on math. That kills the audit's hypothesis and simultaneously proves the main result is not a lucky pick. |
| Component A, the bf16 cross-validation: the precision-gap conjoint FAILED on Spearman (0.204 vs ≥0.95) while the load-bearing Bland-Altman NLL axis PASSED with margin (LoA 0.030 nats vs ≤0.10); the Spearman failure is attributed to cell grains far below the \~250-example reliability floor; reported as falsification plus a separate positive fidelity finding, no rescue.                         | The precision audit split honestly (**FAIL plus a positive finding**): the rank-agreement test failed for a sample-size reason the paper names, while the substantive fidelity test passed comfortably. Both verdicts are printed as-is instead of blending into a soft pass.                                                                                                                                                          |
| Component C recovered the skipped dataset with authentication fixed: the full 45-cell grid ratio is 0.9866, inside the pre-registered PASS interval, confirming the reduced-scope verdict (+0.004 delta); the first merge attempt produced silently degenerate output, discovered before any verdict authorship, fixed at a named commit, the degenerate cascade preserved as the forensic anchor.         | The cut test was completed after the fact (**the recovered grid**) and confirmed the shortened version's verdict almost exactly. Even here the record keeps a wart: the first merge script silently produced garbage, was caught before anyone wrote a conclusion on it, and the garbage is archived with the fix.                                                                                                                     |
| §10.7 names the tooling-orchestration-verification-gap class: six silent-mismatch defects (typed-but-empty reads, no-write flags that wrote) discovered by post-hoc inspection of frozen artifacts, anchoring a planned property-tested harness; §10.8 tallies the \~219 h zero-event substrate record; §10.9 pre-registers five follow-up questions.                                                      | The audit also audited its own tools (**the verification gap**): six bugs that produce plausible-looking nothing were caught by inspecting frozen outputs, and the class now has a named harness planned. The run tally across everything: about 219 error-free hours on one consumer PC.                                                                                                                                              |

#### System Update: July 2026

**The point:** the append-only update: the method scaled to 122B and into prune-then-recover, with the failures carried at full fidelity.

| The technical note                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                               | In plain language                                                                                                                                                                                                                                                                                                                            |
| ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| At 122B: a REAP-style one-shot 25% prune (256 → 192 experts across 48 layers, 23.3 h observer pass, zero thermal events) produced a 186.2 GB checkpoint whose GGUF + matched-imatrix chain (29× calibration wall-clock reduction: 2 h 04 m vs 118 h) serves at 16.65 tok/s vs the unpruned 15.60 at two quantization tiers better fidelity; the sealed quality gate returned PASS (pooled-math +2.5 pp, GPQA-198 +1.01 pp, MATH-500 tie) and the model was promoted to the production research-brain preset; vision was re-enabled by extracting the multimodal projector, the pruned trunk needing zero rework. | The method graduated three sizes up (**the 122B arc**): a quarter of the giant research model deleted, the quality gate passed with the pruned model slightly ahead, and it now serves in production, faster than the original at better fidelity. Even its vision capability was restored afterward without touching the pruned core.       |
| At 35B: two deeper cut geometries (25%/31% → 26.61B/24.59B) extended the destructive-only method with prune-then-recover 200-step fine-tunes (val CE 1.4681 → 1.2100; 1.5113 → 1.2430), the training receipts doubling as the on-box-trainability proof: a 24.6B model trained at \~10.3 GiB on the 12 GB card.                                                                                                                                                                                                                                                                                                  | The second front cuts deeper and then heals (**prune-then-recover**): brief retraining recovers the deeper cuts' quality, and the receipts prove something bigger in passing: a 24.6-billion-parameter model being trained on the little 12 GB card.                                                                                         |
| The honest negatives carried at full fidelity: the fully-resident serving hypothesis FAILED (IQ4\_XS actuals 12.68/13.69 GiB spilled past dedicated VRAM, the WDDM dedicated-VRAM-lies spill signature itself a sealed finding), and a matched-corpus re-calibration experiment concluded at composite FAIL with an operator-ruled reframe: the role corpus was mis-layered into the base-model slot, and the pre-registered forgetting guard caught it.                                                                                                                                                         | And the update ends on two failures kept whole (**the honest negatives**): the deeper cuts still do not fully fit the card (with the misleading-VRAM-counter behavior sealed as its own finding), and one calibration experiment failed because the wrong corpus was loaded into the wrong slot, caught by the guard built for exactly that. |

***

### Test yourself

A short quiz on this chapter, generated in Google's Gemini LM (formerly NotebookLM) from the chapter itself. Work through the notes above first, then check what stuck; the link opens the quiz on Google's site.

{% embed url="<https://notebook.google.com/notebook/28f1f15f-0f00-4353-92ca-6f6746faeb19/artifact/1296b8b0-8957-4a35-a644-de253ec0c384?utm_source=nlm_web_share&utm_medium=google_oo&utm_campaign=art_share_1&utm_content=&utm_smc=nlm_web_share_google_oo_art_share_1>\_" %}

*The quiz is AI-generated: its questions and answer keys are an interpretation of the chapter, not the chapter, and can misstate a detail. Where a question and the text disagree, the written chapter is the authoritative, canonical source:* [*16 · Sovereign Domain Pruning*](/osintelligence/part-iv-the-evidence-what-worked/16-sovereign-domain-pruning.md)*.*
