> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/part-ii-the-discipline/8-sovereign-optimization-flywheel.md).

# 8 · Sovereign Optimization Flywheel

**Sovereign Optimization Flywheel: A Multiplicative-Compound Five-Axis Architecture for Single-Operator AI Deployment on Consumer Hardware**

*Chapter 8 · Part II: The Discipline · Evidence-backed · v1.0.0*

**Author:** Jamey Kistner, OSINTelligence LLC

**Keywords:** multiplicative compound · quantization · KV-cache compression · MoE pruning · in-weights specialization · multi-token prediction · reinforcing feedback · consumer hardware

> **A companion paper.** This is the loop the spine's Governor verifies: the self-improvement flywheel whose integrity *The Sovereign Triad*'s V2 check certifies between cycles. It synthesizes the series' evidence papers into one architecture (*Sovereign Imatrix Calibration* (L₁), *Sovereign Domain Pruning* (L₃), and the LoRA/self-distillation chain including *Sovereign CTI-NER* and *Corpus-Sovereign Self-Distillation* (L₄)) and formalizes why the layers compound multiplicatively rather than adding. §3.5 carries the one-sentence Sovereign Pair formulation the series quotes throughout. Cited in-series by title.
>
> **Status note.** Authored 2026-05-12; the §3.5 Pair-contract section added additively 2026-05-13 with the sealed body preserved. L₂ and L₅ are prediction-stage layers gated on upstream substrate; the maturity ledger in §7 states plainly which layers are sealed, deployed, and predicted.

> **What is new here.** The contribution is a multiplicative rather than additive account of single-operator optimization: five compression and serving axes (quantization calibration, KV-cache compression, expert pruning, in-weights specialization, multi-token prediction) whose ten pairwise cross-couplings are each argued to be monotonically reinforcing, so the portfolio gain is bounded by the product of the per-axis gains, not their sum. The load-bearing structural claim is a shared substrate: one in-distribution corpus measurement supplies both the imatrix calibration signal and the pruning-importance signal at two granularities, which is what makes the axes reinforce rather than merely coexist and is the direct answer to the data-moat critique that portfolios usually only add. Onto that sits the Sovereign Pair failure-mode contract, that in-weights compliance and mechanical gates are non-overlapping in failure mode (one degrades under load at the token layer, the other fires on pattern match at the dispatcher layer regardless of what the model remembers), and the observation that industry ships the probabilistic half of that pair alone.
>
> **Deepest water.** §3.2, the ten-pair cross-coupling matrix and the multiplicative bound it requires (all ten reinforcing is the maximally-reinforcing condition the argument stakes itself on); §3.5, the Pair failure-mode contract and the same-session six-gate-fire receipt; and §3.7 with §5.2, the recursive moat and why compounding, not raw data, is the moat. The honest register is §6: the five-axis compound envelopes are predicted from sealed single-axis results, not yet measured end to end, an architecture with its per-axis anchors sealed and the joint result reserved.

### Abstract

This paper formalizes the **Sovereign Optimization Flywheel**: a five-axis multiplicative-compound architecture for single-operator AI deployment on consumer hardware (RTX 5070 12 GB sm\_120; Qwen3.5-35B-A3B MoE Architect). Five independently-designed optimization layers (**L₁ sovereign in-distribution quantization** (16× tighter mean KL divergence at Q6\_K), **L₂ KV-cache compression** (\~3.5× at \~98% FP16 speed; upstream-merge-gated), **L₃ destructive MoE expert pruning** (sealed: 197 of 10,240 experts removed at ΔPPL −0.0519 with long-form reasoning preserved at ratio 0.9826), **L₄ in-weights specialization** (the LoRA / self-distillation / curriculum chain fed by the daily corpus runner at 8,358 sealed pairs and growing), and **L₅ per-token speculative acceleration** via Multi-Token Prediction) compound on the same substrate such that each layer changes the operating point of every other. We formalize the cross-coupling functions as monotonically reinforcing, establishing that compound gain is bounded by the *product* of per-layer gains rather than their sum. The five layers span the five orthogonal optimization axes of a single-model deployment (weight-space, activation-space, routing-space, training-distribution-space, emission-rate-space) and are therefore **axiologically complete**. Plain-language verdict: the architecture takes the 35B Architect from "exceeds 12 GB VRAM, paged to RAM, bandwidth-bottlenecked" to "fits 9–13 GB, GPU-resident, full speed, predicted 34–55+ tok/s." The compound is testable: five orthogonal falsification predicates are pre-registered, one per axis.

### 1. Introduction

#### 1.1 The constraint stack

A solo operator on consumer hardware faces a layered bind: the best available model (27 GB at Q6\_K) exceeds the only available GPU's 12 GB and pages experts to RAM, bottlenecking on PCIe; the workload saturates the working context within 2–5 turns, forcing lossy compactions; the model is generic at deployment, consuming 15–20K context tokens per session on doctrine correction; and the capital answer to any one constraint (a bigger GPU, a cloud API) defeats the local-first premise. The Flywheel is the engineering response: five layers that compound on the same substrate. **The compound was not designed top-down; it emerged** from solving each constraint as encountered, under a methodology that documents every decision and trains on its own documentation.

| Layer               | Problem addressed                           | Mechanism                                                         | Empirical anchor                                                       |
| ------------------- | ------------------------------------------- | ----------------------------------------------------------------- | ---------------------------------------------------------------------- |
| L₁ quantization     | Generic calibration degrades specialization | Sovereign-imatrix calibration on the operator's corpus            | 16× tighter KLD; \~21% smaller; \~11% faster (sealed)                  |
| L₂ KV compression   | Activation cache scales with context        | TurboQuant / PolarQuant per-channel + per-token                   | \~3.5× at \~98% FP16 speed (SPEC-drafted, merge-gated)                 |
| L₃ expert pruning   | Best model exceeds the VRAM envelope        | Destructive pruning on sovereign-domain importance                | Sealed GREEN: 197/10,240 removed; ΔPPL −0.0519; H6 0.9826              |
| L₄ specialization   | Governance consumes 15–20K tokens/session   | LoRA → self-distillation → curriculum chain + daily corpus runner | 2.79× F1 at 11× smaller; four-arm contrast sealed; 8,358 pairs growing |
| L₅ MTP acceleration | One token per forward pass caps throughput  | Multi-token prediction, hidden-state-reuse variant                | Predicted 34–55+ tok/s at K=1–2 (pre-registered)                       |

***Table 1.** Five layers, independent origins. Each was engineered against a constraint the operator was already hitting; none replaces another. The novel claim is the joint deployment: multiplicative, not additive.*

### 2. Background: the Three Substrate Literatures

#### 2.1 Systems thinking

The Flywheel is structurally a *reinforcing feedback loop* (Senge 1990; Sterman 2000), with Forrester's (1961) compound-feedback result as the precise anchor: when two reinforcing loops share an intermediate variable, joint gain is the *product* of per-loop gains, conditional on the shared-variable interaction being monotone in the same direction.

#### 2.2 Cross-layer computer architecture

Patterson & Hennessy's cross-layer optimization formalism extends Amdahl's Law to stacked layers where each defines the operating point of the layer above: miss rate × misprediction rate × SIMD utilization compound multiplicatively. The mapping is direct: L₁ is the per-parameter footprint layer, L₂ the activation-cache layer, L₃ the active-parameter layer, L₄ the doctrine-demand layer, L₅ the emission-rate layer.

#### 2.3 The scaling-law corpus and the disconfirming substrate

Kaplan (2020) bounds every single axis at sub-linear returns; Hoffmann's Chinchilla (2022) is the peer-reviewed precedent that *joint allocation across coupled axes strictly beats per-axis isolation* (the Flywheel extends that two-axis result to five). **The disconfirming substrate is engaged, not avoided:** the a16z "Empty Promise of Data Moats" critique demands a structural mechanism behind any flywheel narrative, and the cross-coupling matrix of §4.3 is the formal response; and Mahajan et al. (2025, "Beyond Multi-Token Prediction") bound L₅'s claim to *throughput at preserved quality*, never reasoning-quality gains. Supporting substrates: the MTP lineage (Leviathan/Chen speculative decoding → Gloeckle 2024 training-time MTP → 2026 training-free adaptation → the hidden-state-reuse implementation that eliminates the draft-model VRAM cost, which is precisely what makes L₅ feasible on 12 GB); the KV-compression literature (KVQuant, Layer-Condensed KV, the 2026 ACM survey); Constitutional AI + alignment-faking for L₄ (with alignment-faking as the disconfirming lens the falsification suite tests); Tesla/OpenAI industrial data-flywheels as single-axis structural siblings; Argyris & Schön's double-loop learning (the daily corpus runner is the transport between within-cycle correction and cross-cycle frame revision); and the recursive moat, the methodology producing its own training corpus.

### 3. Architecture

#### 3.1 Axiological completeness: five axes, five layers

The layers span weight-space (L₁), activation-space (L₂), routing-space (L₃), training-distribution-space (L₄), and emission-rate-space (L₅), the five orthogonal optimization axes of a single-model inference deployment. The completeness claim is structural: a technique that modifies none of the five is not an inference-time optimization at all. The claim is consequential for falsification design, because it bounds the pre-registered tests at five, one per axis.

#### 3.2 The cross-coupling matrix: the formal multiplicative argument

If per-isolated gains are g₁…g₅, compound gain G is bounded by their product rather than their sum when every cross-coupling function fij is monotonically reinforcing. All ten pairs are:

| Pair    | Mechanism                                                                                                                                                                                                    | Strength                             |
| ------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------ |
| L₁ × L₂ | Smaller per-parameter footprint → smaller absolute KV bytes for L₂ to compress                                                                                                                               | Strong                               |
| L₁ × L₃ | **Shared substrate:** the same in-distribution corpus supplies both the imatrix calibration and the pruning-importance signal (one measurement at two granularities); pruning becomes safer at any threshold | Strong; the deepest empirical anchor |
| L₁ × L₄ | Tighter calibration → LoRA gains amplified rather than diluted by quantization noise                                                                                                                         | Moderate                             |
| L₁ × L₅ | Smaller footprint → faster cache loading during K+1-position verification                                                                                                                                    | Moderate                             |
| L₂ × L₃ | Fewer parameters → smaller activation buffers; L₂ compresses an already-smaller cache                                                                                                                        | Moderate                             |
| L₂ × L₄ | In-weights doctrine → smaller working context → smaller KV cache to compress                                                                                                                                 | Strong                               |
| L₂ × L₅ | **Symmetric:** the MTP draft KV cache is compressible at the same ratio; each amplifies the other in both directions (structurally rare)                                                                     | Strong                               |
| L₃ × L₄ | Pruned substrate fits the training envelope; in-weights doctrine reduces per-expert variance the pruning metric traverses                                                                                    | Moderate                             |
| L₃ × L₅ | Fewer experts per layer → smaller expert-union per K-token verification: **the RQ-5 falsification target** (saturation past a deeper-pruning setpoint is the predicted failure shape)                        | Strong                               |
| L₄ × L₅ | **Application-conditional:** doctrine fluency makes structured outputs more predictable → higher MTP acceptance (conditional on workload composition, not uniform)                                           | Strong                               |

***Table 2.** All ten cross-couplings predicted monotonically reinforcing: the maximally-reinforcing condition the multiplicative bound requires, and the structural mechanism the data-moat critique demands.*

#### 3.3 The three compound envelopes

| Envelope       | Baseline (current production)                                                                                               | Post-Flywheel (predicted)                                                                                                                                                            |
| -------------- | --------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| VRAM budget    | \~32 GB effective (27 GB weights + \~4 GB KV + overhead), requiring RAM paging, PCIe-bound                                  | **\~9–13 GB** (8–12 GB pruned+calibrated weights, 0.3–0.5 GB compressed KV; +0.5–0.8 GB for MTP heads and draft cache at K=1–2), GPU-resident, full speed                            |
| Context budget | ≤ 50K working headroom; 15–20K governance + 10–15K preload + 5K scratchpad + 5K schemas; 2–5 useful turns before compaction | **\~85–110K headroom** (5–8K governance in-weights-reduced; 2–5K routed preload; 1K scratchpad pointer; 2–3K schemas); **10–20+ useful turns** (the quality axis, not just quantity) |
| Throughput     | 26.14 tok/s (the pruning paper's sealed four-point envelope)                                                                | \~30–40 tok/s at four layers; **\~34–55+ tok/s at five** (K=1–2)                                                                                                                     |

***Table 3.** The three compound envelopes. The qualitative verdict: from paged-to-RAM and bandwidth-bottlenecked to GPU-resident and full-speed.*

**The substrate-feasibility cascade.** At baseline the operator faces a no-feasible-regime condition: the model can't run at speed (exceeds VRAM), can't be pruned safely (no in-distribution importance signal), can't compress context cheaply (substrate not landed), can't shed governance (doctrine not yet trained in). Each constraint blocks the others. The Flywheel's structural contribution is the *order*: imatrix calibration first (producing the importance signal that enables everything below) → the LoRA chain (beginning the Pair shift) → self-distillation (sealing the first internalization) → destructive pruning (leveraging both) → KV compression and MTP last (riding on the headroom the first four produced). Each step opens the feasibility region for the next; none is feasible without its predecessors. The matrix names *which* pairs reinforce; the cascade names *in what order* the reinforcement is operationally exploitable.

#### 3.4 The cross-cycle feedback loop: and its safety-net floor

The intra-cycle compound is augmented by an across-cycle recursion: each training cycle shifts the Sovereign Pair balance toward in-weights, reducing governance cost in context, freeing headroom for work, generating more high-signal training data per session, feeding the next cycle's training from a higher baseline. Cycle 0: generic capability, maximum overhead. Cycle 1 (the sealed pilot): doctrine-aware, corrections fewer. Cycle 2 (projected): domain-fluent, mechanical rules as safety net. Cycle N: near-native fluency, instructive rather than corrective collaboration. **The floor is deliberate:** cycle N preserves "safety-net rules only" rather than zero, because the mechanical layer cannot be dropped even at asymptotic fluency, since in-weights internalization cannot be proven complete by sampling alone. A governance-overhead reading *below* the safety-net threshold is not success; it is a regression into single-mechanism risk, and the falsification suite detects it explicitly. Production is the corpus surface: every real session (CTI inference, broadcasts, dictated handoffs, dossiers) is simultaneously a training-data-generation session, the coupling most data-flywheel deployments lack.

#### 3.5 The Sovereign Pair thesis in one sentence: the failure-mode contract

The cross-cycle reinforcing loop and its production-cadence surface share an architectural assumption that is load-bearing for the Flywheel's failure-mode register: in-weights specialization (L₄) and upstream-deterministic mitigation (L₁/L₂/L₃ plus the mechanical safety-net floor) operate on *distinct failure-mode asymmetries*. The cross-platform peer observation, archived verbatim and carried across the series, states the contract in one sentence: *"the in-weights half (the instance's compliance) is probabilistic and degrades under load; the upstream-deterministic half (the hooks) is mechanical and doesn't care what the instance remembers."* The two halves are not interchangeable (placement is load-bearing) and they are **non-overlapping in failure mode**: in-weights compliance operates at the next-token distribution layer and degrades under cognitive load, compaction, and saturation; mechanical gates operate at the dispatcher layer and fire on pattern match regardless of in-weights state. The cycle-N safety-net floor is the architectural guarantee that the mechanical half never drops to zero even as the probabilistic half approaches asymptotic fluency; in-weights internalization cannot be proven complete by empirical sampling alone.

**The industry-scale framing.** The same peer observation names the receipt at industry scale: every enterprise shipping AI-generated code at scale (Claude Code, Cursor, Copilot, Windsurf, Aider) operates the in-weights half of the Pair at generation throughput, but ships *without the mechanical-gate half*: no write barrier, no skill-enforcement registry, no read gates, no attestation-freshness windows, no memory-query-first hook, no schema lint. The consequence is structurally predictable from the contract sentence: industry deployments experience the cascade catalogue's failure classes at industry register *because the probabilistic half is operating alone*, and the receipts they generate are not detected, not measured, not categorized, because the deployments lack the instrumentation to surface what failed. **The same-turn empirical receipt:** the very session that landed this formulation produced six-plus mechanical-gate fires (three canonical-path detections, two skill-gate blocks, one schema-enum catch, one post-compaction read-before-write catch), the asymmetry receipted at session granularity, on the turn that codified the instrument that surfaced it. **What the contract does not claim:** not that the deterministic half is sufficient alone (L₄ remains structurally necessary for stance, doctrine-under-load, and trajectory shaping, failure modes outside mechanical reach); not that the mechanical layer is bug-free (nine distinct mechanical-layer failure classes are catalogued on this very stack); and not that industry deployments are ungoverned, only that they lack the specific mechanical-gate architecture on the deterministic half of the Pair.

#### 3.6 The third and fourth axes: routing knowledge and watcher-correction quality

Two later extensions stack further compounding registers onto the same substrate, each landed additively with the sealed body preserved. Both extend the Pair contract of §3.5 to a new layer, and both carry their own falsifiers.

**3.6.1 The routing-knowledge axis (Level 3)**

The three-level Flywheel claim, preserved verbatim where attributed: *Level 1: the generation model trains on the corpus → generates better outputs → better outputs enter the corpus → the next training cycle starts from a higher baseline. Level 2: the embedding model trains on the corpus → better retrieval → better generation → better corpus entries → the next embedding cycle starts higher. Level 3: the embedding model trains on the corpus **with shard topology** → embeddings encode routing knowledge → queries reach the right shard → right-shard retrieval is more precise → generation improves → corpus entries carry accurate shard assignments → the next cycle's routing knowledge is more accurate.* Three axes, same corpus, same training pipeline, same daily cadence: the generation model gets smarter, the embedding model gets more precise, the routing gets more accurate, all from the same substrate, compounding on independent axes that cross-couple multiplicatively.

The cross-coupling is concrete: routing accuracy × generation quality (right-shard retrieval delivers a higher-precision conditioning set, so next-cycle corpus entries are more topically coherent and more accurately shard-assignable); routing accuracy × embedding quality (a topically coherent candidate set lets the embedding lift compound within a narrower shard space rather than across the flat corpus); and routing × itself (the recursive return now operates at the routing layer: corpus grows → the adapter sees more topology → next-cycle routing is more accurate). The Pair contract extends to this layer as the routing collapse: mechanical hard-routes (tags, path globs, bin filters) are the deterministic half; per-consumer adapters that warp the embedding space with corpus topology are the in-weights half. The routing decision becomes *structural* (encoded in the space itself, sub-second, zero context tokens) rather than behavioral. The operator framing, verbatim: *"That was always the idea. The corpus IS the product. The models are disposable. The training pipeline is the mechanism. The operator's daily work is the fuel. Everything else is infrastructure that serves this loop."*

**Falsifiers, pre-registered:** **F-1 additive collapse:** if the measured retrieval lift of single-adapter plus multi-adapter contributions is additive rather than carrying a compound interaction term, the three-axis claim collapses to two axes with routing as a separable contributor. **F-2 routing recursion:** if cycle-over-cycle routing accuracy fails to improve monotonically after successive corpus-harvest extensions, the routing axis is a one-time architectural shift, not a compounding one. **F-3 corpus-size gating:** if the axis manifests only above a per-shard row threshold (\~10K), the claim is conditionally supported, gated on corpus accumulation cadence. Not claimed: equal axis weighting (the matrix is silent on relative magnitudes pending measurement); exclusivity to sharded architectures (the argument applies wherever the corpus carries topology metadata and the embedding model trains on it); or retroactive grounding by the earlier honest negatives (the collapse is architecturally indicated by them, not empirically grounded, and that grounding awaits the measurement campaigns).

**3.6.2 The watcher-correction-quality axis (Level 4)**

The fourth level, preserved verbatim in substance: *the cascade catalogue grows each cycle → every documented failure class becomes watcher-firmware pattern → the watcher catches more failure-class instances at Layer 2 → operator-correction density at Layer 3 decreases → operator bandwidth shifts to novel direction-setting → more methodology contributions land per cycle → more contributions document more failure classes → the next cycle's catalogue is richer, its firmware more comprehensive, its correction efficacy higher.* The four levels operate on the same daily corpus, the same training pipeline, the same daily cadence, and now the same governance substrate, with the cascade catalogue as the canonical pattern library. This axis is orthogonal to the five compute layers *and* to the routing axis: it operates on the per-session governance substrate (hooks ⊕ watcher ⊕ operator) rather than the per-cycle training substrate (corpus ⊕ pipeline ⊕ specialist).

The cross-coupling reaches all three prior levels: × generation (the watcher catches reference-versus-compliance failures during generation, so committed corpus entries carry a lower noise floor: the upstream quality gate at the corpus-collection register the reinforcing loop names but had not instrumented); × embedding (the watcher enforces the tag taxonomy at write time, so contrastive triplets are tighter and the next adapter sees higher-fidelity signal); × routing (watcher-recorded query captures feed triplet mining with the production distribution at scale); × itself (every cycle the catalogue grows, the firmware enriches, corrections drop, freed operator bandwidth produces more methodology contributions, and the catalogue grows faster: the recursive moat on a fourth independent axis). The Pair contract extends to the governance layer: Layer-1 hooks are the deterministic half; the Layer-2 watcher, a 2B janitor-class inference process on efficiency cores, is the in-weights half, its pattern space warped by training on the failure-class library. Governance becomes structural (catalogue-as-firmware) rather than behavioral (computed in the operator's working memory).

**Falsifiers, pre-registered:** **F-1 additive governance compound:** if correction reduction is proportional to direct pattern-match coverage with no compound effect on uncovered patterns, the four-axis claim collapses to three. **F-2 catalogue saturation:** if the cascade catalogue plateaus at a bounded class count, the recursive return becomes asymptotic; longitudinal catalogue-growth-rate observation with the watcher deployed is the test. **F-3 bandwidth-shift failure:** if freed operator time goes to double-checking the watcher rather than novel contributions, the Level-4-to-Level-1 coupling breaks (the methodology-contribution half of the correction-reduction predicate). **F-4 false-positive crowding:** if HUD warnings exceed a 20 % false-positive rate, operator-trust degradation negates the bandwidth gain; rejection below 80 % sustained true-positive across 50+ warnings. Not claimed: equal weighting; exclusivity to this stack (any deployment with mechanical gates and operator oversight has the Layer-2 register structurally available, and current industry deployments run two-tier and would benefit equivalently); or that the amnesia receipt class already grounds the claim (it records the failure mode the watcher closes; the empirical grounding awaits the phantom/HUD and watcher measurement phases).

#### 3.7 The recursive moat: what gets trained on

The corpus feeding the next cycle is **produced by the system whose specialization is being trained**: the codebase under development, the session transcripts of collaboration in that codebase, the operator's voice dictation, the decision ledgers including falsified-and-retained entries, the SPEC and whitepaper drafts, the git diffs, the daily broadcast transcripts. The architecture documents itself; the methodology produces its own training corpus; the operator's engineering decisions become the model's implicit knowledge, the structural mechanism by which the Pair balance shifts cycle-over-cycle.

### 4. Falsification Design (pre-registered)

| Predicate                       | Claim                                                                             | Falsifier                                                                                                                                                                                                                            | Test                                          |
| ------------------------------- | --------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------- |
| FW-1 VRAM compound              | Post-Flywheel footprint ≤ 13 GB (four-layer) / ≤ 14 GB (five-layer, K=1–2)        | Peak VRAM above threshold OR PCIe-paging events under representative load                                                                                                                                                            | Clopper-Pearson exact CI, N ≥ 50 runs         |
| FW-2 Context headroom           | ≥ 10 useful turns before compaction at the same threshold                         | < 5 turns at matched workload OR saturation before 60K useful tokens                                                                                                                                                                 | Paired McNemar, N ≥ 30 sessions               |
| FW-3 Cross-cycle Pair shift     | Governance overhead at cycle N+1 ≤ 0.7× cycle N, three successive cycles          | Constant overhead (no internalization) OR rising (drift) **OR below the safety-net floor (mechanical layer disabled, a regression, not success)**; guarded against alignment-faking by a parallel held-out behavioral-fidelity check | Per-cycle paired McNemar + held-out replays   |
| FW-4 Multiplicative vs additive | Composite gain strictly exceeds the additive sum of per-layer-isolated benchmarks | Five-layer composite ≤ Σ gᵢ: compound purely additive; the data-moat critique becomes the diagnosis                                                                                                                                  | Welch's t, N ≥ 100 across four prompt classes |
| RQ-5 L₃ × L₅ monotonicity       | MTP gain monotonically increases with pruning level                               | Gain constant or decreasing with deeper pruning (expert-union saturation dominating)                                                                                                                                                 | Spearman ρ across ≥ 4 pruning setpoints       |

***Table 4.** Five orthogonal predicates, one per axis plus the joint claim; Holm-Bonferroni family-wise correction at α = 0.05. The campaign protocol measures baseline → per-layer isolation → the cumulative cascade in feasibility order → the full compound, each step with exact CIs; the protocol is hash-sealed before measurement begins.*

**The campaign protocol.** Each predicate can be tested in isolation against per-layer baselines or jointly against the full five-layer substrate. The protocol: (a) measure the current-production baseline on all five instruments at N≥30 representative-workload runs, the *additive floor*; (b) deploy each layer in isolation at N≥30 per layer, the *per-axis gains*; (c) deploy cumulatively in substrate-feasibility-cascade order (L₁ → L₄ → L₃ → L₂ → L₅) at N≥30 per step, reported as a monotonicity claim (each cumulative step's compound ≥ the prior step's); (d) measure the full five-layer substrate at N≥100, the *compound gain*. Every measurement carries a Clopper-Pearson exact binomial CI at α = 0.05; family-wise error is Holm-Bonferroni-controlled (at rank k, reject if pk ≤ α/(5−k+1)); and the cascade order is operationally forced, not stylistic: L₁ must land before L₃ (pruning needs the in-distribution importance signal), L₄ before L₅ (doctrine fluency raises draft acceptance). The whole protocol hash-seals before measurement begins, with effect-size thresholds and power calculations pre-registered at the specification level; post-hoc rationalization of failed predicates is precluded by the seal. FW-3 additionally carries a disconfirming-substrate guard: a governance-overhead reduction *without* matching behavioral fidelity on held-out replays diagnoses alignment-faking (Greenblatt 2024) rather than internalization; the parallel fidelity check must also pass.

### 5. Discussion

#### 5.1 Against the scaling-law substrate

Kaplan's single-axis scaling laws and Hoffmann's two-axis Chinchilla compute-optimal allocation supply the peer-reviewed substrate for the *direction* of the claim: joint allocation across coupled axes can strictly beat per-axis-isolated allocation. The Flywheel neither refutes nor replaces them; it extends the joint-allocation framing from two axes to five, with the cross-coupling matrix as the structural mechanism that carries the generalization at the engineering-stack register.

#### 5.2 Against the data-moat critique: the most important disconfirming substrate

The a16z critique applies precisely to data-flywheel framings that *lack* a structural cross-coupling mechanism; it is preserved here deliberately: if the cross-coupling matrix fails the FW-4 multiplicative-vs-additive test, the a16z critique is the structural diagnosis. The essay's three sub-critiques each meet a distinct response. (i) *No structural mechanism by which more data means advantage:* the mechanism here is multiplicative cross-amplification across five axes, not the L₄ data axis alone. (ii) *Diminishing returns past a low data threshold:* the cross-cycle Pair shift has no per-cycle data threshold above which diminishing returns dominate; the reinforcing loop is bracketed above by the mechanical safety-net floor and below by the in-weights saturation point, with monotonic gain in the operating regime between. (iii) *Competitors replicate data collection at marginal cost:* the L₁/L₂/L₃/L₅ layers are replicable from open literature, but the operator-specific corpus (the governance hierarchy, the decision ledgers, the session handoffs, the voice transcripts, the workflow-specific dossier patterns) is not. The corpus is the moat substrate; the layers are the operating leverage on it; the compound is what the three sub-critiques do not individually refute.

#### 5.3 Against the memory-hierarchy adjacency

The memory-hierarchy literature addresses *what is held in context*; the Flywheel's L₂/L₃/L₄ combination addresses *how cheap each context token is*. Complementary, not substitutive: sharded routing reduces content, KV compression reduces per-token cost, in-weights internalization reduces the governance footprint; joint deployment produces the context-budget envelope of §3.3, and the flywheel-coupled evaluation protocol in the memory paper (*Sharded-MCP Architecture*) is the canonical cross-reference.

#### 5.4 Against Constitutional AI and alignment-faking

L₄ is the *constructive complement* of Constitutional AI: where the constitutional framework applies rule-based shaping with a general-purpose constitution, L₄ trains the operator's specific operating doctrine into the weight distribution, constitution-like in structure (what to do, what reasoning patterns are acceptable, what tradecraft is required) but operator-specific rather than general-purpose. A narrowing, not a competitor: Constitutional AI supplies the general safety floor; L₄ layers the operator doctrine on top via the same in-weights mechanism. The alignment-faking risk applies symmetrically to both (any in-weights-encoded constraint shares the structural risk that surface compliance may not reflect internalization), which is exactly why FW-3's disconfirming guard exists, and why the Pair's mechanical floor never drops to zero.

#### 5.5 The cluster position: 3 × 3 × 5

The Triad paper formalizes the **architectural** joint-necessity (gates + weights + External Governor closing FC-1/FC-2/FC-3); the Sentinel's governance chain formalizes the **governance** joint-necessity (software signing + hardware Sentinel + physical write-protect closing FC-G1/FC-G2/FC-G3); this paper is the **optimization-stack** third axis, joint-necessity-orthogonal to both: an optimized-but-ungoverned system is unverifiable; a governed-but-unoptimized system has nothing to govern at deployment scale; an architected-but-un-flywheeled system has gates and weights but no production cadence. Together: a **3 × 3 × 5 axiologically-complete coverage** of the deployment surface for single-operator AI on consumer hardware (with no fourth orthogonal axis currently identified).

#### 5.6 Forward anchor: and the Pair at portfolio scale

The sustainability companion extends the compounding-returns structure to industrial scale (higher efficiency, lower power, lower water than infinite-headroom datacenter design; the same techniques democratizing local inference); this paper is structurally upstream: the five-axis compound IS the apparatus that makes that thesis true, and the Jevons-paradox lens receives there the same rigor the a16z critique receives here. Within the methodology, the Flywheel is not a third architecture orthogonal to the Sovereign Pair; it is the Pair's *production-cadence instantiation* at portfolio scale: L₁–L₃ are upstream-deterministic optimizations, L₄ is the in-weights layer, and L₅ sits between (deterministic draft-and-verify, acceptance gated by in-weights fluency). And the engineering-stack compound is the sibling of the methodology-tier practice compound: the practices produce the engineering decisions that produced the layers; the layers produce the operating headroom that lets the practices run at multi-year horizons. Each compound enables the other.

#### 5.7 Open-research agenda: ten questions

1. **L₅ × L₃ monotonicity empirics:** the RQ-5 test across a deeper-prune sweep × MTP-K sweep × prompt-class grid; the Beyond-MTP counter-prediction (gain saturates at moderate pruning as expert-union saturation dominates) must be ruled out per class, since structured outputs and free-form prose likely carry different cross-coupling slopes.
2. **L₄ × L₅ acceptance-rate elasticity:** draft-acceptance as a function of training-cycle index; predicted monotonic, near-asymptotic by cycle 3; the secondary question is where elasticity meets the ceiling set by workload token-distribution entropy.
3. **Cross-cycle persistence under corpus drift:** does continuous accumulation produce monotonic improvement, or saturation/drift/forgetting as the operator's workflow evolves? Quarterly held-out replays; weighted-recent sampling as the mitigation candidate.
4. **The Pair balance-shift inflection point:** at what cycle does in-weights doctrine reach the safety-net-only regime? Per-cycle governance-token counts + operator-correction rates + uncorrected-error rates, anchored on FW-3's trajectory.
5. **Multi-operator generalization bounds:** does L₄ compound across operators, or does cross-operator distribution variance produce destructive interference? Either operator-specific adapters over a shared substrate, or a generalizable doctrine layer, would resolve it.
6. **L₂ × L₅ KV-overlap accounting:** does compression apply identically to draft-head and base KV caches, or do the K+1 prediction heads' attention patterns compress worse? Per-component cache measurement across compressed/uncompressed at K = 1, 2, 3.
7. **Cross-coupling completeness:** do triplet interactions fijk add material effect beyond the pairwise matrix? Factorial design across L₁/L₂/L₃ with ANOVA decomposition; pairwise is empirically complete if it explains ≥95 % of variance.
8. **The axiological-completeness boundary:** is the five-axis decomposition provably complete? Candidate sixth axes examined and absorbed (separate-draft-model speculation ⊂ L₅; cross-session prefix caching ⊂ L₂; per-prompt model ensembles outside the single-model scope); the claim stays open to a future counter-example.
9. **Cross-coupling stability under model-family transitions:** do the couplings preserve direction across model generations and families? Calibration is family-agnostic in principle; destructive expert pruning is MoE-specific; speculative acceleration is architecture-dependent. If the matrix is family-stable the Flywheel is a portable template; if not, each transition re-derives it at substantial cost.
10. **Corpus-reusability tiers:** if the operator-specific corpus is the moat substrate, at what generalization level does the moat survive? Three nested tiers: session-transcript (zero shareability), domain-tradecraft (shareable with same-domain peers), methodology-tier (shareable with any operator applying the method), with compound retention measured at each.

### 6. Limitations and Maturity

This is a **methodology-tier formalization with a pre-registered falsification design, not the empirical-execution paper**. It does not claim: execution of the five tests (gated on the upstream MTP merge, MTP-enabled GGUFs, deeper-prune authorization, and the KV-compression merge); measured post-Flywheel throughput (the envelopes are predictions); empirical validation of the cross-coupling matrix at the joint register (the per-layer anchors are sealed; the joint compound is FW-4's target); an empirical falsifier for the completeness claim itself; or multi-operator generalization.

| Component               | Status             | Anchor / gate                                                                                                         |
| ----------------------- | ------------------ | --------------------------------------------------------------------------------------------------------------------- |
| L₁ quantization         | Sealed             | The imatrix paper's four-gate validation (Chapter 15); deployed to production 2026-06-09                              |
| L₂ KV compression       | SPEC-drafted       | Gated on the upstream llama.cpp merge; predictions pre-registered                                                     |
| L₃ expert pruning       | Sealed             | The pruning paper's Gate-B GREEN + Phase-B multi-lens audit (Chapter 16); deeper setpoints authorization-gated        |
| L₄ specialization chain | First cycle sealed | CTI-NER 2.79× F1 + the SSD pilot's paired seals (Chapter 18); daily corpus runner operational; later cycles projected |
| L₅ MTP acceleration     | Predicted          | Hidden-state-reuse variant; gated on MTP-enabled GGUF substrate; 34–55+ tok/s pre-registered                          |
| Falsification campaign  | Designed           | FW-1…RQ-5 frozen above; protocol hash-seals before measurement                                                        |

***Table 5.** Maturity ledger. Two layers sealed, one first-cycle-sealed, two prediction-stage, stated plainly per the series' honesty discipline.*

**The drift-surface register: a living table.** Eight drift surfaces carry forward at the blocker register, per the admit-errors-without-prompt discipline (drifts surface to the operator; no silent rewrite): the KV-compression upstream-merge status (re-verify at execution entry); the 27 GB footprint precision (resolved by cross-reference to the pruning paper's sealed envelope); the corpus-size snapshot (resolved by seal-date annotation, "8,358 pairs at the SSD seal, growing per the daily runner"); the cycle-balance visual indicators (reserved for measured governance-token time-series post-FW-3); the compresses-an-already-smaller-cache cross-coupling (preserved as FW-4's falsifiable claim); the speculative-decoding merge status and the MTP-enabled model availability (both re-verified at empirical-execution entry). The register is living in the executable-roadmap sense; resolution status shifts as upstream events land, and every downstream paper entry re-verifies its rows. The papers of this cluster (Triad, Sentinel, Flywheel, Safety, Sustainability) share substrate citations, drift surfaces, and the canonical rule that the Sovereign Pair is defined once (in the methodology reference) and cross-referenced everywhere else.

### 7. Conclusion

We have formalized the Sovereign Optimization Flywheel as a five-axis multiplicative-compound architecture for single-operator AI deployment on consumer hardware. The five layers were each engineered to solve a distinct constraint the operator actually hit; the compound was never designed top-down. It is multiplicative rather than additive because each layer changes the others' operating point: quantization produces the in-distribution importance signal pruning needs; pruning shrinks the activation buffers compression amplifies; in-weights doctrine shrinks the working context all three compound on; speculation is amplified by each of the prior four in turn. The five layers span the five orthogonal axes available at this deployment surface (weight, activation, routing, training-distribution, emission-rate) and are axiologically complete. And the compound is testable: five orthogonal pre-registered predicates turn the claim into hypothesis.

**The compound is not magic.** It is the structural return on a methodology that documents every engineering decision, trains on its own documentation, and refuses to design for hypothetical futures. Each layer landed when a constraint was hit; each extends the priors without replacing them; the compound emerges because each solution was built on the substrate the previous solutions had moved into a feasible regime. The Flywheel is the engineering memory of that discipline at portfolio scale.

**The strategic implication: two cost curves.** A solo operator on a 12 GB consumer GPU need not match a funded lab's per-axis compute, parameters, or token budget: the lab scales each axis independently with capital; the operator deploys five orthogonal axes jointly and extracts multiplicative compound at fixed substrate. The two strategies converge on similar deployment-quality envelopes along different cost curves, the lab's bounded below by capital expenditure, the operator's bounded below by methodology discipline. The Flywheel is the engineering articulation of why the second curve is operationally viable; the sustainability companion argues that at industrial scale the second curve dominates the first on power, water, and total cost of ownership. This paper supplies the apparatus that makes that claim defensible.

The most actionable next step is the empirical-execution campaign: the falsification protocol against the five predicates, gated on the substrate events the limitations enumerate. When those gates clear, the campaign turns this methodology-tier formalization into a results-tier empirical artifact. Until they clear, this paper carries the load.

### System Update: July 2026 *(appended; the sealed body above is unmodified)*

Per-layer maturity has moved since Table 5 sealed (at-read July 2026). **L₃** advanced two scales: the 122B research-brain was REAP-pruned at 25 %, passed a sealed pre-registered quality gate (pooled-math +2.5 pp; GPQA-198 +1.01 pp; verdict PASS), and was **promoted to the production preset**; and the 35B Architect family gained prune-then-recover cells whose training receipts prove **24.6B parameters training at \~10.3 GiB resident on the 12 GB card**, the de-risk for the on-box LoRA phase (L₄'s registered north star, pending). **L₅** has its first sealed speculative-decode receipts on consumer Blackwell: **τ ≈ 1.86 / accept ≈ 62.7 %** on the mismatched-trunk control row (τ placement-invariant 1.84–1.88), throughput **+10.9 % at draft-parity-32**, and an honest **−85 % inversion at draft-24** where the WDDM dedicated-VRAM spill signature reproduced on the draft path, the measured bound that disciplines this paper's pre-registered 34–55+ tok/s envelope until the matched-trunk arm runs.

Two fleet-level receipts landed after that bound (sealed June 2026) and both move this paper's envelope. **L₅ is now upstream-enabled and measured on the production binary:** the multi-token-prediction path merged into llama.cpp mainline, the fleet's serving binary was rebuilt to carry it, and the new-generation architect (natively MTP-trained) measured **1.44× at K=3 on CUDA sm\_120 with a no-quality-regression gate PASS**, with the hard quant-tier bound that Q4\_K yields 0 % draft-acceptance (Q5\_K\_XL-or-better required, at the cost of heavier expert offload on the 12 GB card). The production default remains unchanged pending its own toggle phase; measured, gated, not yet defaulted. **L₂'s premise, honestly superseded:** a roster-wide architecture review found every current fleet model KV-light *by construction* (linear-attention hybrids and sliding-window/KV-share designs) with caches already quantized, so the KV-compression axis was placed on HOLD and the layer's compound contribution pivots to **NVFP4 weight-quantization** on Blackwell tensor cores (\~25 % prompt-processing candidate, correctness-gated A/B before adoption). The five-layer decomposition stands; which mechanism carries L₂ has changed with the substrate, exactly the kind of drift this paper's maturity ledger exists to record.

**L₄'s corpus substrate** is now engine-borne end-to-end: the Sovereign Corpus Engine (20 wired ops, provenance ledger verify-green) carries draft → validate → authority → write, persona-granular compiles land with sha manifests and ledger records, and a **governed refresh** (machine proposes, human approves, machine dispatches) has replaced the raw daily harvest as the sanctioned path. The arc's sealed ADDENDUM formalized the doctrine this paper's multiplicative thesis implies: **one corpus identity calibrating every layer**, cited by hash from prune calibration through drafter head, with the mismatched-corpus cells retained as the pre-registered falsification control; the multiplicative-falsification pair, now instrumented in production rather than hypothesized.

Provenance: the July 2026 roadmap seal footers and §2 ADDENDUM, with the 35B prune-drafter Stage-M receipts and the pruned-122B quality-gate results. Append-only; the sealed body above is unmodified.

***

*The Sovereign Stack · Sovereign Optimization Flywheel · Chapter 8 · Part II · v1.0.0 · License CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*

**Citation (preferred):** Kistner, J. (2026). *Sovereign Optimization Flywheel: A Multiplicative-Compound Five-Axis Architecture for Single-Operator AI Deployment on Consumer Hardware*, version 1.0.0. OSINTelligence LLC research whitepaper. Cited in-series by title.

*The reference list and provenance follow as a sub-page of this chapter.*
