> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/evidence-and-seals/35b-prune-experiment-p3-mismatch-control.md).

# 35B Prune Experiment (P3, mismatch control)

*Evidence & seal for* [***16 · Sovereign Domain Pruning***](/osintelligence/part-iv-the-evidence-what-worked/16-sovereign-domain-pruning.md)*. This is the pre-registered experiment that pruned the deployed 35B Architect (`Qwen3.6-35B-A3B`) and measured every flywheel layer, prune scoring, recovery fine-tune, imatrix quant, drafter τ, against the deployed Q5 baseline. It returned a **composite FAIL**, and that is exactly why it matters: it stands as the frozen **mismatch control** (every layer measured on the 122B-slot corpus, not the sovereign one) against which the matched-corpus chain (Chapter 18 evidence) is compared. A pre-registered experiment that fails, kept in the record at full fidelity, is the discipline working as designed.*

*New to how these seals work, read* [***Verifying a Seal***](/osintelligence/evidence-and-seals/verifying-a-seal.md) *first.*

**Experiment id** `osint-suite-dev-2026-07-06-764b` · **pre-registered** 2026-07-07T03:55Z · **final-sealed** 2026-07-12T05:55Z · **verdict composite FAIL** (H3 forecloses H1 ∧ H2 ∧ H3; the incumbent stays deployed).

## The seal (quoted from the sidecar, verify it yourself)

```
49d58009699bde45e60fb46dc12695dcf22902c6bb0479fa9ff8155b56d4e8e8  PRE_REGISTRATION.md
prefix-bytes-hashed: 22696 · raw-file-bytes: 24682 · marker-position-lf: 22697
computed AT 2026-07-07T04:36:31Z (post-DEV-1 re-seal; pre-amendment hash preserved in AMENDMENT_1 §4; canonical-prefix per seal_prereg.py)
```

Run the algorithm on [**Verifying a Seal**](/osintelligence/evidence-and-seals/verifying-a-seal.md) over the frozen text below and you will reproduce this hex and byte count. At final seal the post-run integrity envelope re-computed the same prefix hash and it matched the experiment marker exactly, the methodology was unchanged through the run (freeze HUD read 15/15 frozen paths fresh). The experiment carried **eight amendments** (DEV-1 through DEV-7 lineage), each logged and hash-resealed before the stage it governed; their pre-amendment hashes are preserved in the amendments themselves.

## Verdict (final seal, adjudicated mechanically on the sealed rules)

**P3 composite FAIL — H3 forecloses the conjunction gate (PASS = H1 ∧ H2 ∧ H3); the incumbent `[architect-apex]` UD-Q5\_K\_XL stays deployed.** The two quality hypotheses passed on the thinking-on surface; the mechanical resident-serving gate failed, and under the pre-registered rule that is decisive. No post-hoc discretion was available or used.

| Gate                                       | Rule (sealed)                                                                | Measured (at-seal)                                                    | Verdict                                   |
| ------------------------------------------ | ---------------------------------------------------------------------------- | --------------------------------------------------------------------- | ----------------------------------------- |
| **H1** math retention                      | Δ pooled-math ≥ −5.0 pp (n=80 paired, thinking-on)                           | 66.25% vs 67.50% → **Δ −1.25 pp** (McNemar p=1.0)                     | **PASS**                                  |
| **H2** knowledge retention                 | Δ GPQA-198 ≥ −6.0 pp (thinking-on)                                           | 45.45% vs 29.80% → **Δ +15.66 pp** (p=3.7e-4)                         | **PASS** (truncation-salvage caveat)      |
| **H3** resident serving (mechanical)       | fully resident (no `--n-cpu-moe`), no WDDM spill, ≤ 12,227 MiB, ≥ 63.8 tok/s | IQ4\_XS 12.68 / 13.69 GiB > 12 GB card; WDDM spill; 8.82 / 8.48 tok/s | **FAIL** (both cells; Q4\_K\_M dominated) |
| **H4** MTP branch (selects branch)         | KEEP iff accept ≥ 55% AND Δ tok/s ≥ +10% on the H3-winning resident cell     | precondition absent (no resident cell) → operator RETRAIN-vs-DROP     | branch                                    |
| **H5** drafter three-arm τ (report-graded) | report τ/throughput; recipe branch                                           | A1-only per DEV-7 (A2/A3 → matched-corpus re-run)                     | reported                                  |
| **Composite**                              | H1 ∧ H2 ∧ H3                                                                 | —                                                                     | **FAIL**                                  |

**Why the composite FAIL is the point.** The close-out re-classed these cells as the pre-registered **mismatch control**: every corpus-touching layer was measured on the D3-mix corpus built for the 122B slot, frozen before the sovereign-corpus arm existed. Read honestly, the mismatch is visible in the numbers the experiment banked, the exploratory no-thinking deployment-parity table (which cannot move the verdict) shows AIME24 **−53.3 pp** (p=3e-5), GPQA **−24.7 pp** (p=5e-9), MATH-500 −8.0 (n.s.), with the best-retained domain (MATH-500) being the one the mismatched corpus was densest in. The thinking-on H2 "+15.66 pp win" was truncation salvage (the apex baseline hit 188/198 content-zero at the token budget; the gap dissolves to zero on both arms without thinking) and is reported as such, not banked as a real gain. That candor is the seal's whole purpose.

**What the control still banked (at-seal, honest scope):** the drafter τ statistic 1.86 at 62.7% accept, placement-invariant across offload settings; the placement inversion reproduced on the 35B (+7.1% at n-cpu-moe 40 → +10.9% at 32 → −85.0% at 24, the dedicated-VRAM-lies signature on the draft path); and an on-box-trainability receipt (24.59B and 26.61B cells recovery-fine-tuned on the 12 GB card, \~10.3 GiB steady, val CE −0.26 both). This gate does not isolate prune-only cost (the baseline carries Q5 damage and the pruned arm carries prune+Q4+recovery), by design and by the pre-registered deployment-retention framing.

**Deviations through the run:** seven, all logged and hash-resealed at resolution with no open rows at seal, DEV-1 (vision-preservation + the MTP-is-MoE corrigendum), DEV-2 (score-pass device fix), DEV-3 (grouped-mm eager pin plus the k′-alignment law), DEV-4 (cell-2 ratio), DEV-5 (prebake-window rebase), DEV-6 (the no-thinking supplementary pass, AMENDMENT\_7), DEV-7 (drafter A1-only re-scope, AMENDMENT\_8). Full narratives in the experiment's `decisions.md`.

## Integrity envelope (11 artifacts, SHA-256, computed 2026-07-12T05:49Z)

Each run artifact carries an authoritative `.sha256` sidecar; whole-file hashes at final seal, first 16 hex:

| Artifact                                                              | Bytes   | sha256\[:16]       |
| --------------------------------------------------------------------- | ------- | ------------------ |
| `p3_results.jsonl` (Stage R thinking-on, 556 items)                   | 460,004 | `ec9b664b2f482744` |
| `p3_retention_summary.json` (frozen H1/H2 verdict source)             | 1,968   | `119a044c97a26f71` |
| `p3_bank_manifest.json` (banks + subset SHAs; byte-identical to D7's) | 3,998   | `22531c4d84c8fbcd` |
| `p3_results_nothink.jsonl` (AMENDMENT\_7 supplementary, 556 items)    | 185,096 | `cfb8eb8f01743357` |
| `p3_retention_summary_nothink.json` (exploratory deployment-parity)   | 1,986   | `44d7e0ddd85b7e2b` |
| `p3_bank_manifest_nothink.json`                                       | 3,998   | `22531c4d84c8fbcd` |
| `telemetry/…p3_score35…/score_receipt.json` (Stage P scoring)         | 1,034   | `a90ae7265dddeafe` |
| `telemetry/…p3_recover35…/recover_receipt.json` (slim31p)             | 2,310   | `681ca2effd091667` |
| `telemetry/…p3_recover35…/recover_receipt.json` (slim25p)             | 2,308   | `0ad4a80d8b158e44` |
| `telemetry/…p3_quant35…/quant_receipt.json` (Stage Q / H3)            | 4,453   | `2eab311748d3db39` |
| `telemetry/…p3_mtp35…/mtp_receipt.json` (Stage M A1 τ)                | 23,610  | `91e3a37a9a2caea5` |

Run tally: Stage G 5 gates · Stage P 40-layer score + 2 slim cells + 2 recovery runs (\~33 h) · Stage Q convert/imatrix/quant + H3 receipts · Stage R 2×556 items (thinking-on 14.4 h + no-think 6.0 h, zero errors) · Stage M 8 τ sessions. The 15 frozen artifacts remain on disk with their `.sha256` sidecars as the tamper-evidence record.

*Supporting earlier pruning research (activation profiling, pruning execution, phase-B companion) lives under the experiment's `P8_sovereign_pruning` tree with its own decision ledger; it is the method's development record, not a separate sealed pre-registration.*

## The frozen pre-registration, verbatim

Reproduced exactly as sealed (the bytes the self-hash above is computed over). The strikethrough in §2 (H4) is the AMENDMENT\_1 corrigendum, preserved as the record carries it.

```md
---
doc_class: attestation
status: FROZEN
status_as_of: 2026-07-06
frozen_at: "2026-07-07T03:55Z (SHA-256 canonical-prefix self-hash via /sidecar Mode C at seal moment; hex lives in the .sha256 sidecar + experiment marker — NEVER inlined into this body per the 2026-04-18 Gate-D 815e0a35 anti-propagation lesson)"
verified_against_code_on: 2026-07-06
verified_via: |
  Pre-registration authored 2026-07-07T~03:40Z UTC at CYC41 P3 D1 via /experiment Mode A.
  Grounded in: phase-open brain queries (MoE-Slimming/REAP SOTA 2026-07-01 captures + P3
  METHOD LOCKED decision + 122B P2-D5 MTP receipts a706e8c242104f89 + DFlash three-arm
  capture 2026-07-03 + deep research 2026-07-06), D7 PRE_REG precedent (whole-file read
  this session), p2d5 canary receipt (35B baseline 42.5 / draft-mtp 58.4 tok/s, 8088),
  models.ini deployed presets (whole-file read). Hypotheses H1-H5 await operator
  ratification BEFORE seal (§16 carries no hash until ratified).
signed_at: "2026-07-07T03:40:00Z"
signed_by: "Claude Fable 5 (claude-fable-5) + Jamey Kistner"
authors:
  - Jamey Kistner
  - Claude Fable 5 (claude-fable-5)
bid_relevance: high
related_docs:
  - RESEARCH/Roadmaps/ROADMAP_CYCLE41_SOVEREIGN_BIG_MODEL_COMPRESSION.md
  - RESEARCH/SPECS/RESULTS/cyc41_d7_pruned122b_quality_gate/PRE_REGISTRATION.md
  - DOCKER_ARMORY/reap-forge/run_p2d5_mtp_canary.py
  - REFINERY_MEMORY_SERVICE/SRC/training/eval_bank_runner.py
  - REFINERY_MEMORY_SERVICE/SRC/training/seal_prereg.py
phase_anchor: "CYC41 P3 D1 PRE_REG — 35B prune experiment: method fork (MoE-Slimming primary vs REAP comparator) x ratio -> resident serving + quality + MTP branch + drafter three-arm tau A/B. Experiment-id osint-suite-dev-2026-07-06-764b."
arc_status: active
changelog:
  - {date: "2026-07-06", change: "Authored at CYC41 P3 D1 via /experiment Mode A (DRAFT). H1-H5 + verdict rule pending operator ratification; seal follows ratification."}
  - {date: "2026-07-06", change: "H1-H5 + verdict rule + §9 bars operator-ratified verbatim 'Accepted. Run with it.' (pre-data). status DRAFT→FROZEN; canonical-prefix self-hash sealed via /sidecar Mode C (hex in .sha256 sidecar + experiment marker); freeze-gate armed on PRE_REGISTRATION.md."}
  - {date: "2026-07-07", change: "AMENDMENT_1 (DEV-1, operator-directed pre-primary-data): vision-preservation requirement (D2 retains visual.*; D4 vision smoke cell; mmproj sequencing) + H4 'dense'→MoE corrigendum (mtp block = 256-exp fused-3D MoE per index) + GGUF moe_intermediate_size uniformity risk registered. Pre-amendment self-hash 126a027a… preserved in AMENDMENT_1 §4; post-corrigendum hash recomputed at DEV-1 resolution."}
---

# PRE_REGISTRATION — CYC41 P3: 35B Prune Experiment (method fork × ratio × MTP × drafter τ A/B)

> **Companion docs:** `ROADMAP_CYCLE41` §1 Phase 3 (governing phase) · `D7 PRE_REG` (in-arc precedent: banks/sampling/stats inherited) · `run_p2d5_mtp_canary.py` (two-arm harness pattern + 35B baselines) · `eval_bank_runner.py` (imported extractors/stratifier/grader) · `EXPERIMENT_LEDGER.md` (deviation register) · `decisions.md` (deviation narrative).
> **Reading discipline:** Immutable mid-experiment per §11 once sealed. Post-seal changes ONLY via /experiment Mode C (ledger 🔧 + `decisions.md` + strikethrough-corrigendum). The `.sha256` sidecar is the tamper-evidence envelope (canonical-prefix per `seal_prereg.py`).
> **Spec authority:** `~/.claude/specs/EXPERIMENT_DISCIPLINE.md` v1.1 + `~/.claude/specs/SOVEREIGN_OPERATIONS_SPEC.md` §1.1/§4.3/§4.6/§4.7/§X. `.claude/rules/phase-claim-discipline.md` Rule 2 + `.claude/rules/attestation-class-files.md`.
> **Rules:** `attestation-class-files` · `brain-query-first` · `phase-claim-discipline` · `skill-enforcement`

> **Ground-Truth Attestation**
> - **Doc class:** attestation (pre-registration; SHA-256 canonical-prefix self-hash at seal)
> - **Status as of:** 2026-07-07T03:55Z UTC — FROZEN (H1-H5 operator-ratified pre-data, verbatim "Accepted. Run with it."; sealed)
> - **Verifier:** Claude Fable 5 (claude-fable-5) + Jamey Kistner
> - **Ledger:** `.claude/verified/RESEARCH__SPECS__RESULTS__cyc41_p3_35b_prune_drafter__PRE_REGISTRATION.verified.json` + this file's `.sha256` sidecar at seal
> - **Scope:** Pre-registers preflight gates, method fork, ratio grid, hypotheses, banks, sampling, statistics, thresholds, MTP branch rule, drafter three-arm τ design, frozen set, and deviation policy for the CYC41 P3 35B prune experiment.
> - **Re-verify cadence:** IMMUTABLE mid-experiment per §11 (once sealed).

---

## 1. Abstract

P3 prunes the EXACT deployed Architect base — `Qwen/Qwen3.6-35B-A3B` (MTP-bundled; served today as `[architect-apex]` UD-Q5_K_XL 27.16GB @ n-cpu-moe 32) — to a **fully-GPU-resident, on-box-trainable** size. Method fork (P3-METHOD-LOCKED decision 2026-07-01): **PRIMARY = MoE-Slimming** (arxiv 2606.18304; channel-level expert prune + light recovery-FT; runs NATIVE in venv-training; its recovery-FT is the P4 persona-LoRA double-duty) with **COMPARATOR = REAP-on-35B** (one-shot whole-expert; community pre-pruned GGUF; report-only). Anchor bench (Qwen3-30B-A3B, same family): 25%+4bit → 11.64GB, MATH500 94.5 (+1.7 vs base), MMLU 73.0 (−4.8); 50% craters general (MMLU 61.3) → ratio grid stays 25–33%. Four question families, kept separate: (A) **quality retention** vs the deployed Q5 baseline (deployment-retention framing per D7 — not a pure prune ablation); (B) **resident serving** — does pruned+4bit serve `-ngl 99` with NO `--n-cpu-moe` and beat the offloaded incumbent's 42.5 tok/s materially; (C) **MTP branch** — do the dense ~210M draft heads survive channel-prune (preserve/retrain/drop rule pre-registered; 122B prior: base-trained MTP head on 25%-pruned MoE trunk accepts 65–68.5%); (D) **drafter three-arm τ A/B** — A1 native MTP vs A2 stock `z-lab/Qwen3.6-35B-A3B-DFlash` (trained vs UNPRUNED base = deliberately misfitted) vs A3 DFlash distilled against the pruned target (RECIPE-GATED; branch resolved at preflight). Phase verdict = H1 ∧ H2 ∧ H3; H4 selects the MTP branch; H5 is the D5 evidence table (report-graded).

## 2. Hypotheses (operator-ratified 2026-07-07T03:52Z pre-data, verbatim "Accepted. Run with it.")

| ID | Statement | Primary statistic | Decision rule | Anchor |
|---|---|---|---|---|
| **H1 — math/reasoning retention** | Pruned-35B (primary cell, +recovery-FT, 4-bit) retains generative math vs deployed `[architect-apex]` UD-Q5_K_XL | Pooled accuracy over AIME24 (30/30) + MATH-500 (50/500 stratified) = 80 paired items; Δ = acc(pruned) − acc(baseline) | **PASS** if Δ ≥ −5.0 pp · **FAIL** if Δ < −5.0 pp AND McNemar exact two-sided p < 0.05 · else **INCONCLUSIVE** → operator ruling | MoE-Slimming 25%: MATH500 +1.7 over base; math is the strong axis |
| **H2 — knowledge retention** | Pruned-35B retains knowledge-MC vs the same baseline | GPQA-diamond FULL bank, n=198 paired; Δ = acc(pruned) − acc(baseline) | **PASS** if Δ ≥ −6.0 pp · **FAIL** if Δ < −6.0 pp AND McNemar exact two-sided p < 0.05 · else **INCONCLUSIVE** → operator ruling | Knowledge-MC is the weak axis (paper MMLU −4.8 pp @25%); −6.0 bar gives the documented drop headroom without accepting a crater |
| **H3 — fully-resident serving gate (mechanical)** | At least one grid cell (ratio ≤ 33%, 4-bit imatrix GGUF) serves **fully resident** and materially faster than the incumbent | b9789 on 8088: `-ngl 99`, **NO `--n-cpu-moe`**, c=32768, FA on, KV q8_0/q4_0 → /health 200 + no WDDM spill (dedicated-VRAM used ≤ 12227 MiB budget, no spill-cliff signature) + gen tok/s on the D6-style 3-prompt battery | **PASS** if resident-serve holds AND mean gen tok/s ≥ **63.8** (= 1.5 × 42.5 canary baseline) · **FAIL** if no cell ≤ 33% serves resident OR best resident cell < 63.8 tok/s · directional expectation ≥ 2× | 42.5 tok/s = p2d5 canary baseline (same GGUF class, n-cpu-moe 32); REAP-GGUF community bench +31–119% on residency gains |
| **H4 — MTP branch rule (selects branch; not a phase gate)** | The bundled ~~dense~~ **[corrigendum DEV-1, AMENDMENT_1 §2: MoE — the mtp block carries its own 256-expert fused-3D MoE MLP per base index ground truth 2026-07-07; rule unchanged]** MTP heads, carried through the prune (heads untouched; trunk hidden width unchanged by channel-prune), still draft usefully | On the H3-winning resident cell: `--spec-type draft-mtp --spec-draft-n-max 3` vs same-cell no-spec; draft-accept % (timings `draft_n_accepted/draft_n`) + tok/s delta | **KEEP-MTP** if accept ≥ 55% AND tok/s delta ≥ +10% · **RETRAIN-vs-DROP branch** (operator rules, informed by 122B D5 recipe cost) if accept < 55% OR delta < +10% · MTP tensors quant-pinned Q8_0 in ALL pruned GGUFs (CYC36 Q4 = 0%-accept + 122B D5 precedent) | 35B stock accepts 63–95% (canary); 122B base-trained head on 25%-pruned trunk 65–68.5%; fully-resident target REMOVES the 122B placement-inversion penalty (no GPU-expert competition) |
| **H5 — drafter three-arm τ (D5 evidence; report-graded, no phase-gate weight)** | (a) prune-induced drafter drift is real: τ(A2 stock DFlash, trained vs unpruned base) < τ(A3 DFlash distilled vs pruned target); (b) block-diffusion drafting beats native MTP on the same target: τ(A3) ≥ τ(A1 native MTP) | τ = mean accepted length per verify step (PRIMARY, hardware-independent) + concurrency-1 throughput (secondary); identical prompt battery + T=0 across arms; quality = lossless-by-construction invariant (verify-step guarantees) | Report table τ/throughput per arm. **RECIPE BRANCH (resolved at preflight G4):** z-lab trainer released → full three-arm; NOT released → two-arm A1-vs-A2 + drift documented as the honest bound, (a)/(b) marked UNTESTABLE-THIS-CYCLE | DFlash ICML 2026 (arxiv 2602.06036, MIT): 2.5× vs EAGLE-3, 6× lossless claim; A2 published for our EXACT target; llama.cpp PR #22105 PENDING; Qwen3.6-MoE hybrid-GDN DFlash suboptimal-in-llama.cpp caveat on record |

**P3 verdict rule (frozen at seal):** PASS = H1 PASS ∧ H2 PASS ∧ H3 PASS. Any FAIL → P3 FAIL (findings documented; incumbent `[architect-apex]` stays). Any INCONCLUSIVE (no FAIL) → operator adjudication in `decisions.md` before the verdict is written. H4 selects the MTP branch for D3/D5; H5 grades the D5 table only.

## 3. Preflight gates (pre-registered; each is a STOP, not a silent skip)

| Gate | Content | On FAIL / block |
|---|---|---|
| **G1 — license vet (§1.1)** | MoE-Slimming repo (`yifu-ding/MoE-Slimming`) license read BEFORE any clone/install. Allowlist Apache-2.0/MIT/BSD only. (DFlash already vetted MIT 2026-07-06.) | STOP, surface license, operator say-so |
| **G2 — venv-training import smoke** | MoE-Slimming import + tiny-prune smoke in venv-training (torch 2.10+cu128 / transformers 5.5). Requirements are `>=` lower-bounds — compatible on paper ≠ runs (5.x API-drift risk, flagged at METHOD-LOCK) | **METHOD BRANCH (pre-registered):** REAP-on-35B in the forge becomes PRIMARY; MoE-Slimming documented blocked; H1/H2 bars unchanged; recovery-FT axis reported as absent (REAP = one-shot) |
| **G3 — base checkpoint download** | `Qwen/Qwen3.6-35B-A3B` HF BF16 (~70GB) → `INFRA/UPGRADE_CANDIDATES/` staging (empty; 798.4GB free). SAY-SO-ONLY | No download without operator word |
| **G4 — drafter assets + recipe check** | A2 `z-lab/Qwen3.6-35B-A3B-DFlash` download (MIT; SAY-SO-ONLY) + z-lab training-recipe availability re-checked at D5-open (releases said "soon") → resolves the H5 recipe branch | Recipe absent → two-arm branch per H5 |
| **G5 — D5 runtime gate** | llama.cpp PR #22105 status + fork vet (dflash-llama / beellama.cpp) vs vLLM-in-forge (LINUX-ONLY; reap-forge containerized w/ GPU passthrough; measurement-harness duty ONLY — prod serving stays llama.cpp) | Neither viable → A2/A3 arms deferred with receipts; A1 (native MTP, llama.cpp b9789 proven) always runnable |

## 4. Materials + variables

### 4.1 Arms / artifacts

| Item | Path / source | Role |
|---|---|---|
| Baseline (quality + speed) | `INFRA/MODELS/Qwen3.6-35B-A3B/Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf` 27.16GB (deployed `[architect-apex]`) | H1/H2 paired baseline arm; H3 speed reference is the p2d5 canary 42.5 tok/s (n-cpu-moe 32, 8088) |
| Prune input | `Qwen/Qwen3.6-35B-A3B` BF16 HF checkpoint (G3 download) | MoE-Slimming input; MTP heads carried untouched |
| Primary pruned cells | ratio grid **{25%, 30%}** channel-prune + recovery-FT; **33% contingency** pre-authorized ONLY if 30%+4bit misses the resident line | D2 checkpoints + manifest sha256 each |
| Recovery-FT corpus | D3-mix general/reasoning corpus (CYC41 P1 D3; license-cleared, sovereign) — **NOT Alpaca** (CC-BY-NC fails §1.1) | MoE-Slimming recovery-FT stage |
| Pruned GGUFs | convert (NEXT tree, MTP bundled by default) → imatrix → Q4_K_M with `--tensor-type` Q8_0 pin on MTP/nextn tensors | H3/H4/H5 serving artifacts |
| Comparator | REAP-on-35B community pre-pruned GGUF (DJLougen / barozp; download SAY-SO-ONLY) at nearest matched ratio | Report-only method comparison (§6.4) |
| A2 drafter | `z-lab/Qwen3.6-35B-A3B-DFlash` (MIT), kept **Q8_0** (3.6 sliding-window layers Q4-fragile) | H5 arm |
| Server binary | `INFRA/llama.cpp_128/build_sm120_NEXT/bin/llama-server.exe` b9789 (`--version` in run manifest); same binary all llama.cpp arms | Control |

Hash discipline: every checkpoint/GGUF gets a HASH-FIRST `.sha256` sidecar at creation; driver preflights re-hash and HARD-FAIL on drift (D7 pattern). No hex is propagated symbolically into this doc.

### 4.2 Banks + sampling (inherited verbatim from D7 PRE_REG §3/§4 — one validated harness, zero new scope)

AIME24 30/30 · MATH-500 50/500 stratified (seed 20260422) · GPQA-diamond 198/198 (choice-shuffle `random.Random(seed*9973+row_idx)`). Endpoint `/v1/chat/completions` (deployed jinja surface, thinking-ON); T=0.6, top_p 0.95, top_k 20, seed 20260422, ONE generation per item per arm; budgets AIME24 12288 · MATH-500 6144 · GPQA 4096. Grading: rule-based extractors first, janitor (:8081, T=0) fallback, raw verdicts logged; empty/errored = incorrect (symmetric). Bank manifest with subset SHA-256 emitted at run start. All eval serving on **8088** (PROD 8080/8081/8082/8085 untouched); SynthLock per stage; per-item JSONL + `--resume`.

### 4.3 H4/H5 speed battery

The p2d5 canary 3-prompt battery (T=0, 256 tok, warmup 1×32) extended to 8 prompts spanning code/reasoning/summarize; single-sequence (spec path has no batching). τ computed from server timings (`draft_n`, `draft_n_accepted`) per verify step; vLLM arms (if G5 selects) report their native acceptance metrics with the cross-runtime caveat logged — **cross-runtime τ is compared descriptively only, never gated.**

## 5. Frozen set / gate discipline

`frozen_paths` at seal: **this `PRE_REGISTRATION.md` only** (the P3 drivers do not exist yet). Pre-registered policy: each eval/prune/serve driver is added to `frozen_paths` via an AMENDMENT co-seal (Mode C sub-flow) at its authoring, BEFORE the run it governs — mirrors D7 AMENDMENT_1 mechanics. Run outputs are never frozen. Cold-start resume per EXPERIMENT_DISCIPLINE §3.2.

## 6. Analysis plan + post-hoc policy

1. H1/H2: paired contingency → McNemar exact two-sided (scipy binomtest on discordant pairs); Wilson 95% CI per arm per bank; H1 pools 80 items, per-bank splits descriptive.
2. H3: tok/s mean over battery + VRAM receipt (`nvidia-smi` dedicated-used series) + spill signature check (the 44-cell 3.45 tok/s cliff class from P2 tune).
3. H4: accept% + delta vs same-cell no-spec; placement is FIXED fully-resident (no n-cpu-moe sweep — the 122B inversion driver is absent by design; if resident fails and offload cells are ever evaluated, a placement sweep is REQUIRED, pre-authorized).
4. Comparator (REAP-35B): same banks, report-only table; directional expectation MoE-Slimming(+FT) ≥ REAP one-shot at matched ratio. No gate weight.
5. **Post-hoc policy:** anything not named here is labeled `EXPLORATORY (post-hoc)` and cannot move the verdict. Pre-authorized descriptives: truncation rate, thinking-length, per-domain GPQA splits, per-prompt τ variance.
6. No mid-run peeking: no accuracy aggregation until an arm completes.

## 7. Procedure (stages; each under SynthLock, receipts + telemetry per stage)

1. **Stage G (preflight):** G1→G5 in order; each outcome logged in `decisions.md`. G2 FAIL executes the method branch immediately (no re-design).
2. **Stage P (prune):** MoE-Slimming attribution + channel-prune at 25%/30% + recovery-FT (D3-mix) → checkpoints + manifests (D2).
3. **Stage Q (quant/serve):** convert → imatrix → Q4_K_M (MTP pinned Q8_0) → H3 resident-serve receipts on 8088 (D4 evidence).
4. **Stage R (retention):** D7-pattern paired eval, pruned-primary-cell vs `[architect-apex]` baseline, banks per §4.2 (D3 evidence → H1/H2).
5. **Stage M (MTP):** H4 two-arm on the H3 winner (D3 MTP verdict).
6. **Stage T (drafter):** G4/G5-resolved arms; τ battery per §4.3 (D5 evidence → H5).
7. **Post-run:** Mode D — POST_RUN_INTEGRITY + FINAL_SEAL + verdict; roadmap ticks; brain finding; de-register.

## 8. Sample-size rationale

Identical to D7 (n=198 GPQA paired: ~70-80% power for a true 7-8 pp shift at plausible discordance; pooled-math n=80 point-estimate rule with McNemar guarding only the FAIL branch; single-seed acknowledged, 3-seed sensitivity pre-authorized as EXPLORATORY on any INCONCLUSIVE). τ battery n=8 prompts × 256 tok: τ is a per-token-step statistic (thousands of verify steps per arm) — prompt-level n is not the power constraint; per-prompt variance reported.

## 9. Thresholds (frozen at seal)

| Gate | Bar |
|---|---|
| H1 | Δ pooled-math ≥ −5.0 pp (FAIL needs Δ < −5.0 AND p < 0.05) |
| H2 | Δ GPQA-198 ≥ −6.0 pp (FAIL needs Δ < −6.0 AND p < 0.05) |
| H3 | resident serve (no `--n-cpu-moe`, no spill, ≤ 12227 MiB) AND mean gen tok/s ≥ 63.8 |
| H4 | KEEP-MTP iff accept ≥ 55% AND tok/s delta ≥ +10% (else branch to operator) |
| H5 | none (report-graded; recipe branch per §2) |
| Verdict | PASS = H1 ∧ H2 ∧ H3 |

## 10. Blinding / bias notes

No human grading (extractors + janitor; janitor never sees arm identity). Arms sequential on identical substrate, fresh server process per arm, identical prompts/seeds. No mid-run peeking (§6.6). The comparator and H5 table are report-only — no incentive gradient on the phase verdict.

## 11. Deviations policy

Immutable once sealed. Deviations: `/experiment --deviate osint-suite-dev-2026-07-06-764b` → narrate in `decisions.md` → flip the `EXPERIMENT_LEDGER.md` row 🔧 → strikethrough-corrigendum → resolve ✅. Post-data adjudications land in `## 99. Corrigendum` below the seal. Gate-D §11 norm: **document, don't silently patch — the discipline is the finding.**

## 12. Wall-clock + infra plan

- Stage P: paper prune cost 5.23 GPU-hr (their HW) + recovery-FT; on-box VRAM-to-prune is UNSTATED by the paper (honest unknown) — G2 tiny-smoke also measures it; offload-assisted attribution acceptable (quality-neutral, wall-clock cost only). Budget expectation: hours-to-a-day per cell.
- Stage R: ~3-6h/arm at 35B speeds (42+ tok/s) with thinking budgets; per-item resume makes interruption cheap.
- Stages Q/M/T: ~1-3h each. Disk: 798.4GB free vs ~70GB base + ~50GB checkpoints/GGUFs — ample. All GPU work under SynthLock with governor deferral; llama exes via PowerShell-tool/subprocess (never foreground Bash).

## 13. Pre-registered deviations from roadmap/precedent text

1. **Quality baseline = served UD-Q5_K_XL** (deployment-retention framing, D7 precedent) — NOT a BF16 or same-quant baseline; "prune-only cost" is not isolated (baseline carries Q5 damage, pruned arm carries prune+Q4+FT). The question answered: "what does the deployed Architect slot gain/lose by the swap."
2. **Knowledge bar −6.0 pp (vs D7's −5.0).** MoE-Slimming's own 25% cell shows −4.8 pp MMLU; a −5.0 bar would gate the method's documented best case on noise. −6.0 accepts the paper prior, still rejects a crater. (Operator may tighten at ratification.)
3. **Recovery-FT corpus = D3-mix, not paper's Alpaca/C4** — Alpaca is CC-BY-NC (§1.1 fail); D3-mix is sovereign, validated, and matches the research/general profile.
4. **H5 τ cross-runtime caveat** — if A2/A3 run on vLLM-in-forge while A1 runs on llama.cpp, τ comparisons across runtimes are descriptive only (different verify implementations); within-runtime pairs carry the load.

## 14. Limitations (honest scope)

1. Not a pure pruning ablation (§13.1); single-operator, single-box, single-seed primary battery.
2. MoE-Slimming benched Qwen3-30B-A3B-Thinking, NOT the 3.6-MTP — MTP survival through channel-prune is genuinely untested upstream (that's H4's job); the dense-head + unchanged-hidden-width prior is an argument, not evidence.
3. A2 is deliberately misfitted (trained vs unpruned base) — that mismatch IS the measurement (drafter-drift), not a flaw.
4. Qwen3.6-MoE hybrid-GDN DFlash currently SUBOPTIMAL in llama.cpp (upstream caveat); G5 may route drafter arms to vLLM-in-forge (measurement-only; prod stays llama.cpp).
5. 33% contingency cell, if fired, weakens comparability to the paper's 25% anchor; reported as its own row, never pooled.

## 15. Cross-references

- Experiment-id: `osint-suite-dev-2026-07-06-764b` · marker `~/.claude/hooks/experiments/osint-suite-dev-2026-07-06-764b.experiment.json` · registry `~/.claude/hooks/experiments/REGISTRY.json`
- Brain anchors: P3 METHOD LOCKED (2026-07-01T06:58Z decision) · pruning SOTA `70c9cc67`-era findings (2026-07-01T05:13Z/05:26Z) · 122B MTP receipts `a706e8c242104f89` · DFlash three-arm capture (2026-07-03T08:54Z) · triple-seal boundary `8377edde559a43c6`
- Precedents: D7 PRE_REG (banks/stats/sampling) · p2d5 canary receipt (`p2d5_mtp_canary_20260706T063641Z`) · p2d5 122B smoke (`p2d5_mtp_smoke122_20260706T160147Z`) · CYC36 Q4-0%-accept ruling

## 16. Seal

Seal state: **SEALED 2026-07-07T03:55Z** — §2 hypotheses H1-H5 + verdict rule + §9 bars operator-ratified pre-data (verbatim "Accepted. Run with it."). The canonical-prefix self-hash (LF-norm bytes `0..^## 16. Seal`, rstrip+`\n`, SHA-256 per `seal_prereg.py`) was computed via `Skill(/sidecar --hash-first)` at this moment; the hex lives in the `.sha256` sidecar beside this file + the experiment marker `~/.claude/hooks/experiments/osint-suite-dev-2026-07-06-764b.experiment.json` — NEVER inlined here (Gate-D `815e0a35` anti-propagation). This section sits below the hash boundary; frozen content is everything above the `## 16. Seal` line. Freeze-gate armed on `PRE_REGISTRATION.md` (drivers added via amendment co-seals per §5). Post-data adjudications land in `## 99. Corrigendum` below this section.
```

***

*Evidence & seals · 35B Prune Experiment (P3, mismatch control) · experiment `osint-suite-dev-2026-07-06-764b` · pre-registration reproduced verbatim from the sealed record; self-hash quoted from its sidecar · CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*
