> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/evidence-and-seals/pruned-122b-quality-gate-d7.md).

# Pruned-122B Quality Gate (D7)

*Evidence & seal for* [***17 · Sovereign Big-Model Compression***](/osintelligence/part-iv-the-evidence-what-worked/17-sovereign-big-model-compression.md)*. This is the pre-registered gate that cleared the REAP-0.25-pruned Qwen3.5-122B to replace the served research-brain: does the pruned model, at the quant tier one consumer GPU can serve it (Q4\_K\_M, 56.6 GB), retain the capability of the best full 122B the same box can serve (UD-Q2\_K\_XL, 41.8 GB)? Methodology frozen before any data; verdict adjudicated mechanically against the frozen bars.*

*New to how these seals work, read* [***Verifying a Seal***](/osintelligence/evidence-and-seals/verifying-a-seal.md) *first.*

**Experiment id** `osint-suite-dev-2026-07-04-4bd1` · **pre-registered** 2026-07-04T01:35Z · **final-sealed** 2026-07-06T04:00Z · **verdict D7 PASS** (H1 ∧ H2 ∧ H3).

## The seal (quoted from the sidecar, verify it yourself)

The pre-registration below is frozen. Its canonical-prefix self-hash, computed at seal time and quoted verbatim from `PRE_REGISTRATION.md.sha256`:

```
76b3bd1ca4c4753bedd876b2b97c47acae2d9f40da51d59577ff2c4ab22ec77c  PRE_REGISTRATION.md
prefix-bytes-hashed: 24516 · raw-file-bytes: 26599 · marker-position-lf: 24517
computed AT 2026-07-04T01:33:16Z (canonical-prefix per seal_prereg.py; hex lives in the sidecar + experiment marker, never inlined in the body)
```

Run the algorithm on [**Verifying a Seal**](/osintelligence/evidence-and-seals/verifying-a-seal.md) over the frozen text and you will reproduce this hex and byte count exactly. The post-run integrity envelope re-computed the same prefix hash at final seal (2026-07-06T03:55Z) and it matched the experiment marker exactly, the methodology was unchanged through the run. **AMENDMENT\_1** (DEV-1, operator-ratified before any primary data) re-tiered the H3 fidelity bars from the Q8\_0-era thresholds to Q4\_K\_M-appropriate ones; its pre-amendment self-hash `c0882784…` is preserved in the amendment. The re-tier is shown as a strikethrough corrigendum in §2/§9 of the frozen text below, never silently merged.

## Verdict (final seal, adjudicated on the clean dataset against the sealed rules)

**D7 PASS — H1 ∧ H2 ∧ H3 all PASS.** The pruned model retains or exceeds the best servable full 122B on every pre-registered bank, while serving faster (D6: 16.65 vs 15.60 tok/s) and at two quant tiers of better fidelity. Cleared to replace the served research-brain.

| Gate                                 | Rule (sealed / AMENDMENT\_1)                                                                               | Measured (at-seal)                                                        | Verdict        |
| ------------------------------------ | ---------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------- | -------------- |
| **H1** math retention                | Δ pooled-math ≥ −5.0 pp (n=80 paired)                                                                      | pruned 67.5% vs full 65.0% → **Δ +2.5 pp** (McNemar p=0.754)              | **PASS**       |
| **H2** knowledge retention (primary) | Δ GPQA-diamond ≥ −5.0 pp (n=198 paired)                                                                    | pruned 41.92% vs full 40.91% → **Δ +1.01 pp** (discordant 39/37, p=0.909) | **PASS**       |
| **H3** quant fidelity                | KL mean ≤ 0.05 · max ≤ 7.6 · p95 ≤ 0.5 · top-1 ≥ 90% · \|PPL Δ\| ≤ 5% (Q4 vs Q8\_0, wikitext2, 100 chunks) | 0.0240 · 3.23 · 0.088 · 93.21% · +1.47%                                   | **PASS (5/5)** |
| H4 (report-only)                     | expectation PPL ratio ≤ 1.10                                                                               | wikitext2 1.354 (missed) · held\_out\_sovereign 1.027 (met)               | reported       |

Per-bank splits (at-seal): AIME24 23.3% vs 16.7% (Δ +6.7 pp, n=30) · MATH-500 94.0% vs 94.0% (Δ 0.0, n=50) · GPQA 41.9% vs 40.9%.

**Honest scope (pre-authorized, descriptive):** (1) GPQA generations hit the 4096-token budget on \~100% of items in both arms and AIME \~90% at 12288, so absolute scores understate both models symmetrically, the paired Δ is the valid signal. (2) H4 divergence is the known REAP pattern: general-web-text PPL is 35% higher while domain text is at parity (1.027) and task accuracy is retained everywhere measured, perplexity overstates task damage. (3) This gate compares deployable-vs-deployable (prune+Q4 vs Q2); it does not isolate prune-only cost (a full-Q4 baseline was disk-infeasible, pre-registered limitation §14.1).

**Deviations through the run:** DEV-1 (pre-primary-data, operator-ratified) re-tiered the H3 bars and fixed a parse path, co-sealed as AMENDMENT\_1. DEV-2 (run-integrity): 45 full-arm GPQA cells were error-poisoned by an author server-kill under a mis-diagnosed detached driver; they were quarantined (the poisoned summary preserved as INVALID, a byte-backup kept), 45 cells cleanly re-run, and H2 adjudicated only on the clean set. Both are narrated in the experiment's `decisions.md`; no open ledger rows at seal.

## Integrity envelope (11 artifacts, SHA-256, computed 2026-07-06T03:55Z)

Each run artifact carries a `.sha256` sidecar (authoritative); the whole-file hashes recorded at final seal, first 16 hex:

| Artifact                                                           | Bytes   | sha256\[:16]       |
| ------------------------------------------------------------------ | ------- | ------------------ |
| `d7_results.jsonl` (clean; 598 records / 556 unique cells)         | 590,408 | `8a6dcfb53ecae28c` |
| `d7_retention_summary.json` (clean, adjudication source)           | 1,869   | `7e85ae47dd1b83ed` |
| `d7_fidelity_summary.json` (Stage F, H3/H4)                        | 1,902   | `93d87e65ca9bd9c3` |
| `d7_bank_manifest.json` (banks + subset SHAs)                      | 3,998   | `22531c4d84c8fbcd` |
| `d7_results.quarantine_DEV2.jsonl` (53 error records)              | 27,705  | `5354015a1761f3e7` |
| `d7_results.pre-DEV2-backup.jsonl` (pre-remediation byte copy)     | 603,424 | `b3358adc5ea102eb` |
| `d7_retention_summary.INVALID-DEV2.json` (poisoned; preserved)     | 1,869   | `cd963b6677d17c35` |
| `telemetry/d7_apex_…/host_telemetry.csv`                           | 21,988  | `977330f523f061d7` |
| `telemetry/d7_retention_…020909Z/run_manifest.json` (run 1)        | 1,510   | `38bedf7a93b2e654` |
| `telemetry/d7_retention_…184221Z/run_manifest.json` (resume)       | 1,510   | `315765cf3a1c13bd` |
| `telemetry/d7_retention_…004246Z/run_manifest.json` (DEV-2 re-run) | 1,510   | `2ad6ff1604452337` |

Run tally: 556 unique cells (598 records) + a 7-run fidelity battery; wall ≈ 46 h under SynthLock across four fires.

## The frozen pre-registration, verbatim

Reproduced exactly as sealed (the bytes the self-hash above is computed over). The strikethrough passages in §2 and §9 are the AMENDMENT\_1 corrigendum, preserved as the record carries them.

```md
---
doc_class: attestation
status: FROZEN
status_as_of: 2026-07-04
frozen_at: "2026-07-04T01:35Z (SHA-256 canonical-prefix self-hash via /sidecar Mode C at seal moment; hex lives in the .sha256 sidecar + experiment marker — NEVER inlined into this body per the 2026-04-18 Gate-D 815e0a35 anti-propagation lesson)"
verified_against_code_on: 2026-07-04
verified_via: |
  Pre-registration authored 2026-07-04T~01:10Z UTC at CYC41 P1 D7 via /experiment Mode A +
  research-first reads (eval_bank_runner.py argparse/probes/extractors/grader read at lines
  100-410, 2866-3005, 3040-3220, 3634-3763; llama_server_client.py whole-file; Cycle-25 Gate-C
  APEX precedent + D4 ratio research + D6 receipts recalled from MCP-Memory). Hypotheses H1-H4
  await operator ratification BEFORE seal (§16 carries no hash until ratified). Arms are the
  D6-delivered GGUFs; their D6 sha256 sidecars are cross-verified by the driver preflight
  re-hash at run start (HARD-FAIL on drift) — no hash in this doc is propagated without a
  re-runnable verification path.
signed_at: "2026-07-04T01:10:00Z"
signed_by: "Claude Fable 5 (claude-fable-5) + Jamey Kistner"
authors:
  - Jamey Kistner
  - Claude Fable 5 (claude-fable-5)
bid_relevance: high
related_docs:
  - RESEARCH/Roadmaps/ROADMAP_CYCLE41_SOVEREIGN_BIG_MODEL_COMPRESSION.md
  - DOCKER_ARMORY/reap-forge/run_d7_retention.py
  - DOCKER_ARMORY/reap-forge/run_d7_apex_fidelity.py
  - DOCKER_ARMORY/reap-forge/run_d6_serve_test.py
  - REFINERY_MEMORY_SERVICE/SRC/training/eval_bank_runner.py
  - REFINERY_MEMORY_SERVICE/SRC/training/llama_server_client.py
  - REFINERY_MEMORY_SERVICE/SRC/training/seal_prereg.py
  - DOCKER_ARMORY/reap-forge/pruned-out/Qwen3.5-122B-A10B-REAP-0.25/d5_prune_manifest.json
phase_anchor: "CYC41 P1 D7 PRE_REG — pruned-122B quality gate: deployment-retention (AIME24/GPQA-diamond-198/MATH-500 vs full-UD-Q2_K_XL on 8088) + APEX quant-fidelity (Q4/Q5 vs Q8_0 reference via llama-perplexity --kl-divergence). Experiment-id osint-suite-dev-2026-07-04-4bd1."
arc_status: active
changelog:
  - {date: "2026-07-04", change: "Authored at CYC41 P1 D7 via /experiment Mode A. Three roadmap-text deviations pre-registered in §13 (MMLU-Pro→GPQA-198 proxy; KL reference = Q8_0 host; /v1/chat/completions endpoint vs P8's raw /completion)."}
  - {date: "2026-07-04", change: "H1-H4 + verdict rule operator-ratified verbatim 'Prereg looks good. Approved.' (pre-data). status DRAFT→FROZEN; canonical-prefix self-hash sealed via /sidecar Mode C (hex in .sha256 sidecar + experiment marker); freeze-gate armed on 3 frozen paths."}
  - {date: "2026-07-04", change: "AMENDMENT_1 (DEV-1, operator-ratified pre-primary-data): §2 H3 + §9 bars re-tiered Q8_0-era→Q4_K_M-appropriate (strikethrough corrigendum; smoke disclosure in AMENDMENT_1 §2 + decisions.md). Pre-amendment self-hash c0882784… preserved in AMENDMENT_1 §4; post-corrigendum hash recomputed at DEV-1 resolution."}
---

# PRE_REGISTRATION — CYC41 D7: Pruned-122B Quality Gate (deployment retention + quant fidelity)

> **Companion docs:** `ROADMAP_CYCLE41` §1 D7 (governing deliverable) · `run_d7_retention.py` (retention driver; FROZEN at seal) · `run_d7_apex_fidelity.py` (fidelity driver; FROZEN at seal) · `eval_bank_runner.py` (imported components: stratifier L2940-3005 / extractors L2866-2899 / janitor grader L2902-2937) · `llama_server_client.py` (`LlamaServerClient`) · `run_d6_serve_test.py` (server-lifecycle pattern + COMMON_ARGS provenance) · `EXPERIMENT_LEDGER.md` (deviation register) · `decisions.md` (deviation narrative).
> **Reading discipline:** This pre-registration is **immutable mid-experiment** per §11. Post-seal changes land ONLY via /experiment Mode C (ledger row 🔧 + `decisions.md` narrative + strikethrough-corrigendum). The `.sha256` sidecar beside this file is the tamper-evidence envelope (canonical-prefix algorithm per `seal_prereg.py`).
> **Spec authority:** `~/.claude/specs/EXPERIMENT_DISCIPLINE.md` v1.1 + `~/.claude/specs/SOVEREIGN_OPERATIONS_SPEC.md` §1.1/§4.3/§4.6/§4.7/§X. `.claude/rules/phase-claim-discipline.md` Rule 2 (HASH-FIRST) + `.claude/rules/attestation-class-files.md`.
> **Rules:** `attestation-class-files` · `brain-query-first` · `phase-claim-discipline` · `skill-enforcement`

> **Ground-Truth Attestation**
> - **Doc class:** attestation (pre-registration; SHA-256 canonical-prefix self-hash at seal)
> - **Status as of:** 2026-07-04T01:35Z UTC — FROZEN (H1-H4 operator-ratified pre-data; sealed)
> - **Verifier:** Claude Fable 5 (claude-fable-5) + Jamey Kistner
> - **Ledger:** `.claude/verified/RESEARCH__SPECS__RESULTS__cyc41_d7_pruned122b_quality_gate__PRE_REGISTRATION.verified.json` + this file's `.sha256` sidecar at seal
> - **Scope:** Pre-registers hypotheses, arms, banks, sampling, statistics, thresholds, frozen set, and deviation policy for the CYC41 D7 quality gate on the REAP-0.25-pruned Qwen3.5-122B-A10B. Two claim families kept SEPARATE per brain `747af810`: (A) deployment retention (pruned-Q4_K_M vs full-UD-Q2_K_XL, both served); (B) quant fidelity (Q4/Q5 vs pruned-Q8_0 reference).
> - **Re-verify cadence:** IMMUTABLE mid-experiment per §11.

---

## 1. Abstract

CYC41 D5 pruned Qwen3.5-122B-A10B to 192/256 experts (REAP saliency, ratio 0.25, text-only; manifest sidecar `07fd3600…` at-D5-seal); D6 delivered imatrix-quantized GGUFs and an operational serve receipt (pruned-Q4_K_M 16.65 tok/s vs full-UD-Q2_K_XL 15.60 tok/s on test-bind 8088). D7 is the pre-registered **quality gate**: does the pruned model, at the quant tier this box can serve it (Q4_K_M, 56.6GB), retain the capability of the best FULL 122B this box can serve (UD-Q2_K_XL, 41.8GB)? This is explicitly a **deployment-retention** comparison — the full arm carries Q2-tier quantization damage and the pruned arm carries prune-plus-Q4 damage; the question answered is "what does this box gain or lose by swapping the served research-brain," NOT "what does pruning alone cost" (that isolation would require a full-122B Q4_K_M baseline, disk-infeasible at 198GB free; acknowledged limitation §14). A second, separate claim family gates **quant fidelity**: Q4_K_M/Q5_K_M vs the pruned-Q8_0 reference under the Cycle-25 APEX battery (KL/PPL/top-1 via `llama-perplexity --kl-divergence`). Expectations are anchored by D4 ratio research (REAP paper Table 2, Qwen3-30B-A3B @25%: math +0.1, MC −5.6): generative reasoning ~lossless, knowledge-MC modestly down — hence GPQA-diamond at full n=198 as the knowledge probe with real paired power. Verdict = H1 ∧ H2 ∧ H3.

## 2. Hypotheses (operator-ratified 2026-07-04 pre-data, verbatim "Prereg looks good. Approved.")

All directional, falsifiable, thresholds fixed before any data. Retention family (H1, H2) adjudicated per-hypothesis (two pre-registered tests, no family-wise correction needed for a conjunction gate — BOTH must individually pass; this is conservative). H3 is a mechanical gate. H4 is registered exploratory (report-only).

| ID | Statement | Primary statistic | Decision rule | Anchor |
|---|---|---|---|---|
| **H1 — math/reasoning retention** | Pruned-Q4_K_M retains generative math reasoning vs full-UD-Q2_K_XL | Pooled accuracy over AIME24 (30/30) + MATH-500 (50/500 stratified) = 80 paired items; Δ = acc(pruned) − acc(full) | **PASS** if Δ ≥ −5.0 pp · **FAIL** if Δ < −5.0 pp AND McNemar exact two-sided p < 0.05 · else **INCONCLUSIVE** → operator ruling | REAP@25% math ≈ lossless (D4, brain `747af810`); Q4-vs-Q2 fidelity gap favors pruned arm |
| **H2 — knowledge retention (PRIMARY)** | Pruned-Q4_K_M retains knowledge-MC vs full-UD-Q2_K_XL | GPQA-diamond FULL bank accuracy, n=198 paired items; Δ = acc(pruned) − acc(full) | **PASS** if Δ ≥ −5.0 pp · **FAIL** if Δ < −5.0 pp AND McNemar exact two-sided p < 0.05 · else **INCONCLUSIVE** → operator ruling | Knowledge-MC is REAP's weak axis (−4 to −6 pp @25% vs unquantized full); n=198 paired gives usable power where n=30 stratified could not |
| **H3 — quant fidelity (mechanical gate)** | Q4_K_M preserves the pruned model's next-token distribution within the APEX envelope vs the Q8_0 reference | `llama-perplexity --kl-divergence` on wikitext2_test (100 chunks, ctx 512): KL mean / KL max / KL p95 / top-1 / PPL ratio | ~~PASS iff KL mean ≤ 0.008 AND KL max ≤ 7.6 AND KL p95 ≤ 0.5 AND \|PPL Δ\| ≤ 1.0% (Cycle-25 Gate-C bars, frozen). Top-1 agreement report-only.~~ **[AMENDMENT_1, DEV-1, operator-ratified pre-primary-data]** PASS iff KL mean ≤ 0.05 AND KL max ≤ 7.6 AND KL p95 ≤ 0.5 AND same-top-1 ≥ 90% AND \|PPL Δ\| ≤ 5.0%. Q5_K_M: report-only (no gate) | ~~Cycle-25 APEX 4/4 precedent (2B: KL mean 0.001004); bars transfer as-is~~ Cycle-25 bars were Q8_0-tier (KL mean ~0.001); Q4_K_M-tier bars per AMENDMENT_1 §2 (community-typical Q4_K_M: KL mean ~0.02–0.04, top-1 ~93%) |
| **H4 — deployment PPL (exploratory, report-only)** | Pruned-Q4 language-modeling quality is comparable to full-Q2 | Plain PPL, both arms, wikitext2_test + held_out_sovereign (100 chunks, ctx 512 each) | No gate. Directional expectation: PPL(pruned-Q4) ≤ PPL(full-Q2) × 1.10. Reported with both values + ratio | Cross-model PPL on identical tokenizer is comparable; no precedent bar exists → honest report-only |

**D7 verdict rule (frozen):** PASS = H1 PASS ∧ H2 PASS ∧ H3 PASS. Any FAIL → D7 FAIL (roadmap D7 flips 🚫 with findings; pruned model does NOT replace the served research-brain). Any INCONCLUSIVE (with no FAIL) → operator adjudication, documented in `decisions.md` before the verdict is written.

## 3. Variables

| Role | Name | Operationalization |
|---|---|---|
| Independent | Served model identity | `pruned-Q4_K_M` (Qwen3.5-122B-A10B-REAP-0.25-Q4_K_M.gguf) vs `full-UD-Q2_K_XL` (Qwen3.5-122B-A10B-UD-Q2_K_XL.gguf) |
| Dependent (H1) | Pooled math accuracy | Correct/total over 80 paired items (rule-based extractor first; janitor fallback per §4.4) |
| Dependent (H2) | GPQA-diamond accuracy | Correct/total over 198 paired items (letter extractor; janitor fallback) |
| Dependent (H3) | APEX KL/PPL battery | Parsed from `llama-perplexity --kl-divergence` stdout |
| Dependent (H4) | Plain PPL | Parsed from `llama-perplexity` stdout per arm/corpus |
| Controls | Server binary | `INFRA/llama.cpp_128/build_sm120_NEXT/bin/llama-server.exe` b9789 (`--version` recorded in run manifest); same binary both arms |
| Controls | Server args | D6 COMMON_ARGS verbatim: `--port 8088 --host 127.0.0.1 -ngl 99 --n-cpu-moe 48 -c 32768 --flash-attn on --jinja -t 8 -b 2048 -ub 2048 --cache-type-k q8_0 --cache-type-v q4_0` (mirrors `[research-brain]` deployment preset; KV-quant applied identically to both arms) |
| Controls | Endpoint + sampling | `/v1/chat/completions` (server-applied jinja chat template = deployment path); temperature 0.6, top_p 0.95, top_k 20 (Qwen3.5 recommended thinking-mode sampling), seed **20260422**, ONE generation per item per arm |
| Controls | Token budgets (`max_tokens`) | AIME24 12288 · MATH-500 6144 · GPQA 4096 (thinking traces included in budget; truncation counts as-graded per extractor fallback, symmetric across arms) |
| Controls | Bank subset seed | 20260422 (stratifier seed; subset SHA-256 emitted in bank manifest at run start) |
| Controls | GPQA choice-shuffle | Deterministic `random.Random(seed*9973 + row_idx)` — verbatim `eval_bank_runner.py` L3171-3178, identical across arms |
| Confounds (acknowledged) | Quant-tier asymmetry | Full arm = Q2-tier damage, pruned arm = prune+Q4 damage. Inherent to deployment framing (§14.1) |
| Confounds | Sampled decoding | T=0.6 single-seed: per-item stochasticity; mitigated by pairing (same seed + same server sampler both arms) + McNemar on paired outcomes |
| Confounds | Janitor grader | Qwen3.5-2B judge on :8081 for extractor-ambiguous items only; identical grader both arms; grader verdicts logged raw |

## 4. Materials

### 4.1 Arms (artifacts on disk; driver preflight re-hashes and HARD-FAILS on sidecar mismatch)

| Arm | Path | Size | sha256 sidecar (at-D6-seal 2026-07-03) |
|---|---|---|---|
| pruned-Q4_K_M | `INFRA/MODELS/Qwen3.5-122B-Research-REAP/Qwen3.5-122B-A10B-REAP-0.25-Q4_K_M.gguf` | 56.6GB | `6cee1b6e…` — re-verified by preflight |
| full-UD-Q2_K_XL | `INFRA/MODELS/Qwen3.5-122B-Research/Qwen3.5-122B-A10B-UD-Q2_K_XL.gguf` | 41.8GB | preflight computes + records (no prior sidecar requirement) |
| Q8_0 reference (H3 only) | `INFRA/MODELS/Qwen3.5-122B-Research-REAP/Qwen3.5-122B-A10B-REAP-0.25-Q8_0.gguf` | 99.0GB | `a246093b…` — re-verified by preflight |
| Q5_K_M (H3 report-only) | `INFRA/MODELS/Qwen3.5-122B-Research-REAP/Qwen3.5-122B-A10B-REAP-0.25-Q5_K_M.gguf` | 66.3GB | `2e2e1782…` — re-verified by preflight |

Hash discipline: the hexes above are pointers to D6 sidecars, not claims; the binding claim is the **preflight re-hash at run start** (chunked-stream SHA-256, recorded in `run_manifest.json`, HARD-FAIL on mismatch). Per `phase-claim-discipline` Rule 2 this gives every artifact a re-runnable verification path without propagating symbolic hashes.

### 4.2 Task banks (HF datasets, loaded exactly as `eval_bank_runner.py` L3090-3103)

| Bank | HF id / split | N used | Sampling |
|---|---|---|---|
| AIME24 | `HuggingFaceH4/aime_2024` train | 30/30 | all |
| GPQA-diamond | `Idavidrein/gpqa` `gpqa_diamond` train (HF_TOKEN-gated; token present, verified 2026-07-04) | 198/198 | all (deviation from the 30-stratified P8 grid — §13.1) |
| MATH-500 | `HuggingFaceH4/MATH-500` test | 50/500 | stratified by level, seed 20260422 (`_tts_compute_stratified_subset` imported verbatim) |

Bank manifest (`d7_bank_manifest.json`) emitted at run start: per-bank row counts, subset index lists, subset SHA-256, per-bank content SHA-256 (sorted canonical JSON of used rows). Prompts: verbatim `eval_bank_runner.py` L3153-3195 templates (boxed-answer instruction per task).

### 4.3 Harness (frozen at seal)

- `DOCKER_ARMORY/reap-forge/run_d7_retention.py` — serves each arm sequentially on **8088** (PROD 8080/8081/8082/8085 untouched), SynthLock window, PID-captured teardown in `finally`, file-handle logs (never PIPE), per-item JSONL checkpointing (`d7_results_<arm>.jsonl`) with `--resume` (completed `(arm, task, idx)` cells skipped), host telemetry (psutil 5s CSV) + `nvidia-smi dmon`.
- `DOCKER_ARMORY/reap-forge/run_d7_apex_fidelity.py` — `llama-perplexity` chains (b9789, NEXT/bin) via subprocess with logged output; Stage F1 baseline `.kld` from Q8_0; F2/F3 Q4/Q5 `--kl-divergence` runs; F4 plain-PPL runs (pruned-Q4 + full-Q2 × 2 corpora). GPU exes launched per Cycle-25 execution finding (PowerShell-tool / subprocess, never foreground Bash).
- Imported (NOT copied) from `eval_bank_runner.py`: `_tts_compute_stratified_subset`, `_aime24_extract_integer`, `_gpqa_extract_letter`, `_math_extract_pred`, `_math_is_equiv`, `_janitor_grade`. Imported from `llama_server_client.py`: `LlamaServerClient`. Any import-surface change is a frozen-path deviation on the D7 drivers only (upstream files are NOT frozen — P8 artifacts stay untouched).

### 4.4 Grading (identical both arms)

1. Rule-based extractor per task (boxed-first, fallback regex — imported fns above); MATH-500 equivalence via `_math_is_equiv`.
2. If the extractor returns `None` (or MATH equivalence is not established), `_janitor_grade` (janitor :8081, T=0, 8 tokens) adjudicates; raw verdict logged.
3. An item is `correct ∈ {0,1}`; empty/errored generations grade incorrect (fail-soft symmetric, per P8 precedent).

### 4.5 Fidelity corpora

`INFRA/MODELS/TRAINING_SETS/perplexity_test/wikitext2_test.txt` (1,299,261 B) + `held_out_sovereign.txt` (10,441,707 B); sizes verified 2026-07-04; SHA-256 of both recorded in `run_manifest.json` at run start. H3 gate adjudicated on **wikitext2 only** (Cycle-25 precedent); held_out_sovereign KL/PPL report-only.

## 5. Frozen set / gate discipline

`frozen_paths` (experiment_gate.py enforces post-seal): this `PRE_REGISTRATION.md` · `run_d7_retention.py` · `run_d7_apex_fidelity.py`. Run outputs (JSONL, manifests, telemetry, verdict docs) are NOT frozen. Post-seal edits to frozen paths ONLY via Mode C (ledger 🔧 + `decisions.md` + corrigendum). Cold-start resume per EXPERIMENT_DISCIPLINE §3.2: any session resuming the run re-reads this doc + the ledger first.

## 6. Analysis plan + post-hoc policy

1. Per-bank contingency: paired item outcomes (both-correct / both-wrong / pruned-only / full-only) → McNemar exact two-sided (scipy binomtest on discordant pairs, computed by the analysis block in `run_d7_retention.py`).
2. Wilson 95% CI per arm per bank (reported alongside every accuracy).
3. H1 pools AIME24+MATH-500 (80 items); per-bank splits reported descriptively.
4. H3/H4 parsed from `llama-perplexity` stdout; raw logs retained in telemetry dir.
5. **Post-hoc policy:** any analysis not named here is labeled `EXPLORATORY (post-hoc)` in the findings doc and cannot move the verdict. Truncation-rate, thinking-length, and per-domain GPQA splits are pre-authorized descriptive statistics (no gate weight).
6. No mid-run peeking rule: per-item JSONL accumulates, but no accuracy is aggregated or reported until an arm completes (monitor surfaces progress counts + errors only).

## 7. Procedure

1. Preflight: arm re-hash vs sidecars, corpora hash, VRAM/disk check, janitor :8081 health, HF_TOKEN presence, bank load + manifest emit. HARD-FAIL on any mismatch.
2. **Stage R (retention):** SynthLock(`reap_d7_retention`) → serve pruned-Q4_K_M on 8088 → warmup 1 gen → AIME24(30) → MATH-500(50) → GPQA(198) with per-item JSONL → teardown → serve full-UD-Q2_K_XL → same banks, same order → teardown → analysis block → `d7_retention_summary.json`.
3. **Stage F (fidelity):** SynthLock(`reap_d7_apex`) → F1 Q8_0 baseline `.kld` (wikitext2, 100 chunks, ctx 512; ~13GB, disk-checked) → F2 Q4_K_M `--kl-divergence` → F3 Q5_K_M `--kl-divergence` → F4 plain PPL: pruned-Q4 + full-Q2 on wikitext2 + held_out_sovereign; held_out KL optional-if-disk (report-only either way) → `d7_fidelity_summary.json`. Stages R and F are independent and may run in either order (F first is cheaper smoke of the substrate).
4. Post-run: /experiment Mode D — POST_RUN_INTEGRITY.md (sha envelope) + FINAL_SEAL.md + verdict; roadmap D7 tick; brain finding; de-register.

## 8. Sample-size rationale

- GPQA n=198 paired: McNemar power depends on discordant rate; at a plausible 25% discordance (~50 pairs), detecting a true 7-8 pp one-sided shift at α=0.05 has ~70-80% power; n=30 (P8 grid) would have been decorative for H2 — this is why the full bank is pre-registered.
- Math pooled n=80: adequate for the −5 pp bar as a point-estimate rule; McNemar guards only the FAIL branch (requiring significance to declare failure protects against noise-driven kills of a passing model).
- Single-seed sampled decoding: variance acknowledged; pairing + identical seed/sampler across arms removes the between-arm sampler asymmetry. A 3-seed sensitivity re-run is pre-authorized as EXPLORATORY if any hypothesis lands INCONCLUSIVE.

## 9. Thresholds (frozen at seal)

| Gate | Bar |
|---|---|
| H1 | Δ pooled-math ≥ −5.0 pp (FAIL needs Δ < −5.0 pp AND p < 0.05) |
| H2 | Δ GPQA-198 ≥ −5.0 pp (FAIL needs Δ < −5.0 pp AND p < 0.05) |
| H3 | ~~KL mean ≤ 0.008 · KL max ≤ 7.6 · KL p95 ≤ 0.5 · \|PPL Δ\| ≤ 1.0%~~ **[AMENDMENT_1]** KL mean ≤ 0.05 · KL max ≤ 7.6 · KL p95 ≤ 0.5 · top-1 ≥ 90% · \|PPL Δ\| ≤ 5.0% (Q4_K_M vs Q8_0, wikitext2, 100 chunks ctx 512) |
| H4 | none (report-only; directional expectation PPL ratio ≤ 1.10) |
| Verdict | PASS = H1 ∧ H2 ∧ H3 |

## 10. Blinding / bias notes

No human grading in the loop (rule-based + janitor; janitor sees only problem/reference/candidate text, never the arm identity). Arms run sequentially on identical substrate; order fixed (pruned first) and identical prompts/seeds — order effects on a stateless server restart are nil (fresh process per arm). The author adjudicates nothing mid-run (§6.6 no-peeking).

## 11. Deviations policy

This document is immutable mid-experiment. Deviations: `/experiment --deviate osint-suite-dev-2026-07-04-4bd1` → narrate in `decisions.md` → flip the artifact row in `EXPERIMENT_LEDGER.md` to 🔧 → corrigendum edit (strikethrough, never mutate) → resolve ✅. Post-data scientific adjudications land in `## 99. Corrigendum` below the seal. The P7 Gate-D §11 norm governs: **document, don't silently patch — the discipline is the finding.**

## 12. Wall-clock + infra plan

- Stage F: ~2-4h (Q8 baseline ~20-30 min + 2 KL runs + 4 PPL runs at Q8-class speeds measured in D6).
- Stage R: expected ~10-14h per arm at 15-17 tok/s gen (D6 receipt) with thinking-mode budgets; worst-case ~25h/arm if every gen exhausts budget. SynthLock `expected_sec=129600` (36h) per stage-R fire; per-item resume makes interruption cheap. Governor defers crons/self-heal for the window; monitors follow the D5/D6 pattern.
- Disk: 198GB free (verified 2026-07-04); baseline `.kld` ~13GB; margin ample. R: not used for load-bearing artifacts (ephemeral).

## 13. Pre-registered deviations from roadmap/precedent text

1. **MMLU-Pro → GPQA-diamond full-198.** Roadmap D7 names "MMLU-Pro/GPQA". MMLU-Pro has no processor in the eval harness (adding one mid-arc = unvalidated scope) and the existing MMLU probe is HF-in-process (cannot drive GGUF arms). GPQA-diamond at full n=198 is the validated knowledge-MC probe with materially better paired power than the P8 30-stratified grid. MMLU-Pro remains open for a follow-up cycle.
2. **KL reference = pruned-Q8_0, not pruned-BF16.** BF16-direct hosting measured 1063.78 s/pass at D6 (≈118h-wall class) — infeasible. Q8_0 is near-lossless (Cycle-25 APEX: KL mean ~0.001 vs f16, two orders below the H3 bar), so bars retain meaning; the H3 claim is stated as "Q4 vs Q8_0 reference" — never silently upgraded to "vs BF16".
3. **`/v1/chat/completions` instead of P8's raw `/completion`.** The P8 grid sent raw prompts (base-model era). The 122B research-brain is deployed behind the jinja chat template; evaluating through the chat endpoint measures the deployed surface (thinking-mode included). Both arms identical, so the comparison is internally valid; cross-experiment comparability to P8 absolute numbers is NOT claimed.

## 14. Limitations (honest scope)

1. **Not a pure pruning ablation.** Quant tiers differ across arms by necessity (§1); "prune-only cost" is unmeasurable on this box without a full-Q4 baseline (disk-infeasible now; possible after Q8_0 deletion post-D7 — registered as optional follow-up, not part of this gate).
2. Single-operator, single-box, single-seed; no multi-seed CI on generative banks in the primary battery.
3. GPQA-diamond ≠ MMLU-Pro breadth; knowledge coverage is science-domain-weighted.
4. Janitor (2B) grading of extractor-ambiguous math is imperfect; identical across arms, raw verdicts retained for audit.
5. τ/speculative-drafter questions are OUT of scope (P3 three-arm design registered separately, brain `8a184516`).

## 15. Cross-references

- Experiment-id: `osint-suite-dev-2026-07-04-4bd1` · marker `~/.claude/hooks/experiments/osint-suite-dev-2026-07-04-4bd1.experiment.json` · registry `~/.claude/hooks/experiments/REGISTRY.json`
- Brain anchors: D6 receipt `14fb747f8e8c4eda` · Q8_0 pivot `bc40547d0f764775` · D4 ratio research `747af810` · MTP ruling `599622b341ee431e`
- Precedents: Cycle-25 Gate-C APEX (bars) · P8 EE.3 long_form grid (banks/extractors/grader) · D-P8-021/022 (llama-server pivot) · AMENDMENT_3/4 (caps + GPQA recovery)

## 16. Seal

Seal state: **SEALED 2026-07-04T01:35Z** — §2 hypotheses + verdict rule operator-ratified pre-data (verbatim "Prereg looks good. Approved."). The canonical-prefix self-hash (LF-norm bytes `0..^## 16. Seal`, rstrip+`\n`, SHA-256, per `seal_prereg.py`) was computed via `Skill(/sidecar --hash-first)` at this moment; the hex lives in the `.sha256` sidecar beside this file + the experiment marker `~/.claude/hooks/experiments/osint-suite-dev-2026-07-04-4bd1.experiment.json` — NEVER inlined here (Gate-D `815e0a35` anti-propagation). This section sits below the hash boundary; frozen content is everything above the `## 16. Seal` line. Freeze-gate armed on `PRE_REGISTRATION.md` · `run_d7_retention.py` · `run_d7_apex_fidelity.py`. Post-data adjudications land in `## 99. Corrigendum` below this section.
```

***

*Evidence & seals · Pruned-122B Quality Gate (D7) · experiment `osint-suite-dev-2026-07-04-4bd1` · pre-registration reproduced verbatim from the sealed record; self-hash quoted from its sidecar · CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*
