> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/evidence-and-seals/apex-imatrix-calibration-p5.md).

# APEX Imatrix Calibration (P5)

*Evidence for* [***15 · Sovereign In-Distribution Imatrix Calibration***](/osintelligence/part-iv-the-evidence-what-worked/15-sovereign-imatrix-calibration.md)*. This is the validation record behind the chapter's headline: a sovereign-corpus imatrix calibration of the 35B Architect that came out 21% smaller, \~11% faster, and statistically indistinguishable from the Q6\_K baseline across perplexity, KL divergence, and an audit-pass battery. Its mean KL divergence of 0.0103 is the sixteen-fold-tighter result the chapter reports.*

*How this evidence is classed, and how to read a seal generally, is on* [***Verifying a Seal***](/osintelligence/evidence-and-seals/verifying-a-seal.md)*.*

**Phase run** 2026-04-11 → 2026-04-13 · **verdict all gates GREEN** (V1–V4) · artifact promoted to the default `[architect]` preset.

## Evidence class (read this first)

This is **report-class evidence, not a pre-registration seal.** P5 ran 2026-04-11 through 04-13, which is *before* the hash-provenance correction of 2026-04-18 that created this program's canonical-prefix sealing discipline (that correction, and why it matters, is documented on the [Self-Distillation Pilot (P7)](/osintelligence/evidence-and-seals/self-distillation-pilot-p7.md) page). P5 therefore has no canonical-prefix self-hash to quote, and this page does not manufacture one. What it has instead is a frozen final report with a fixed four-gate acceptance table, a live tracking log with full line-referenced provenance, and a machine-readable GGUF verification report, the evidence that existed under the discipline of the time. It is reproduced here verbatim and labelled honestly for what it is.

## Verdict (four validation gates, all green)

The artifact **Qwen3.5-35B-A3B-APEX-I-Balanced.gguf** cleared every pre-set gate with margin and was promoted to the default Architect preset. The limits were fixed in the P5 spec before the validation run; the actuals are read from the tracking log (line ranges quoted per gate in the report below).

| Gate | Metric                 | Limit            | Actual                 | Verdict |
| ---- | ---------------------- | ---------------- | ---------------------- | ------- |
| V1   | sovereign-corpus PPL Δ | ≤ ±1.0%          | **−0.082%**            | PASS    |
| V1   | wikitext-2 PPL Δ       | ≤ ±3.0%          | **−0.107%**            | PASS    |
| V2   | KL mean vs baseline    | ≤ 0.30           | **0.0103**             | PASS    |
| V2   | KL max                 | ≤ 5.0            | **1.52**               | PASS    |
| V2   | same top-1             | informational    | 96.08%                 | —       |
| V3   | SAT audit APPROVE      | ≥ baseline − 2pp | **53.3% (= baseline)** | PASS    |
| V3   | SAT audit REJECT       | ≤ baseline + 2pp | **0% (= baseline)**    | PASS    |
| V4   | VRAM peak              | ≤ 23.6 GB        | **11.2 GB**            | PASS    |
| V4   | throughput ratio       | ≥ 0.90×          | **1.11×**              | PASS    |

The headline: the sovereign-imatrix quant is **smaller (22.8 GB vs 28.9 GB, −21%), faster (1.11×), and quality-indistinguishable** (PPL within a tenth of a percent, KL mean 0.0103, identical SAT-audit distribution). The KL mean is **\~16× below** typical Q6\_K community benchmarks (\~0.20), the chapter's central claim that in-distribution calibration data measurably beats a generic corpus.

**A methodology correction is on the record inside V2**, and it belongs to the same verify-don't-trust ethic as the seals: two early KL attempts using API `top_logprobs` produced inflated divergence through top-k truncation, and the error was caught and corrected to the native full-vocabulary `llama-perplexity --kl-divergence-base` tool (found via llama.cpp Discussion #4110 and community benchmarks) rather than reported as-is. The report names the wasted compute and the lesson; it is reproduced verbatim below.

## Artifact provenance (quoted from the verify report)

The GGUF verification report (`P5_apex_verify_report.json`) machine-checked the BF16 source shards the quant was built from; every field matched (`ok: true` overall):

| Check                           | Expected                            | Actual    | ok |
| ------------------------------- | ----------------------------------- | --------- | -- |
| `general.architecture`          | qwen35moe                           | qwen35moe | ✓  |
| `qwen35moe.block_count`         | 40                                  | 40        | ✓  |
| `qwen35moe.expert_count`        | 256                                 | 256       | ✓  |
| `qwen35moe.expert_used_count`   | 8                                   | 8         | ✓  |
| `general.base_model.0.repo_url` | huggingface.co/Qwen/Qwen3.5-35B-A3B | (match)   | ✓  |

Source shards: `Qwen3.5-35B-A3B-BF16-00001-of-00002.gguf` (49,826,698,560 B) + `-00002-of-00002.gguf` (19,549,939,264 B), 733 tensors, quantized-by Unsloth. Imatrix source: `imatrix_apex_v6.dat`, a 511-breadcrumb sovereign corpus (Hindsight operator breadcrumbs + Architect dog-food synthesis). Full per-gate provenance lives in the 1,013-line `P5_apex_tracking.md` (line ranges cited in the report).

## The final report, verbatim

Reproduced exactly as sealed at phase completion (`P5_apex_report.md`).

```md
# P5 APEX Architect — Final Report

**Spec:** [SPEC_P5_apex_architect.md](../SPEC_P5_apex_architect.md)
**Parent roadmap:** [SPEC_P0_master_roadmap.md](../SPEC_P0_master_roadmap.md) (velvet-swimming-parrot)
**Live tracking log:** [P5_apex_tracking.md](P5_apex_tracking.md) (1013 lines, full provenance)
**Phase status:** ✅ **COMPLETE — ALL GATES GREEN**
**Kickoff:** 2026-04-11
**Completion:** 2026-04-13
**Operator decision at kickoff:** Option 1A (Unsloth BF16 source) + Option 2b (deferred sovereign imatrix corpus)

---

## Executive Summary

P5 delivers **Qwen3.5-35B-A3B-APEX-I-Balanced.gguf** — a sovereign-imatrix-calibrated quantization of the architect model that is **smaller, faster, and qualitatively indistinguishable** from the Q6_K baseline. All four validation gates (V1–V4) passed with substantial margin. The artifact is ready to be promoted to the default `[architect]` preset.

| Gate | Metric | Limit | Actual | Verdict |
|------|--------|-------|--------|---------|
| V1 (sovereign PPL) | Δ ≤ ±1.0% | ±1.0% | **−0.082%** | ✅ PASS |
| V1 (wikitext-2 PPL) | Δ ≤ ±3.0% | ±3.0% | **−0.107%** | ✅ PASS |
| V2 (KL mean) | ≤ 0.30 | 0.30 | **0.0103** | ✅ PASS |
| V2 (KL max) | ≤ 5.0 | 5.0 | **1.52** | ✅ PASS |
| V2 (same top-1) | informational | — | **96.08%** | — |
| V3 (SAT APPROVE) | ≥ baseline − 2pp | 51.3% | **53.3%** (= baseline) | ✅ PASS |
| V3 (SAT REJECT) | ≤ baseline + 2pp | 2.0% | **0%** (= baseline) | ✅ PASS |
| V4 (VRAM peak) | ≤ 23.6 GB | 23.6 GB | **11.2 GB** | ✅ PASS |
| V4 (TPS ratio) | ≥ 0.90× | 0.90× | **1.11×** | ✅ PASS |

**Overall: GREEN.**

---

## Artifact

| Property | Value |
|----------|-------|
| Path | `INFRA/MODELS/Qwen3.5-35B-A3B/Qwen3.5-35B-A3B-APEX-I-Balanced.gguf` |
| Size | 22.8 GB (baseline Q6_K: 28.9 GB — **−21%**, **−6.1 GB**) |
| Architecture | `qwen35moe`, 40 layers, 256 experts, 8 active |
| Imatrix source | `imatrix_apex_v6.dat` — sovereign corpus: 511 Hindsight operator breadcrumbs + Architect dog-food synthesis |
| Quant config | APEX I-Balanced: layer-wise `--tensor-type` schedule (attention/MoE gates at higher precision; expert FFN at IQ4_XS) |
| Provenance | Quantized from `unsloth/Qwen3.5-35B-A3B-GGUF` BF16 shards (49.8 GB + 19.5 GB) |

---

## Gate-by-Gate Results

### V1 — Perplexity Comparison (C.6)

Full details in tracking log lines 871–906.

- **Sovereign corpus (10.0 MB, 500 chunks):** APEX 13.0821 vs Baseline 13.0928 → **Δ −0.082%** (APEX fractionally better, statistically indistinguishable within ±stderr)
- **Wikitext-2 (1.3 MB, 580 chunks):** APEX 6.7244 vs Baseline 6.7316 → **Δ −0.107%**
- **Throughput bonus:** APEX ~30% faster on perplexity passes (128–129 vs 98–100 tok/s) due to smaller weights

### V2 — KL Divergence (full-vocab)

Full details in tracking log lines 947–995.

- **Methodology:** `llama-perplexity --save-all-logits` baseline dump → `--kl-divergence --kl-divergence-base` APEX comparison (native llama.cpp, full 151k-token vocabulary — no top-k truncation)
- **Mean KLD 0.0103 ± 0.0003** — **16× better** than typical Q6_K community benchmarks (~0.20)
- **Same top-1 96.08%** — models agree on the most-likely token 96 times out of 100
- **Methodology lesson:** Two early attempts using API `top_logprobs` produced inflated KL due to top-k truncation + autoregressive divergence. Web research (llama.cpp Discussion #4110, localbench, HuggingFace KLD-guided quantization) revealed the native flag exists. Principle captured in `memory/feedback_research_first.md`: *stand on the shoulders of giants — always lead with research.*

### V3 — SAT Audit Pass Rate (C.7)

Full details in tracking log lines 748–868.

- **15 prompts** through APEX and Q6_K baseline, Orchestrator judges each code response APPROVE/FLAG/REJECT
- APEX APPROVE 8/15 (53.3%) = Baseline APPROVE 8/15 (53.3%) — **identical distribution**
- Zero REJECTs on either side
- **Quality preserved exactly**; APEX is 6–8% faster at steady state

### V4 — VRAM / Throughput (C.8)

Full details in tracking log lines 909–944.

- **Peak VRAM:** APEX 11.2 GB vs Baseline 11.5 GB — **−324 MB**, 52% of the 23.6 GB ceiling unused
- **Throughput:** APEX 26.3 vs Baseline 23.8 tok/s median — **1.11× ratio**, 11% faster
- Both fit comfortably on the RTX 5070 12 GB with KV cache headroom

---

## Operational Impact

1. **Smaller & faster at equal quality.** APEX delivers ~11% throughput gain, ~324 MB VRAM savings, 6.1 GB disk savings, with zero measurable quality regression across PPL, KLD, and SAT audit.
2. **Sovereign imatrix validated.** The 511-breadcrumb Hindsight + dog-food calibration corpus produced KLD **16× below** standard-corpus Q6_K community benchmarks. In-distribution calibration data matters.
3. **Router Mode headroom.** The 22.8 GB footprint leaves more VRAM for KV cache under long-context workloads (research-brain sessions, architect SAT audits).
4. **Training-data flywheel confirmed.** Operator breadcrumbs harvested from Hindsight banks produced measurably better calibration than generic corpora. Every session is fuel.

---

## Promotion Plan (Post-P5)

1. Update `INFRA/llama.cpp/models.ini` — promote `[architect-apex]` preset to the default `[architect]` alias; keep Q6_K registered as `[architect-legacy]` for A/B fallback.
2. Verify Router Mode LRU handles the swap cleanly (SHUSH=0 cold-load test).
3. Archive `Qwen3.5-35B-A3B-Q6_K.gguf` to `INFRA/MODELS/ARCHIVE_GEN1/` once a week of production use confirms stability.

---

## Handoff to P6

- **P6 is CTI NER LoRA training** on Qwen3.5-0.8B-Base (762 labeled NER pairs staged by P3).
- P5 unblocks P6 because the Architect SAT loop now runs on the faster, smaller APEX — SAT audit throughput directly affects training iteration velocity.
- **venv-training** is confirmed operational (Python 3.10.11, torch 2.10.0+cu128, unsloth 2026.4.4).

---

## Lessons Captured

- `feedback_research_first.md` — stand on the shoulders of giants; web-search established methodology before inventing from scratch. (Forged during V2 KL methodology thrash: two flawed API top-k implementations wasted ~20 min of compute before native `llama-perplexity --kl-divergence-base` was discovered.)
- `feedback_venv_training.md` — venv-training directory location + usage (hidden by deny-list to tooling; operator-visible).

---

## Artifacts Index

- Quant artifact: `INFRA/MODELS/Qwen3.5-35B-A3B/Qwen3.5-35B-A3B-APEX-I-Balanced.gguf`
- Imatrix: `INFRA/MODELS/Qwen3.5-35B-A3B/imatrix_apex_v6.dat`
- Tensor schedule: `INFRA/MODELS/Qwen3.5-35B-A3B/apex_i_balanced_tensor_types.txt`
- Validation harness: `TEST_SUITE/spec_p5_apex_validation.py`
- Live log: `RESEARCH/SPECS/RESULTS/P5_apex_tracking.md`
- V3 baseline fixture: `RESEARCH/SPECS/RESULTS/P5_v3_baseline_results.json`
- Verify report: `RESEARCH/SPECS/RESULTS/P5_apex_verify_report.json`
```

***

*Evidence & seals · APEX Imatrix Calibration (P5) · report-class validation record reproduced verbatim; pre-dates the canonical-prefix sealing discipline (P7, 2026-04-18), so classed honestly as a validated report rather than a pre-registration seal · CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*
