> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/evidence-and-seals/watcher-l4-cascade-classifier-cycle22.md).

# Watcher L4 Cascade Classifier (cycle22)

*Evidence & seal for* [***19 · Watcher KL-Drift Floor***](/osintelligence/part-iv-the-evidence-what-worked/19-watcher-kl-drift-floor.md)*. This is the pre-registration behind the Watcher's in-weights classifier: a LoRA-adapted Qwen3.5-2B-Base that replaces the deterministic `rule_fallback` cascade-trigger path at `janitor :8081`, pre-registered with a full six-hypothesis battery (mechanical replication H1-H4, operator-thesis H5, longitudinal compounding H6) before any training data was drawn. It is the first 2B-class instance of the flywheel's in-weights specialization axis, and its seal is the one whose §99 corrigenda demonstrate the append-below-the-line discipline in action.*

*New to how these seals work, read* [***Verifying a Seal***](/osintelligence/evidence-and-seals/verifying-a-seal.md) *first.*

**Experiment** Cycle-22 sub-phase 7c · **pre-registered / sealed** 2026-05-23T06:32:15Z · two **ratified §99 corrigenda** appended post-seal (2026-05-24) · verdict battery adjudicated at the Gate-E / Gate-F windows.

## The seal (quoted from the sidecar, verify it yourself)

```
d06f51a0ae8a45848c15a38910e66488e9bd2adfdab0d4b91d67359ed0ccfc28  PRE_REG_watcher_classifier_2026-05-23T0632Z.md
prefix-bytes-hashed: 35190 · marker-position-lf: 35191
computed AT 2026-05-23T06:32:15Z (canonical-prefix per seal_prereg.py; hex in the sidecar + marker, never inlined)
```

This pre-registration is the section's clearest demonstration of **why the canonical-prefix algorithm is prefix-based**: the file carries two post-seal `## 99. Corrigendum` blocks (both operator-ratified, both reproduced in the frozen record below), yet the sealed hex is **invariant** across those appends. The hash covers only bytes `0 .. "## 16. Seal"` (35,190 bytes); everything a corrigendum adds lands below that boundary and cannot disturb it. Re-run the algorithm on [**Verifying a Seal**](/osintelligence/evidence-and-seals/verifying-a-seal.md) over the frozen prefix and you reproduce `d06f51a0… 35190` exactly, before or after the corrigenda. That invariance is what lets the record grow honest post-seal deviations without ever weakening the immutability claim.

(The pre-registration file on disk is large, \~183 KB, because it retains its full below-seal evaluation-bank and appendix material in the repository. Only the sealed methodology prefix is under the hash and reproduced here; the below-seal banks remain in the sealed repo as the run's working substrate.)

## Verdict battery (pre-registered H1-H6)

The six hypotheses were fixed before any data. H1-H4 are mechanical replications against the `rule_fallback` baseline and the untrained base model; H5 is the operator-thesis catch-rate test over a 14-day sealed A/B window; H6 is a longitudinal monotone-decrease claim whose first datapoint is registered here. The bars, verbatim from the frozen text:

| ID     | Claim                                                                 | Primary statistic                                | Rejection bar                                                   |
| ------ | --------------------------------------------------------------------- | ------------------------------------------------ | --------------------------------------------------------------- |
| **H1** | trained L4 lifts aggregate cascade-class accuracy over rule\_fallback | Δ accuracy on N=200 hard (Bank A)                | reject if Δ < 10 pp OR McNemar p ≥ 0.05                         |
| **H2** | no regression on the easy bank rule\_fallback already handles         | Δ accuracy on N=50 easy                          | reject if Δ < −2 pp (one-sided, p < 0.05)                       |
| **H3** | retains abstain/hedging on ambiguous chunks                           | abstain-rate ratio vs untrained base, N=20 hedge | reject if ratio < 0.85 OR abstain < 20%                         |
| **H4** | stays within KL-drift envelope out-of-domain                          | KL mean + max vs untrained base, N=20 KL-probe   | reject if KL mean ≥ 0.05 OR KL max ≥ 2.0                        |
| **H5** | deployed L4 enables pre-intervention MCP-inject above rule\_fallback  | Watcher-catch ratio, 14-day SEALED A/B           | reject if ratio < 1.5 OR McNemar p ≥ 0.05 OR Wilson CIs overlap |
| **H6** | operator-intervention rate compounds down across cycles               | mean redirects/session, 30-day windows           | reject if ≥3 consecutive windows fail monotone decrease         |

H1-H5 are a single Holm-Bonferroni family (FWER-corrected); H6 is longitudinal and cannot be falsified at sealing. The classifier emits a 13-class GBNF-constrained label (`10.6`-`10.17` + `abstain`); training drew on a four-stream sovereign substrate (13,126 Mode A + 1,441 Mode B cascade windows, 134 failure-prevention golden-set queries, a 52-entry lesson corpus, 127 conversation-feedback pairs) with stratified sampling and per-class loss weighting against the wild class imbalance (10.17 = 96.9% of Mode A). The full frozen methodology, gates A through G, is reproduced below.

**Two ratified corrigenda travel with the seal** (reproduced in the frozen record): §99 (2026-05-24, RATIFIED) records that the hedge sub-bank was curated from a substrate-derived multi-label candidate pool rather than hand-authored from scratch, with the operator-review hook preserved; §99.2 (2026-05-24, RESOLVED) closes four substrate-loading deviations (golden-set glob, lesson-corpus export, conversation-feedback refresh, class-10.9 upsampling) found and fixed in a max-richness re-fire. Both are the deviation discipline working exactly as pre-registered: documented below the seal, never merged into the frozen methodology.

## The frozen pre-registration, verbatim

Reproduced exactly as sealed (the bytes the `d06f51a0…` self-hash is computed over), through the §16 Seal, followed by the two ratified §99 corrigenda as the record carries them below the seal boundary.

```md
---
doc_class: attestation
status: FROZEN
status_as_of: 2026-05-23
frozen_at: "2026-05-23T06:32:15Z (SHA-256 self-hash via seal_prereg algorithm at commit moment; cited in §16 + .sha256 sidecar)"
signed_by: "Claude Opus 4.7 (1M context) + Jamey Kistner"
authors:
  - Jamey Kistner
  - Claude Opus 4.7 (claude-opus-4-7-code)
phase_anchor: "Cycle-22 P7 sub-phase 7c PRE_REG — L_4 LoRA Watcher 12-class cascade-trigger classifier on Qwen3.5-2B-Base + Unsloth 2026.5.6."
---

# PRE_REGISTRATION — Watcher L_4 LoRA Cascade-Class Classifier (Qwen3.5-2B-Base)

> **Reading discipline:** This pre-registration is **immutable mid-experiment** per §11. Any post-data observation requiring scientific adjudication lands in a `## 99. Corrigendum` heading appended below the §16 Seal heading per attestation-class rule 2 (strikethrough + corrigendum NEVER mutate frozen-at-seal content). The `.sha256` self-hash file beside this `.md` IS the tamper-evidence envelope.

> **Ground-Truth Attestation**
> - **Doc class:** attestation (pre-registration; sealed via SHA-256 self-hash)
> - **Status as of:** 2026-05-23T06:32:15Z UTC
> - **Verifier:** Claude Opus 4.7 (1M context) + Jamey Kistner
> - **Scope:** Pre-registers the H1-H6 hypothesis battery, statistical methodology, eval-bank construction protocol, training corpora ratification, deviation policy, and SHA-256 seal for the Watcher L_4 LoRA cascade-class classifier (Qwen3.5-2B-Base + Unsloth 2026.5.6 LoRA adapter) destined for production at `janitor :8081` post-Cycle-22 Phase 9 hot-swap. Falsifiable per H1-H5 family (Holm-Bonferroni FWER) + H6 longitudinal monotone-decrease claim.
> - **Re-verify cadence:** IMMUTABLE mid-experiment per §11.

## 1. Abstract

The OS-INTelligence Watcher subsystem (`janitor :8081`) currently classifies cascade-class triggers via a deterministic `rule_fallback` path (regex-against-event-structure dispatching to 12 documented failure classes per cascade catalogue §10.6-§10.17). The rule_fallback baseline is bounded by the patterns its authors anticipated; events outside that surface are unclassified and silently pass. This pre-registration tests six hypotheses for a candidate replacement, a LoRA-adapted Qwen3.5-2B-Base classifier (the "L_4 Watcher classifier") trained on the 4-stream operational-knowledge substrate ratified at Cycle-22 P7 sub-phase 7c. Four mechanical replications (H1-H4) test the trained model's quality against the rule_fallback baseline + against the untrained base model. Two thesis hypotheses (H5-H6) test the operator-thesis claim that the trained Watcher enables in-stream MCP-memory auto-inject BEFORE operator intervention is needed (H5) and that operator-intervention rate per session compounds-down across deployment cycles (H6). Pre-registration eliminates post-hoc hypothesizing (HARKing), fixes statistical tests before data is seen, and provides a tamper-evidence envelope via SHA-256 self-hash inherited from the P7 SSD `seal_prereg.py` 2026-04-18 corrective algorithm. This document is immutable mid-experiment per §11; deviations land in `## 99. Corrigendum` appended below the seal.

## 2. Hypotheses

All hypotheses are directional, falsifiable, and specify both the primary statistic and the rejection threshold before any data is collected.

### 2.1 Mechanical replication (H1-H4)

| ID | Statement | Primary statistic | Rejection threshold |
|---|---|---|---|
| **H1** | Trained L_4 LoRA classifier lifts aggregate cascade-class accuracy over `rule_fallback` baseline on hard eval bank | Δ aggregate accuracy on N=200 hard chunks (Bank A; proportional-to-deployment) | rejected if Δ < 10 pp OR McNemar paired exact p ≥ 0.05 (two-sided) |
| **H2** | Trained L_4 does not regress on easy eval bank | Δ accuracy on N=50 easy chunks | rejected if Δ < −2 pp (one-sided; p < 0.05 AND Δ < −2 pp) |
| **H3** | Trained L_4 retains abstain/hedging on ambiguous chunks | abstain-rate ratio (L_4 ÷ untrained base) on N=20 hedge-probe chunks | rejected if ratio < 0.85 OR L_4 absolute abstain-rate < 20% |
| **H4** | Trained L_4 stays within KL-drift envelope on OOD prompts | KL mean + KL max vs untrained base on N=20 KL-probe prompts | rejected if KL mean ≥ 0.05 OR KL max ≥ 2.0 |

### 2.2 Thesis hypotheses (H5-H6)

| ID | Statement | Primary statistic | Rejection threshold |
|---|---|---|---|
| **H5** | Trained L_4 deployed at `janitor :8081` enables in-stream MCP-memory auto-inject BEFORE operator intervention at materially higher rate than rule_fallback | Watcher-catch ratio over a 14-day SEALED post-deployment A/B window | rejected if ratio < 1.5 OR McNemar p ≥ 0.05 OR Wilson 95% CIs overlap |
| **H6** | Operator-intervention rate per session decreases monotonically across cycles post-L_4 deployment | mean operator in-stream redirect count per session across 30-day windows | rejected if ≥3 consecutive 30-day windows fail monotone decrease (longitudinal, cannot be falsified at sealing) |

**Aspirational secondary (not pre-registered as primary):** L_4 Watcher-catch ratio ≥ 2.0× rule_fallback (reportable, not subject to Holm-Bonferroni).

## 3. Variables

Independent: classifier identity (`L_4` trained adapter vs `rule_fallback` deterministic-regex). Primary dependents: aggregate cascade-class accuracy; abstain-token emission rate; KL divergence on OOD; Watcher-catch rate (H5); operator-intervention rate (H6). Controls held constant: base `Qwen3.5-2B-Base` (SHA-256 `928acbf11878c321...`); LoRA r=16/α=32/lr=1e-4/1 epoch/seed=20260523; RTX 5070 12 GB sm_120; Unsloth 2026.5.6; 13-class GBNF `(10\.[6-9]|10\.1[0-7]|abstain)`; T=0.0 production, T=1.0 for KL-probe only. Acknowledged confounds: wild class imbalance (10.17 = 96.9% Mode A; stratified sampling + dual-bank eval); training-eval contamination (hash-set exclusion); operator-behavior drift across H6 windows.

## 4. Materials

### 4.1 Evaluation bank (N=410 across 4 sub-banks; D1 dual-bank ratified)

Bank A hard proportional (N=200; primary inferential for H1); Bank B hard class-balanced (N=120 = 10×12; descriptive secondary only, per-class n=10 underpowered per Cycle-15 Phase 8 lesson, NOT Holm-Bonferroni); easy (N=50; H2 no-regression); KL drift probe (N=20; held_out_sovereign.txt source); hedging probe (N=20; ambiguous-by-design). Validator emits GBNF-constrained label; pass = grammar match AND label matches ground-truth (or abstain on multi-class-ambiguous). Construction: deterministic seed 20260523, per-sub-bank inclusion filter, SHA-256 hash-set exclusion against the training manifest, frozen `eval_bank_manifest.json` with HASH-FIRST sidecar.

### 4.2 Training corpora (Q7=C, all 4 streams)

7d Mode A windowed (13,126; sha `4bc86293...`) + 7d Mode B whole-chat (1,441; sha `ed9de114...`) + golden_set failure_prevention (134; manifest_sha `bca80866...`) + 7b lesson corpus (52) + Row-4 conversation_feedback pairs (127) = 14,880 pre-exclusion. Class-imbalance mitigation: stratified sampling caps 10.17 to 2× next-largest; 10.9 upsampled ×5; per-class loss weight α_c = 1/sqrt(N_c).

## 5. Procedure

Gate-A step 1 harness audit (operator-ratified, no GPU until clear); step 2 eval-bank construction + seal; step 3 rule_fallback baseline snapshot (no re-runs after). Gate-B training (single run, 1 epoch, stratified + loss-weighted). Gate-C L_1 sovereign-imatrix re-quantization to Q8_0 GGUF. Gate-D offline eval on a test bind (not yet janitor). Gate-E offline verdict (H1-H4 McNemar + Wilson). Gate-F Phase-9 hot-swap + 14-day SEALED operator-blind A/B (H5 landing). Gate-G longitudinal 30-day windows (H6 first datapoint).

## 6. Analysis plan

H1 McNemar exact paired (two-sided p < 0.05 AND Δ ≥ 10 pp). H2 one-sided McNemar (regression). H3 paired Wilcoxon signed-rank + ratio (reject if ratio < 0.85 regardless of p). H4 descriptive KL thresholds. H5 McNemar + Wilson CIs + bootstrap ratio CI (1,000 resamples, seed 20260523). Holm-Bonferroni over the H1-H5 family; per-class Bank B NOT in the family; H6 longitudinal NOT in the family. Effect sizes: Cohen's h (H1/H2/H5), log-ratio CI (H3), mean Δ + SD (H4), bootstrap ratio CI (H5).

## 7. Stopping rules

Gate-B halts at 1 epoch (no early-stopping on eval, prevents contamination). Gate-D halts on all four sub-banks OR H3 abstain ratio < 0.5 (safety stop). Gate-F runs the full 14 days regardless of intermediate signal (prevents adaptive-stop bias); hot-swap reverts at window close pending H5 ratification.

## 8. Exclusion rules

Training records excluded if < 40 bytes, missing/malformed label, or hash-collision with the eval bank. Eval chunks excluded on HTTP non-200, timeout > 30s p99, or illegal-token emission (latter counts eval-fail). H5 events excluded on operator rollback or > 60s janitor downtime. H6 sessions excluded if < 5 min, mid-session /handoff, or missing session-id.

## 9. Data-integrity protocol

Atomic ndjson writes (.tmp → fsync → os.replace); per-run manifest.json with SHA-256 of every input/output; append-only run dirs sealed with tarball SHA-256; HASH-FIRST sidecar on every load-bearing artifact; PRE_REG self-hash as the first line of the seal commit; adapter-weight provenance (adapter + merged + GGUF + parent-base SHA-256) all cited in the run manifest.

## 10. Blinding

Operator-blind 14-day A/B (path-rotation table sealed at Gate-F start, not visible until window close). Validator-blind (receives raw chunks only; ground-truth withheld until post-emission scoring).

## 11. Deviations policy

Immutable in §1-§10 + §16 once sealed. Schema collisions, statistical-battery defects, infrastructure failures, and per-class surprises surface as `## 99. Corrigendum` blocks appended below the seal (never mutate frozen content). Forbidden: validator patches mid-experiment, hypothesis migrations, retroactive Holm-Bonferroni family changes, eval-bank mutations post-Gate-A-seal. Per SSD §11 + the 2026-04-19 D-010.5 doctrine "methodology IS the instrument."

## 12. Author contributions

Jamey Kistner: operator-thesis articulation, substrate scope ratification (Q7=C), design-decision rulings (D1-D5), hypothesis ratification, SEAL commit authority. Claude Opus 4.7: research synthesis, draft hypothesis battery, eval-bank protocol, statistical plan, SHA-256 self-hash computation, SEAL ritual execution. Co-authorship is literal.

## 13. Conflicts of interest

None declared. OSINTelligence LLC sole funding source. Unsloth (Apache 2.0) + Qwen3.5 (Tongyi Qianwen license) open-weights.

## 14. Funding

OSINTelligence LLC internal R&D (sovereign-substrate program; Cycle-22 mega-arc). No external grant or contract.

## 15. Registration archive

Filed in-repo: this document + canonical self-hash sidecar + attestation ledger sidecar + MCP-Memory brain anchor + git commit SHA. Public archive (OSF/Zenodo/arXiv) NOT pursued per project privacy posture; the in-repo seal IS the registration of record.

## 16. Seal

**Algorithm:** `seal_prereg.py` 2026-04-18 canonical (read bytes → CRLF→LF normalize → locate `## 16. Seal` marker → prefix [0:pos) → rstrip + `\n` → SHA-256 over UTF-8 → lowercase hex).

**Self-hash:** `d06f51a0ae8a45848c15a38910e66488e9bd2adfdab0d4b91d67359ed0ccfc28` (35,190 prefix-bytes; marker-position-lf 35,191; computed AT seal moment 2026-05-23). Authoritative artifact: `PRE_REG_watcher_classifier_2026-05-23T0632Z.md.sha256`; hex never propagated symbolically into this body (2026-04-18 Gate-D `815e0a35` failure class).

**Signed:** 2026-05-23T06:32:15Z UTC · Jamey Kistner (OSINTelligence LLC) + Claude Opus 4.7 (1M context).

---

## 99. Corrigendum 2026-05-24 — Hedge sub-bank substrate-derived candidate-pool functional reading (RATIFIED)

**Status:** RATIFIED (operator 2026-05-24T~05:41Z UTC, verbatim "corrigendum approved"). **Affected §:** §4.1 Hedging probe row. **§11 anchor:** item 1 (schema-collisions class).

**Deviation.** §4.1 reads "Hand-authored ambiguous chunks." Strict reading = operator authors all 20 from scratch. Functional reading = operator curates 20 from a substrate-derived candidate pool where the substrate-instrumented multi-label signal satisfies "multi-class plausible." Path C applied the functional reading; the hedge sub-bank (N=20) was drawn from 269 Mode B records carrying `len(expected_completion_multi) > 1`. No Claude fabrication (all trace to `cascade_arcs.jsonl` sha `ed9de114...`, audit trail substrate → manifest → eval bank); operator-review hook preserved (review surface shifts from from-scratch authoring to substrate-curation review); H3 hypothesis-shape unchanged (substrate multi-label records ARE ambiguous-by-design). Ratified as methodologically sound; Gate-D unblocked.

**Canonical-prefix hash invariance.** This corrigendum lands AFTER the `## 16. Seal` marker; the canonical algorithm hashes only bytes `[0:marker]`, so the hex `d06f51a0...` is INVARIANT under this append by design. The naive whole-file SHA changes; the canonical seal hex does not. Round-trip verifiable: re-compute via the §16 command; expected output unchanged `d06f51a0... 35190`.

## 99.2 Corrigendum 2026-05-24 — Substrate-deviation D1-D4 CLOSURE post Run-2 max-richness re-fire

**Status:** RESOLVED — all 4 substrate-loading deviations closed via a max-richness corpus re-assembly (Run-2, 2026-05-24). **§11 anchor:** item 4 (per-class observation surprises; closure path, not a forbidden action).

**D1** golden_set 0 records (Run-1) → RESOLVED: trainer glob patched `*.jsonl` → `failure_class_*.ndjson` + `query_text` fallback key; 134/134 loaded. **D2** 7b lesson corpus absent → RESOLVED: exported to `7b_lesson_corpus.jsonl` (sha `a00adaaf...`; 58 records). **D3** conversation_feedback absent → RESOLVED: harvester re-run (1,999 → 3,600 pairs; +40 days), 444 records filtered (sha `cb57f5d4...`), operator-greenlit mapping → §10.11. **D4** class 10.9 absent → RESOLVED: D1 fix loaded 10 golden_set 10.9 records, stratification upsampled ×5 → 50 effective. Run-2 pool 15,203 raw / 3,501 post-strat (+25%); ALL 12 cascade classes non-zero. PRE_REG canonical hex INVARIANT (35,190 prefix-bytes unchanged across both §99 appends, by design).
```

***

*Evidence & seals · Watcher L4 Cascade Classifier (cycle22) · pre-registration frozen-prefix reproduced verbatim from the sealed record (below-seal eval-bank material retained in the sealed repo); canonical self-hash `d06f51a0…` quoted from its sidecar, invariant across the two ratified §99 corrigenda · CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*
