> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/part-i-the-architecture/2-the-external-sentinel/references-and-provenance.md).

# References & provenance

*Companion to Chapter 2 · The External Sentinel · v1.0.0. The July 2026 System Update, the reference list, and provenance.*

### System Update: July 2026 (appended; the sealed body above is unmodified)

At-read July 2026: **the hardware Sentinel remains unbuilt and unprocured; the horizon discipline is holding.** What has landed is a software rehearsal of this paper's smallest check class: a **seal-integrity daemon**, staged on disk pending the operator's sign-off to wire into the boot sequence, which re-verifies every sealed research artifact against its SHA-256 sidecar on a 60-second cycle and raises a non-destructive alarm on mismatch. It is deliberately alarm-only and deliberately not called a Governor: it executes on the same substrate as the system it watches, and this paper's §3 is precisely about why that substrate-sharing disqualifies it from the FC-3 role. Its value is operational (drift detection at the artifact layer) and rehearsal-grade (the V1 hash-comparison discipline, exercised in software before it is committed to silicon).

Meanwhile the operator-executed Governor function has deepened its tooling: frozen-path experiment markers, amendment co-seals, canonical-prefix re-hash verification, and post-run integrity envelopes now make the per-cycle verification partially mechanical while keeping the human as the trust root: the migration path this paper specifies, walked in the intended order. Provenance: ECOSYSTEM\_INVENTORY §3.2 · HANDOFF\_2026-07-14 · Cycle-41 seal machinery. Append-only; the sealed body above is unmodified.

### 9. References

Anderson, R. (2008). *Security Engineering: A Guide to Building Dependable Distributed Systems* (2nd ed.). Wiley.

Aumasson, J.-P. (2017). *Serious Cryptography: A Practical Introduction to Modern Encryption.* No Starch Press.

Bai, Y., Kadavath, S., Kundu, S., et al. (2022). "Constitutional AI: Harmlessness from AI Feedback." [arXiv:2212.08073](https://arxiv.org/abs/2212.08073).

Bell, D. E., & LaPadula, L. J. (1973). *Secure Computer Systems: Mathematical Foundations.* MITRE Technical Report 2547.

Bernstein, D. J., Duif, N., Lange, T., Schwabe, P., & Yang, B.-Y. (2012). "High-speed high-security signatures." *Journal of Cryptographic Engineering* 2(2):77–89 (Ed25519).

Biba, K. J. (1977). *Integrity Considerations for Secure Computer Systems.* MITRE Technical Report 3153.

Bommasani, R., Klyman, K., Longpre, S., et al. (2023). "The Foundation Model Transparency Index." Stanford CRFM. [arXiv:2310.12941](https://arxiv.org/abs/2310.12941).

Boneh, D., Lynn, B., & Shacham, H. (2001). "Short signatures from the Weil pairing." *Advances in Cryptology, ASIACRYPT 2001.* Springer LNCS 2248.

Bostrom, N. (2014). *Superintelligence: Paths, Dangers, Strategies.* Oxford University Press.

Bytecode Alliance (2024). *Wasmtime: A standalone runtime for WebAssembly.* <https://wasmtime.dev/>

Carlsmith, J. (2022). "Is Power-Seeking AI an Existential Risk?" Open Philanthropy report; [arXiv:2206.13353](https://arxiv.org/abs/2206.13353).

Christiano, P. F., Leike, J., Brown, T. B., et al. (2017). "Deep Reinforcement Learning from Human Preferences." NIPS 2017.

Christiano, P., Cotra, A., & Xu, M. (2021). "Eliciting Latent Knowledge: How to Tell If Your Eyes Deceive You." Alignment Research Center technical report.

Clark, D. D., & Wilson, D. R. (1987). "A Comparison of Commercial and Military Computer Security Policies." IEEE Symposium on Security and Privacy 1987.

European Union (2024). "Artificial Intelligence Act (Regulation 2024/1689)." [Official Journal of the European Union](https://eur-lex.europa.eu/eli/reg/2024/1689/oj).

European Union (2024). "Right to Repair Directive 2024/1799." [Official Journal of the European Union](https://eur-lex.europa.eu/eli/dir/2024/1799/oj).

Gödel, K. (1931). "Über formal unentscheidbare Sätze der Principia Mathematica und verwandter Systeme I." *Monatshefte für Mathematik und Physik* 38:173–198.

Greenblatt, R., Denison, C., Wright, B., et al. (2024). "Alignment Faking in Large Language Models." [arXiv:2412.14093](https://arxiv.org/abs/2412.14093).

Hadfield-Menell, D., Russell, S., Abbeel, P., & Dragan, A. (2016). "Cooperative Inverse Reinforcement Learning." NIPS 2016. [arXiv:1606.03137](https://arxiv.org/abs/1606.03137).

Hadfield-Menell, D., Dragan, A., Abbeel, P., & Russell, S. (2017). "The Off-Switch Game." IJCAI 2017.

Hendrycks, D., Carlini, N., Schulman, J., & Steinhardt, J. (2022). "Unsolved Problems in ML Safety." [arXiv:2109.13916](https://arxiv.org/abs/2109.13916).

Hubinger, E., Denison, C., Mu, J., et al. (2024). "Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training." [arXiv:2401.05566](https://arxiv.org/abs/2401.05566).

IEC 61508 (2010). *Functional Safety of Electrical/Electronic/Programmable Electronic Safety-related Systems.* International Electrotechnical Commission.

Islam, M. S., Alouani, I., & Khasawneh, K. N. (2024). "Hardware Support for Trustworthy Machine Learning: A Survey." 2024 25th International Symposium on Quality Electronic Design (ISQED). [DOI:10.1109/ISQED60706.2024.10528373](https://doi.org/10.1109/ISQED60706.2024.10528373).

ISO/IEC 11889 (2015). *Information technology, Trusted Platform Module Library (TPM 2.0)* (parts 1–4).

Juvenal (c. AD 100). *Satirae* (Satire VI), lines 347–348. "Sed quis custodiet ipsos custodes?"

Kistner, J. (2026). [*Sixteen Practices for Sovereign Human-AI Collaboration: A Practitioner's Methodology Formalized.*](/osintelligence/part-ii-the-discipline/5-sixteen-practices.md) OSINTelligence LLC.

Kistner, J. (2026). [*The Sovereign Triad: An Architectural Ethics for Self-Improving AI Systems.*](/osintelligence/part-i-the-architecture/1-the-sovereign-triad.md) OSINTelligence LLC.

Klein, G., Elphinstone, K, Heiser, G., Andronick, J., Cock, D., Derrin, P., Elkaduwe, D., Engelhardt, K., Kolanski, R., Norrish, M., Sewell, T., Tuch, H., & Winwood, S. (2009). "seL4: Formal Verification of an OS Kernel." [SOSP 2009](https://doi.org/10.1145/1629575.1629596).

Kocaoğullar, C., Marjanov, T., Petrov, I., Laurie, B., Cutter, A., Kern, C., Hutchings, A., & Beresford, A. R. (2024). "Confidential Computing Transparency." [arXiv:2409.03720](https://arxiv.org/abs/2409.03720).

Lakatos, I. (1970). "Falsification and the Methodology of Scientific Research Programmes." In *Criticism and the Growth of Knowledge*, eds. Lakatos, I. & Musgrave, A. Cambridge University Press.

Lampson, B. W. (1973). "A Note on the Confinement Problem." *Communications of the ACM* 16(10):613–615.

Leroy, X. (2009). "Formal Verification of a Realistic Compiler." *Communications of the ACM* 52(7):107–115.

McDonald, G., & Bar Or, J. (2025). "Whisper Leak: A Side-Channel Attack on Large Language Models." [arXiv:2511.03675](https://arxiv.org/abs/2511.03675).

Montana Senate Bill 212 (2025). *Right to Compute Act.* Signed April 2025.

NIST (2023). [*Artificial Intelligence Risk Management Framework (AI RMF 1.0).*](https://doi.org/10.6028/NIST.AI.100-1) National Institute of Standards and Technology.

National Institute of Standards and Technology (2006). *Recommendation for Obtaining Assurances for Digital Signature Applications.* [NIST Special Publication 800-89](https://doi.org/10.6028/NIST.SP.800-89) (E. Barker).

Ouyang, L., Wu, J., Jiang, X., et al. (2022). "Training Language Models to Follow Instructions with Human Feedback." [NeurIPS 2022](https://arxiv.org/abs/2203.02155) (InstructGPT).

Parno, B., McCune, J. M., & Perrig, A. (2011). *Bootstrapping Trust in Modern Computers.* Springer SpringerBriefs in Computer Science.

Plato (c. 380 BC). *Republic* Book III.

Popper, K. R. (1959). *The Logic of Scientific Discovery* (English translation of Logik der Forschung, 1934). Hutchinson & Co.

Reason, J. (1990). *Human Error.* Cambridge University Press.

Russell, S. (2019). *Human Compatible: Artificial Intelligence and the Problem of Control.* Viking.

Sailer, R., Zhang, X., Jaeger, T., & van Doorn, L. (2004). "Design and Implementation of a TCG-based Integrity Measurement Architecture." USENIX Security 2004.

Saltzer, J. H., & Schroeder, M. D. (1975). "The Protection of Information in Computer Systems." *Proceedings of the IEEE* 63(9):1278–1308.

Schneier, B. (2000). *Secrets and Lies: Digital Security in a Networked World.* Wiley.

Shamir, A. (1979). "How to Share a Secret." *Communications of the ACM* 22(11):612–613.

Summers, A. E. (2013). *Safety Controls, Alarms, and Interlocks as Independent Protection Layers.* SIS-Tech Solutions.

Trusted Computing Group (2014). *TPM 2.0 Library Specification* (parts 1–4). <https://trustedcomputinggroup.org/>

W3C (2019). *WebAssembly Core Specification.* W3C Recommendation. <https://www.w3.org/TR/wasm-core-1/>

Westfall, P. H., & Young, S. S. (1993). *Resampling-Based Multiple Testing: Examples and Methods for p-Value Adjustment.* Wiley.

### Acknowledgments

The design formalized here grew out of running a self-improving system on the first author's own hardware and confronting, in practice, the question this paper is about: how to keep such a system from evolving past the operator's ability to control it. The operator-as-last-line-defense discipline, a human deciding by hand whether each self-modification may proceed, is the working configuration that produced the paper, and the Sentinel is the attempt to give that configuration a substrate that does not depend on the human. The device was discovered as a need in practice before it was specified as a design in theory.

Claude (Anthropic) was used as a research tool in the preparation of this chapter, assisting with drafting, structuring and successive revision. The specific model versions were not recorded at the time of authorship and so are not named here. No AI system is listed as an author or credited as a contributor, in line with COPE and ICMJE guidance: an AI system cannot take responsibility for the work, cannot assert competing interests, and cannot enter a licence agreement. The author verified every claim in this chapter against the sealed artifacts and is solely accountable for it.

**Citation (preferred):** Kistner, J. (2026). *The External Sentinel: A Hardware Substrate for the External Governor of a Self-Improving AI System*, version 1.0.0. OSINTelligence LLC.

**Companion papers (this series):** [*The Sovereign Triad*](/osintelligence/part-i-the-architecture/1-the-sovereign-triad.md) (Chapter 1; specifies the External Governor as its third component, FC-3, the function this device is built to perform) · [*The Drift Taxonomy*](/osintelligence/part-iii-the-evidence-what-broke/9-the-drift-taxonomy.md) (Chapter 9; the failure classes an out-of-reach control is needed to catch) · [*Sovereign Domain Pruning*](/osintelligence/part-iv-the-evidence-what-worked/16-sovereign-domain-pruning.md) (Chapter 16; the sealed-invariant discipline the Sentinel's attestation checks). Cited in-series by title.

**License:** CC BY 4.0 (text). All referenced code artifacts are MIT licensed unless otherwise noted.

**Corresponding author:** Jamey Kistner, <jamey.kistner@osintelligence.io>, OSINTelligence LLC (Columbus, OH).

***

*The Sovereign Stack · The External Sentinel · Chapter 2 · Part I · v1.0.0 · License CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*
