> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/part-ii-the-discipline/8-sovereign-optimization-flywheel/references-and-provenance.md).

# References & provenance

### References

**Systems + architecture:** Senge, P. M. (1990). *The Fifth Discipline.* Doubleday · Sterman, J. D. (2000). *Business Dynamics.* McGraw-Hill · Forrester, J. W. (1961). *Industrial Dynamics.* MIT Press · Hennessy, J. L., & Patterson, D. A. (2019). *Computer Architecture: A Quantitative Approach* (6th ed.). Morgan Kaufmann.

**Scaling laws:** Kaplan, J., et al. (2020). "Scaling Laws for Neural Language Models." [arXiv:2001.08361](https://arxiv.org/abs/2001.08361) · Hoffmann, J., et al. (2022). "Training Compute-Optimal Large Language Models." [arXiv:2203.15556](https://arxiv.org/abs/2203.15556) (Chinchilla).

**MTP + speculative decoding:** Leviathan, Y., et al. (2023). "Fast Inference from Transformers via Speculative Decoding." *ICML 2023* · Chen, C., et al. (2023). "Accelerating LLM Decoding with Speculative Sampling." [arXiv:2302.01318](https://arxiv.org/abs/2302.01318) · Gloeckle, F., et al. (2024). "Better & Faster LLMs via Multi-token Prediction." [arXiv:2404.19737](https://arxiv.org/abs/2404.19737) · Samragh, M., et al. (2025). "Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential." [arXiv:2507.11851](https://arxiv.org/abs/2507.11851) (gated-LoRA MTP adaptation of pretrained models) · Mahajan, D., et al. (2025). "Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries." [arXiv:2510.14751](https://arxiv.org/abs/2510.14751); *ICLR 2026* (the disconfirming bound on quality claims).

**KV compression:** Hooper, C., et al. (2024). "KVQuant." [arXiv:2401.18079](https://arxiv.org/abs/2401.18079) · Wu, H., & Tu, K. (2024). "Layer-Condensed KV Cache for Efficient Inference of Large Language Models." *ACL 2024*; [arXiv:2405.10637](https://arxiv.org/abs/2405.10637) · Liu, Y., et al. (2025). "KV Cache Compression for Inference Efficiency in LLMs: A Review." [arXiv:2508.06297](https://arxiv.org/abs/2508.06297) · TurboQuant (Zandieh, Daliri, Hadian & Mirrokni; *ICLR 2026*; [arXiv:2504.19874](https://arxiv.org/abs/2504.19874)) / PolarQuant (*AISTATS 2026*).

**Alignment + moats:** Bai, Y., et al. (2022). "Constitutional AI." [arXiv:2212.08073](https://arxiv.org/abs/2212.08073) · Greenblatt, R., et al. (2024). "Alignment Faking in Large Language Models." [arXiv:2412.14093](https://arxiv.org/abs/2412.14093) · a16z (2019). "The Empty Promise of Data Moats" · Argyris, C., & Schön, D. A. (1978). *Organizational Learning.* Addison-Wesley.

**In-series companions:** [*Sovereign In-Distribution Imatrix Calibration*](/osintelligence/part-iv-the-evidence-what-worked/15-sovereign-imatrix-calibration.md) (Chapter 15; L₁ sealed) · [*Sovereign Domain Pruning*](/osintelligence/part-iv-the-evidence-what-worked/16-sovereign-domain-pruning.md) (Chapter 16; L₃ sealed) · [*Sovereign CTI-NER*](/osintelligence/part-iv-the-evidence-what-worked/14-sovereign-cti-ner.md) + [*Corpus-Sovereign Self-Distillation*](/osintelligence/part-iv-the-evidence-what-worked/18-corpus-sovereign-self-distillation.md) (Chapter 18; the L₄ chain) · [*The Sovereign Triad*](/osintelligence/part-i-the-architecture/1-the-sovereign-triad.md) (Chapter 1; the Governor whose V2 verifies this loop) · [*Watcher KL-Drift Floor*](/osintelligence/part-iv-the-evidence-what-worked/19-watcher-kl-drift-floor.md) (Chapter 19; the Level-4 axis's classifier under training) · [*Sixteen Practices*](/osintelligence/part-ii-the-discipline/5-sixteen-practices.md) (Chapter 5, §5.34; the methodology-tier statement) · [*Stateless by Construction*](/osintelligence/part-i-the-architecture/4-stateless-by-construction.md) (Chapter 4; the context-discipline complement) · [*Sovereign Big-Model Compression*](/osintelligence/part-iv-the-evidence-what-worked/17-sovereign-big-model-compression.md) (Chapter 17; the L₁ and L₃ axes carried to 122B: a one-shot expert prune plus a sovereign-calibrated imatrix, cleared by a pre-registered served quality gate and now the deployed research brain).

### AI-assistance disclosure

Large language models were used as research tools in the preparation of this chapter: Claude Opus 4.7 (Anthropic); with a cross-platform peer-observation contribution from Claude Opus 4.7 (web instance). The model versions and roles named above keep the provenance of this chapter auditable. No AI system is listed as an author or credited as a contributor, in line with COPE and ICMJE guidance: an AI system cannot take responsibility for the work, cannot assert competing interests, and cannot enter a licence agreement. The author verified every claim in this chapter against the sealed artifacts and is solely accountable for it.

**Citation (preferred):** Kistner, J. (2026). *Sovereign Optimization Flywheel: A Multiplicative-Compound Five-Axis Architecture for Single-Operator AI Deployment on Consumer Hardware*, version 1.0.0. OSINTelligence LLC research whitepaper. Cited in-series by title.

**License:** CC BY 4.0 (text). Code and data artifacts MIT per repository license.

**Corresponding author:** Jamey Kistner, <jamey.kistner@osintelligence.io>, OSINTelligence LLC (Columbus, OH).

***

*The Sovereign Stack · Sovereign Optimization Flywheel · Chapter 8 · Part II · v1.0.0 · License CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*
