> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/part-iv-the-evidence-what-worked/16-sovereign-domain-pruning/references-and-provenance.md).

# References & provenance

### References

Bean, et al. (2025). Construct-validity analysis for LLM benchmarks. *NeurIPS 2025 Datasets & Benchmarks.* [arXiv:2511.04703](https://arxiv.org/abs/2511.04703).

Bland, J. M., & Altman, D. G. (1986). "Statistical methods for assessing agreement between two methods of clinical measurement." *The Lancet* 327(8476):307–310.

Cassano, F., et al. (2023). "MultiPL-E: A Scalable and Polyglot Approach to Benchmarking Neural Code Generation." [arXiv:2208.08227](https://arxiv.org/abs/2208.08227).

Cohen, J. (1988). *Statistical Power Analysis for the Behavioral Sciences* (2nd ed.). Lawrence Erlbaum.

Dettmers, T., et al. (2023). "QLoRA: Efficient Finetuning of Quantized LLMs." *NeurIPS 2023.*

Efron, B. (1979). "Bootstrap methods: Another look at the jackknife." *Annals of Statistics* 7(1):1–26.

Fisch, A., et al. (2024). "Stratified Prediction-Powered Inference (StratPPI)." *NeurIPS 2024.* [arXiv:2406.04291](https://arxiv.org/abs/2406.04291).

Fogliato, R., et al. (2024). Evaluation-methodology reference. [arXiv:2406.07320](https://arxiv.org/abs/2406.07320).

Frantar, E., & Alistarh, D. (2023). "SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot." *ICML 2023.*

Kargaran, A. H. (2025). ICLR rejection-ground analysis (benchmark-selection transparency; pre-registration backflow as Rejection Ground #2). [arXiv:2511.15462](https://arxiv.org/abs/2511.15462).

Kurtic, E., et al. (2024). "Give Me BF16 or Give Me Death? Accuracy-Performance Trade-Offs in LLM Quantization." [arXiv:2411.02355](https://arxiv.org/abs/2411.02355). The Component A protocol-class anchor; Table 4 supplies the ≤ 0.10 nats LoA envelope.

Liang, P., et al. (2023). "Holistic Evaluation of Language Models (HELM)." [arXiv:2211.09110](https://arxiv.org/abs/2211.09110).

Ma, X., et al. (2023). "LLM-Pruner: On the Structural Pruning of Large Language Models." *NeurIPS 2023.*

Men, X., et al. (2024). "ShortGPT: Layers in Large Language Models are More Redundant Than You Expect." [arXiv:2403.03853](https://arxiv.org/abs/2403.03853).

MLPerf (2019). "MLPerf Inference Benchmark." [arXiv:1911.02549](https://arxiv.org/abs/1911.02549).

Perlitz, Y., et al. (2024). "BenchBench: Benchmark Agreement Testing." [arXiv:2407.13696](https://arxiv.org/abs/2407.13696).

Lasby, M., Lazarevich, I., Sinnadurai, N., Lie, S., Ioannou, Y., & Thangarasa, V. (2025). "REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression." *ICLR 2026*; [arXiv:2510.13999](https://arxiv.org/abs/2510.13999).

Rein, D., et al. (2023). "GPQA: A Graduate-Level Google-Proof Q\&A Benchmark." [arXiv:2311.12022](https://arxiv.org/abs/2311.12022).

Spearman, C. (1904). "The proof and measurement of association between two things." *American Journal of Psychology* 15:72–101.

Srivastava, A., et al. (2023). "Beyond the Imitation Game (BIG-bench)." [arXiv:2206.04615](https://arxiv.org/abs/2206.04615).

Su, Z., Li, Q., Zhang, H., Qian, Y., Xie, Y., & Yuan, K. (2025). "Unveiling Super Experts in Mixture-of-Experts Large Language Models." *ICLR 2026*; [arXiv:2507.23279](https://arxiv.org/abs/2507.23279).

Sun, M., et al. (2024). "A Simple and Effective Pruning Approach for Large Language Models (Wanda)." *ICLR 2024.*

Wang, K., Lyu, T., Su, G., Geiping, J., Yin, L., Canini, M., & Liu, S. (2025). "When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs." [arXiv:2510.22228](https://arxiv.org/abs/2510.22228).

Yauney, G., et al. (2025). Benchmark-reliability floor analysis (\~250-example floor for stable pass-rate ranking). [arXiv:2510.08730](https://arxiv.org/abs/2510.08730).

Zhuo, T. Y., et al. (2024). "BigCodeBench: Benchmarking Code Generation with Diverse Function Calls." [arXiv:2406.15877](https://arxiv.org/abs/2406.15877).

Kistner, J. (2026). [*Sovereign In-Distribution Imatrix Calibration*](/osintelligence/part-iv-the-evidence-what-worked/15-sovereign-imatrix-calibration.md) (Chapter 15) and [*Cross-Toolchain Falsification and Surgical Patch-Pick Resolution of a Production M-RoPE Heap Over-Read in llama.cpp*](/osintelligence/part-iv-the-evidence-what-worked/20-mrope-serving-path-rca.md) (Chapter 20), in-series companions; the latter supplies the patched serving binary (commit 822047a0a) this paper's eval-bank runs on. [*Sovereign Big-Model Compression*](/osintelligence/part-iv-the-evidence-what-worked/17-sovereign-big-model-compression.md) (Chapter 17) is the successor arc: the expert-pruning thesis first tested here, carried to a 122B mixture-of-experts and cleared on the deployed surface by a pre-registered served quality gate.

**Evidence & seal.** The pre-registration behind this chapter, the frozen mismatch-control experiment reproduced verbatim with the canonical self-hash quoted from its sidecar, its composite-FAIL verdict table, and the integrity envelope, is in the record's evidence section: [35B Prune Experiment (P3, mismatch control)](/osintelligence/evidence-and-seals/35b-prune-experiment-p3-mismatch-control.md).

### AI-assistance disclosure

Large language models were used as research tools in the preparation of this chapter: Claude Opus 4.7 (Anthropic). The model versions and roles named above keep the provenance of this chapter auditable. No AI system is listed as an author or credited as a contributor, in line with COPE and ICMJE guidance: an AI system cannot take responsibility for the work, cannot assert competing interests, and cannot enter a licence agreement. The author verified every claim in this chapter against the sealed artifacts and is solely accountable for it.

**Citation (preferred):** Kistner, J. (2026). *Sovereign Domain Pruning: Destructive Expert-Pruning of a 35B-A3B MoE Against a Sovereign CTI Corpus*, version 1.0.0. OSINTelligence LLC research whitepaper. Cited in-series by title.

**License:** CC BY 4.0 (text). Code and data artifacts MIT per repository license.

**Corresponding author:** Jamey Kistner, <jamey.kistner@osintelligence.io>, OSINTelligence LLC (Columbus, OH).

***

*The Sovereign Stack · Sovereign Domain Pruning · Chapter 16 · Part IV · v1.0.0 · License CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*
