> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/part-iv-the-evidence-what-worked/17-sovereign-big-model-compression/references-and-provenance.md).

# References & provenance

### References

Lasby, M., Lazarevich, I., Sinnadurai, N., Lie, S., Ioannou, Y., & Thangarasa, V. (2025). "REAP the Experts: Why Pruning Prevails for One-Shot MoE Compression." [arXiv:2510.13999](https://arxiv.org/abs/2510.13999) (Cerebras Systems; University of Calgary).

Ding, Y., Wang, J., Yang, G., Jing, Y., Guo, J., Liu, X., & Tao, D. (2026). "Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression." [arXiv:2606.18304](https://arxiv.org/abs/2606.18304). (Channel-level structural MoE pruning; the comparator method in the 35B study.)

Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., & Dean, J. (2017). "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer." [arXiv:1701.06538](https://arxiv.org/abs/1701.06538).

Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2023). "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers." *ICLR 2023*; [arXiv:2210.17323](https://arxiv.org/abs/2210.17323).

Gloeckle, F., Youbi Idrissi, B., Rozière, B., Lopez-Paz, D., & Synnaeve, G. (2024). "Better & Faster Large Language Models via Multi-token Prediction." [arXiv:2404.19737](https://arxiv.org/abs/2404.19737).

Gerganov, G., et al. (2023–). *llama.cpp*: importance-matrix (imatrix) quantization and KL-divergence measurement tooling. Open-source project.

Que, H., Liu, J., Zhang, G., et al. (2024). "D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models." *NeurIPS 2024*; [arXiv:2406.01375](https://arxiv.org/abs/2406.01375) (validated on Qwen-family models).

Gu, J., Yang, Z., Ding, C., Zhao, R., & Tan, F. (2024). "CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language Models." *EMNLP 2024*; [arXiv:2407.17467](https://arxiv.org/abs/2407.17467).

Kistner, J. (2026). [*Sovereign In-Distribution Imatrix Calibration*](/osintelligence/part-iv-the-evidence-what-worked/15-sovereign-imatrix-calibration.md) (Chapter 15). OSINTelligence LLC.

Kistner, J. (2026). [*Sovereign Domain Pruning*](/osintelligence/part-iv-the-evidence-what-worked/16-sovereign-domain-pruning.md) (Chapter 16). OSINTelligence LLC.

**In-series companions.** *Sovereign Domain Pruning* (Chapter 16) is the direct predecessor, establishing destructive expert pruning at 35B scale; *Sovereign In-Distribution Imatrix Calibration* (Chapter 15) supplies the calibration technique this arc carries to 122B, and is the sibling axis of the same thesis, that the operator’s own corpus decides what to keep. [*Sovereign Optimization Flywheel*](/osintelligence/part-ii-the-discipline/8-sovereign-optimization-flywheel.md) (Chapter 8) places both as compounding axes of one optimization loop, and this paper is their joint instance at the largest scale the series has attempted. [*Sovereign Sustainability*](/osintelligence/part-v-the-frontier/22-sovereign-sustainability.md) (Chapter 22) takes the deployment result as its democratization datum, a frontier-class mixture-of-experts served for a single operator on a single consumer card. [*Sixteen Practices*](/osintelligence/part-ii-the-discipline/5-sixteen-practices.md) (Chapter 5) carries the pre-registration and falsification discipline that the quality gate of §6 and the reported failure of §8 are run under, including the rule that a failed experiment is published beside the passing one.

*Related one-shot MoE-pruning literature surveyed during method selection: Chen et al. (2022) task-specific expert pruning (*[*arXiv:2206.00277*](https://arxiv.org/abs/2206.00277)*); Chowdhury et al. (2024) provably-effective expert pruning (*[*arXiv:2405.16646*](https://arxiv.org/abs/2405.16646)*); Xie et al. (2024) MoE-Pruner (*[*arXiv:2410.12013*](https://arxiv.org/abs/2410.12013)*). Every citation above was web-verified against its primary source; audit trail in the companion ledger. No internal file, path, or hash identifiers appear in the body.*

**Evidence & seal.** The pre-registered served quality gate behind this chapter, its frozen pre-registration reproduced verbatim with the canonical self-hash quoted from its sidecar, the PASS verdict table, and the 11-artifact integrity envelope, is in the record's evidence section: [Pruned-122B Quality Gate (D7)](/osintelligence/evidence-and-seals/pruned-122b-quality-gate-d7.md).

### AI-assistance disclosure

Claude (Anthropic) was used as a research tool in the preparation of this chapter, assisting with drafting, structuring and successive revision. The specific model versions were not recorded at the time of authorship and so are not named here. No AI system is listed as an author or credited as a contributor, in line with COPE and ICMJE guidance: an AI system cannot take responsibility for the work, cannot assert competing interests, and cannot enter a licence agreement. The author verified every claim in this chapter against the sealed artifacts and is solely accountable for it.

**Citation (preferred):** Kistner, J. (2026). *Sovereign Big-Model Compression: Porting One-Shot Expert Pruning onto a 122B Mixture-of-Experts and Serving It, Trainable, on a Single Consumer GPU*, version 1.0.1. OSINTelligence LLC.

**License:** CC BY 4.0 (text). All referenced code artifacts are MIT licensed unless otherwise noted.

**Corresponding author:** Jamey Kistner, <jamey.kistner@osintelligence.io>, OSINTelligence LLC (Columbus, OH).

***

*The Sovereign Stack · Sovereign Big-Model Compression · Chapter 17 · Part IV · v1.0.1 · License CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*
