> For the complete documentation index, see [llms.txt](https://osintelligence-llc.gitbook.io/osintelligence/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://osintelligence-llc.gitbook.io/osintelligence/part-iii-the-evidence-what-broke/10-the-guard-changes-at-23-26z/references-and-provenance.md).

# References & provenance

### References

\[1–2] Anthropic (2026). "Introducing Claude Opus 4.7" (News, 2026-04-16); Claude Opus 4.7 product page.

\[3] OSINTelligence LLC (2026). P7 SSD D-003.2 abort note (the ratified-and-retained falsification precedent; sovereign repository, RESULTS tree).

\[4] OSINTelligence LLC (2026). Forgejo sovereign git forge deployment record.

\[5] Kistner, J. (2026). [*Sixteen Practices for Sovereign Human–AI Collaboration.*](/osintelligence/part-ii-the-discipline/5-sixteen-practices.md) The series' methodology reference (Chapter 5).

\[6] Shadish, W. R., Cook, T. D., & Campbell, D. T. (2002). *Experimental and Quasi-Experimental Designs for Generalized Causal Inference.* Houghton Mifflin.

\[7] Gao, Z., Bird, C., & Barr, E. T. (2017). "To Type or Not to Type: Quantifying Detectable Bugs in JavaScript." *ICSE 2017.*

\[8] Bird, C., Rigby, P. C., Barr, E. T., et al. (2009). "The Promises and Perils of Mining Git." *MSR 2009.*

\[9] Devanbu, P., Zimmermann, T., & Bird, C. (2016). "Belief & Evidence in Empirical Software Engineering." *ICSE 2016.*

\[10–13] Release-week coverage: XBOW visual-acuity benchmark; CNBC and Axios release reporting (2026-04-16); The AI Corner migration guide.

\[14] Takerngsaksiri, W., et al. (2024). "Human-In-the-Loop Software Development Agents (HULA)." [arXiv:2411.12924](https://arxiv.org/abs/2411.12924).

\[15] Shao, E., et al. (2025). "SciSciGPT: Advancing Human–AI Collaboration in the Science of Science." *Nature Computational Science* ([arXiv:2504.05559](https://arxiv.org/abs/2504.05559)).

\[16] Chambers, C. D. (2013). "Registered Reports: A New Publishing Initiative at Cortex." *Cortex* 49(3).

\[17] Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). "The Preregistration Revolution." *PNAS* 115(11).

\[18] Pineau, J., et al. (2021). "Improving Reproducibility in Machine Learning Research." *JMLR* 22.

\[19] OSINTelligence LLC (2026). P7 SSD pre-registration, SHA-256 815e0a35… (sovereign repository).

\[20] OSINTelligence LLC (2026). Bash-on-Windows path-conversion quirk, session log 2026-04-16.

\[21] Evans, R., et al. (2024). "Evaluating Human-AI Collaboration: A Review and Methodological Framework." [arXiv:2407.19098](https://arxiv.org/abs/2407.19098).

\[22] Fragiadakis, G., et al. (2024). "Human-AI collaboration is not very collaborative yet." *Frontiers in Computer Science* 6:1521066.

\[23] Abasi-amefon, A., et al. (2026). "Advancing Decision-Making through AI-Human Collaboration." *Group Decision and Negotiation* (Springer).

\[24] Panke, S. (2025). "How Can (A)I Research This? An Autoethnographic Exploration of Generative AI." *Qualitative Social Research* (SAGE).

\[25] Wiles, F. (2025). "Recursive Cognition in Practice." *International Journal of Qualitative Methods* (SAGE).

\[26] Koopman, W. J., Watling, C. J., & LaDonna, K. A. (2020). "Autoethnography as a Strategy for Engaging in Reflexivity." *Qualitative Health Research.*

\[27] Wataoka, K., Takahashi, T., & Tsuchiya, R. (2024). "Self-Preference Bias in LLM-as-a-Judge." [arXiv:2410.21819](https://arxiv.org/abs/2410.21819).

\[28] Wallace, E., et al. (2024). "The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions." [arXiv:2404.13208](https://arxiv.org/abs/2404.13208).

\[29] Semmelrock, H., et al. (2025). "Reproducibility in Machine Learning-based Research." *AI Magazine* (Wiley).

\[30] Lee, S., et al. (2025). "Facilitating Longitudinal Interaction Studies of AI Systems." *UIST 2025 Adjunct.* [DOI 10.1145/3746058.3758469](https://doi.org/10.1145/3746058.3758469).

\[N1] Jin, J., et al. (2026). "Capable but Unreliable: Canonical Path Deviation as a Causal Mechanism of Agent Failure in Long-Horizon Tasks." [arXiv:2602.19008](https://arxiv.org/abs/2602.19008).

\[N2] Xu, B., et al. (2025). "AgentIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios." [arXiv:2505.16944](https://arxiv.org/abs/2505.16944).

\[N3] Zilian, R., et al. (2025). "Evaluating Goal Drift in Language Model Agents." [arXiv:2505.02709](https://arxiv.org/abs/2505.02709).

\[N4] HORIZON Benchmark Authors (2026). "The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break." [arXiv:2604.11978](https://arxiv.org/abs/2604.11978).

\[N5] Du, M., He, F., Zou, N., Tao, D., & Hu, X. (2024). "Shortcut Learning of Large Language Models in Natural Language Understanding." *CACM*; [arXiv:2208.11857](https://arxiv.org/abs/2208.11857).

\[N6] IBM Research (2025). "Towards Enforcing Company Policy Adherence in Agentic Workflows." [arXiv:2507.16459](https://arxiv.org/abs/2507.16459); *EMNLP 2025 Industry Track.*

\[N7] Hierarchical Safety Benchmark Authors (2025). "Evaluating LLM Agent Adherence to Hierarchical Safety Principles." [arXiv:2506.02357](https://arxiv.org/abs/2506.02357).

\[N8] Microsoft (2025). "Conversation Compaction in the Microsoft Agent Framework." Microsoft Learn.

\[N9] LangChain (2025). "Your Harness, Your Memory." Engineering blog.

\[N10] Raschka, S. (2025). "Components of a Coding Agent." *Ahead of AI.*

\[N11] JetBrains Research (2025). "Cutting Through the Noise: Smarter Context Management for LLM-Powered Agents."

\[N12] Weaviate (2025). "Context Engineering: LLM Memory and Retrieval for AI Agents."

Provenance of record (this paper's own subject matter): the composition of record ran under Claude Opus 4.7 with Claude Sonnet (web) contributing status tracking and routing, all under Jamey Kistner's direction as sole human author. Scaffold committed 2026-04-16T23:52Z, the first full-length post-transition artifact. All pre-transition citations refer to sealed artifacts composed under Claude Opus 4.6; those attributions are preserved as correct research provenance. The §10 catalogue grew append-only through Cycle-14; the literature-enhancement pass (refs \[21]–\[30]) and the \[N] agentic-reliability set were added under the operator's directive with no rewriting of pre-existing narrative.

### AI-assistance disclosure

Large language models were used as research tools in the preparation of this chapter: Claude Opus 4.7 (Anthropic; post-transition, paper composition of record); Claude Opus 4.6 (Anthropic; pre-transition baseline); Claude Sonnet, web (Anthropic; status tracking, literature pulls, feedback routing). The model versions and roles named above keep the provenance of this chapter auditable. No AI system is listed as an author or credited as a contributor, in line with COPE and ICMJE guidance: an AI system cannot take responsibility for the work, cannot assert competing interests, and cannot enter a licence agreement. The author verified every claim in this chapter against the sealed artifacts and is solely accountable for it.

**Citation (preferred):** Kistner, J. (2026). *The Guard Changes at 23:26Z: An Instrumented Natural Experiment in Model-Over-Model Performance for Sustained Solo-Operator Human–AI Research Collaboration*, version 1.0.0. OSINTelligence LLC research whitepaper. Cited in-series by title.

**License:** CC BY 4.0 (text). Code and data artifacts MIT per repository license.

**Companion papers (this series):** [*Sixteen Practices*](/osintelligence/part-ii-the-discipline/5-sixteen-practices.md) (Chapter 5; the collaboration method under test) · [*The Drift Taxonomy*](/osintelligence/part-iii-the-evidence-what-broke/9-the-drift-taxonomy.md) (Chapter 9; the field record this catalogue seeded) · [*The Sovereign Triad*](/osintelligence/part-i-the-architecture/1-the-sovereign-triad.md) (Chapter 1; the architecture the §12 proposals point toward). Cited in-series by title.

***

*The Sovereign Stack · The Guard Changes at 23:26Z · Chapter 10 · Part III · v1.0.0 · License CC BY 4.0 · © Jamey Kistner, OSINTelligence LLC*
