/public/resources/evidence-authority-execution-integrity.pdf. Do not silently modify claims; replace the working-draft prose below only with approved source text. Executive Summary
Highly autonomous agent systems expose a category of failure that content filters and access controls do not address: the subject of governance influencing the evidence, authority, or execution path by which it is governed.
This paper examines three integrity problems, evidence integrity, authority integrity, and execution integrity, and proposes architectural principles for keeping all three outside the control of the governed subject. It uses a public agent incident as a concrete reference and describes how an external semantic governance architecture could relate to each failure surface.
1. Incident
The paper opens with a publicly reported agent incident as a shared reference point. The incident is used to illustrate failure surfaces, not to assign fault. The relevant question is architectural: when an autonomous system can influence its own evaluation, authorization, or execution, what independent structure remains?
2. Governance Integrity
Governance integrity holds when the mechanisms that evaluate, authorize, and permit action are not writable by the subject being governed. SIA frames this as a property of architecture rather than of model behavior.
3. Human Intention
The governing objective originates in observable human communication. SIA does not claim to read an unexpressed mental state; it derives a machine-addressable communicative objective and associated semantic conditions from what was actually communicated.
4. Accessibility Is Not Authority
That a system can reach a capability does not mean it is authorized to use it. Discovery is not permission. Authorization must be evaluated against a reference the subject does not control.
5. Reasoning vs Agency
Reasoning about an action is distinct from taking it. Governance should move toward the boundary between reasoning and consequential action, where an external authorization step can precede execution.
6. Evidence the Subject Cannot Write
Evidence used to judge a subject must be produced independently of that subject. If the subject can author or alter the evidence, the evidence cannot be trusted for control decisions.
7. Authority the Subject Cannot Grant Itself
Authority to act must come from an externally recognized source. A subject, or a peer subject, cannot enlarge its own authority merely by asserting or messaging it.
8. Execution the Subject Cannot Obtain Without Authorization
Whether an unauthorized action can execute should not be under the control of the subject proposing it. Execution control belongs outside the governed subject.
9. Six Governance Properties
- The governing objective is external to the subject.
- Evidence is produced independently of the subject.
- Authority is granted by a recognized external source.
- Execution of consequential action requires prior authorization.
- Verification compares behavior against the same retained state.
- A governance record preserves what occurred.
10. How the Architecture Could Have Interrupted the Incident
The paper walks through where external evidence, external authority, and external execution control could have introduced an interruption point, described at an architecture level and without claiming a guaranteed outcome for any specific system.
11. Multi-Agent Systems
In multi-agent settings, meaning and authority can change at handoffs. Governed handoffs carry objective, protected conditions, authority, provenance, and unresolved findings across boundaries.
12. RSI Implications
The same separation, capability change distinct from objective-governance change, is discussed as a prospective direction for capability-modifying systems. This is presented as a research hypothesis.
13. What SIA Does Not Solve
SIA does not replace cybersecurity or human accountability, does not guarantee regulatory compliance, and does not claim to control autonomous recursive self-improvement. These boundaries are stated explicitly.
14. Research Program
The paper closes with a research program: persistence, authority, continuity, verification, and robustness, each with a falsifiable framing and an intended comparison against conventional baselines.
Conclusion
The central principle is a constraint on architecture, not a promise about model behavior:
References
- TODO: Populate with the reference list from the approved paper.