Making agent behavior provable after the fact
2026-10-06 — an incident you cannot reconstruct is an incident you cannot fix. When a coding agent deletes a directory, opens a key file, or pipes a remote script into a shell, the interesting moment is already over by the time a human looks: the terminal transcript is scrollback gone, the shell history does not attribute, and the CI log describes the build, not the process. The cousin of the same problem shows up on the server side, where operators now struggle to say which agent traffic did what at all. This post is about the narrow property we build first: provability.
What "provable" concretely means
- One line per observed action in a local, append-only JSONL audit record: agent family, process lineage, the command shape, a timestamp, and the policy verdict that was attached at the moment of the action
- Malformed events are rejected before they reach the log — zero bytes written. The record is trustworthy because a strict validator gates it, not because the reporter promised to be honest
- Secrets never land: every string that comes from the operating system passes a redaction layer before it is logged, covering credentials, key-value secrets, and key-file paths
- The record stays yours: owner-only file permissions, size and daily rotation, and a collector that makes no outbound network calls at all — the dependency closure is machine-checked in CI
The verdict you get today: would_block
Phase 0 is observe, evaluate, audit. A high-risk action is recorded with a would_block verdict and then still proceeds — proof comes before prevention, and the evaluation set is published with the claim: a frozen 40-case run gives 90 percent detection on the danger class, zero false-positive verdicts on the benign class, and zero credential leaks. Enforcement is a later phase with its own announcement; nothing on this blog gets ahead of the code.
Proof also means "what-if", replayable
Landed on main after the current tag, not yet in the shipped binaries: a replay engine reads an existing audit corpus against a candidate rule file and emits per-event verdicts — deterministically, twice byte-identical, with zero side effects on the running system. That answers the question every post-incident review actually asks: what would tomorrow's policy have said about last Tuesday's events? Reproducibly, in seconds.
Try it on your own machine
- The read-only control surface in v0.5.0 — status, audit-tail, timeline — shows you the record exactly as it is written
- Downloads page: the full build matrix with per-file SHA-256, and first-run guidance for unsigned builds (we say so plainly rather than pretending SmartScreen and Gatekeeper will not prompt)
- Source, tags, and issues: github.com/411160007/20131-agentruntime
If you run coding agents and have an incident-shape you currently cannot prove after the fact, write to us: <a href="mailto:[email protected]">[email protected]</a> or via the <a href="/en/support/">support page</a>. Every email is read by a human, and the cases you describe are what the next rule file and the next evaluation set are built from — that sentence is the whole point of this page.