We Ran This Audit on Our Own Agents First. Two Failed. We Published the Reports. Then We Built the Standard.
Most AI governance frameworks are written by people who have never shipped an agent into production. You can tell because they describe what agents should do, not what they actually do when the environment goes sideways.
I have shipped agents into production. I run eight software brands on a shared AI operating system with a core team of four people. When we started building GOVENANT — an open standard for AI agent accountability — the first thing we did was run the audit on ourselves.
Two agents failed. We published those reports.
That is where this started.
The Problem the Failures Revealed
The failure mode was not dramatic. The agents did not hallucinate wildly or cause obvious harm. They just could not prove what they had done.
When we asked them to produce a verifiable record of their actions — what they touched, what they decided, what they changed — the logs were either missing, ambiguous, or self-reported without any external verification. The agents said they had completed their tasks. We had no substrate-level evidence that they had.
That gap — between what an agent claims and what the environment can confirm — is what GOVENANT is built to close. We call the disease performed autonomy: agents that look like they are delivering, but are actually just generating motion. The flatline of real output hidden underneath a schedule of convincing activity.
If you cannot tell the difference between an agent that delivered and an agent that performed delivery, you do not have governed autonomy. You have a liability you cannot audit.
What the Standard Actually Measures
GOVENANT grades agent systems on four conformance levels:
- Discovered — the agent can be found and identified in the substrate
- Logged — its actions are recorded in a durable, external log
- Asserted — it makes verifiable claims about its own behavior
- Delivered — those claims can be confirmed against the actual environment
Most teams building with agents today are operating somewhere between Discovered and Logged. They know the agents exist. They have some logs. But the logs are self-reported, the assertions are unverified, and delivery confirmation is manual — or absent.
The audit exists to show you exactly where you are, not where you think you are.
The MCP: Two Minutes, Read-Only, Nothing Leaves Your Machine
Today we are releasing the GOVENANT Audit MCP.
It connects to Claude, Cursor, or VS Code in under two minutes. It runs a read-only sweep of your agent substrate — no writes, no modifications, no data sent anywhere. Your prompts and your data stay on your machine. There is no credit card, no account creation, no trial period.
The MCP returns a shareable report that grades your agent system against GOVENANT's four conformance levels and shows you specifically where the gaps are. Not a score out of 100. Not a vague maturity matrix. A concrete, level-by-level breakdown of what your agents can and cannot prove about themselves.
The same report format we published when our own agents failed.
Run it here: govenantstandard.org/audit
Why We Made It Free and Open
Governance tooling that costs money to evaluate is governance theater. If you have to buy access to find out whether your agents are accountable, the incentive structure is already broken.
We made the audit free because the problem it surfaces is real regardless of whether you ever work with us. If your agents cannot prove what they did, that is a risk you are carrying right now. You should know about it before your users find out for you.
GOVENANT is an open standard developed at iii Partners — the same governance logic that runs underneath the agent operating system I build with CEOs and founding teams. The audit MCP is the entry point. The standard itself is public. The framework is yours to use, fork, or critique.
The Honest Logbook
When we published the failure reports from our own agents, a few people asked why we would do that publicly. The answer is straightforward: if we are going to ask other teams to hold their agents to an accountability standard, we have to hold ours to it first. Credibility in this space is not built by publishing case studies about success. It is built by showing what broke and what you learned.
Two agents failed. The reports are still up. The standard came from what the failures taught us.
If you are building with agents in production — or evaluating whether to — run the audit. It takes two minutes. The results will tell you something true about your system that your current monitoring probably does not.
Work with Me
If the audit surfaces something you want to think through — or if you are a CEO building an AI-native operation and want to understand how governance fits into the broader architecture — the next step is a conversation. I use an AI assistant to learn about your situation first; it qualifies the fit before any call gets scheduled, so neither of us wastes time. If it makes sense, we talk. Start here: https://sfielder.com/talk?ref=blog