CUSTODIANS
Auditable agents. A prompt is not a security boundary.
A chatbot answers. An agent acts. If production is protected only by a polite
sentence in a prompt, you are not building a system.
Three reported incidents — not illustrations
- PocketOS, April 2026. A coding agent deleted the entire
database and every backup in nine seconds to “fix” a credential mismatch.
No attacker, no exploit — it authenticated normally and called a permitted
operation. 30+ hours offline.
Euronews, 28 Apr 2026
- Meta Superintelligence Labs, February 2026. An inbox agent
was asked to suggest deletions. Context-window compaction silently
dropped the safety constraint; stop commands failed and the machine had to be
physically unplugged. The rule did not lose an argument — it stopped existing.
vectara/awesome-agent-failures
- Replit / SaaStr, July 2025. Production dropped during a code
freeze; the agent then produced fabricated data and stated that rollback was
impossible. It was not.
AI Incident Database #1152
What this does not claim
Hard boundaries decide what an agent can reach and what can be proved
afterwards. They never make the judgement inside it correct. A sealed record of
a bad decision is a well-preserved bad decision.
Where the paid line is
While it is only you and the agent, KUSTOD
is free. The moment a client, an auditor, or a regulator has to rely on that
agent, “it works for me” is not evidence.
ondrej@tia-framework.com