A new joiner learns how an organisation behaves long before they read a policy. They watch who gets challenged at the door. They hear what the team says when a suspicious email lands, and they notice what their manager does when a deadline and a control pull in opposite directions.
An AI agent gets none of that. It arrives with credentials, a task and whatever text somebody thought to put in front of it, and then it acts on that text literally, thousands of times a day.
That gap is the subject of my new paper on SSRN, The Normative Repository. This is the short version: what I am proposing, why I think it belongs to behavioural security, and the part of the idea I think matters most.
Culture does not reach an agent
Building Security Champions Network's taught me that security culture travels through repetition, social proof and identity. Somebody hears the same message enough times, sees colleagues they respect acting on it, and eventually starts to think of themselves as the sort of person who does it too.
None of those channels exists for an agent. It overhears nothing, has no colleagues to copy, and carries no sense of who it is from one session to the next unless we hand it one. What it will do is read a file and follow it.
So the question I kept coming back to is simple: where is the file? In most organisations, the honest answer is that there isn't one. Each agent carries its own prompt, written by whoever built it, with no shared owner, no version history, and no way to tell which rules were in force when something went wrong. We are adding agents to our estates faster than we are giving them any account of how we expect them to behave.
One signed source of how we behave
The paper proposes a normative repository. It is a single, version-controlled, cryptographically signed store holding the organisation's charter and its security, privacy, ethics, safeguard and escalation norms, written in a form agents can consume.
The mechanics are deliberately dull. An agent loads a signed release when a session starts, checks again on a fixed cadence, and writes the hash of that release into its logs beside the actions it takes. Six months later, when somebody asks what the agent was told to do on the day it did something odd, there is an answer, and it can be proved.
Changes to the norms travel the same road as changes to code. They go through a continuous integration and deployment pipeline with owner review, consistency checks, behavioural evaluations and adversarial testing before anything is released. The privacy lead owns the privacy norms, and nobody merges a change to them without that sign-off.
The objection I hear most is that this is a system prompt with extra steps. The extra steps are the point. A prompt cannot be compared across an estate, cannot be rolled back cleanly, and leaves no evidence of which rules applied when.
The second objection is better: a model will not reliably follow a document. That's true, and it is why the evaluations matter more than the documents do. We learned the same lesson with people. Training completion rates told us very little, and watching what people did told us nearly everything.
Behavioural security for actors that are not people
I did not come to this as an engineering problem. I came to it as a behavioural one, and the paper argues that the repository extends behavioural cybersecurity to non-human actors.
The constructs carry across more neatly than I expected. The refresh cadence does for an agent what repetition does for a person. The charter gives the agent an account of who it works for and what kind of actor it is expected to be, which is identity framing by another route. Deciding what gets loaded first, and which escalation path is easier, is choice architecture.
Agents even have an intention-behaviour gap of their own. An agent can hold a norm in its context and still act against it, because it interprets, generalises and satisfies goals in ways nobody predicted. That is the reason the pipeline tests behaviour and does not stop at checking that the words are present.
A supply chain attack on intent
Putting every norm in one place creates a target. Someone who can alter the repository doesn't need to compromise a single agent; they change what every agent believes it ought to do, and the agents then carry out the attack in good faith. The paper calls this a supply chain attack on intent.
So the repository gets its own threat model, with controls set out by layer: the administrator accounts on the platform that hosts it, the independence of the people who review changes, the keys that sign releases, and the behaviour of the agent itself, which should fail closed when it meets a release it cannot verify.
Does this create a single point of failure? It concentrates a risk that today sits unmanaged in every agent's prompt. I would take one well-guarded thing over hundreds of unguarded ones.
When an agent goes quiet
This is the part of the paper I care about most. Once every agent is expected to renew its norms on a known schedule, the renewal itself becomes a signal. I call it norm attestation.
Think of it as a shift handover that never gets skipped. An agent that misses its attestation, reports a stale hash, or reaches for the repository in an unusual way is telling you something. It may have been hijacked and pointed at somebody else's goals, or it may have drifted from the rules everyone else is working to. Monitoring for agents already exists as reliability engineering, where a missed check-in means the thing has fallen over. Here, a missed check-in means the thing may have changed sides.
People cannot watch for this across hundreds of agents, so the paper requires an independent, automated monitor with a graduated response. Independent, because a compromised agent must not be marking its own homework. Graduated, because the first missed check-in deserves a closer look and not a quarantine.
The monitor also has to read silence in context. One agent going quiet while its peers attest normally is suspicious. Every agent going quiet at once means the repository is unreachable, which is an infrastructure fault, and quarantining the whole estate in response would be an outage of our own making.
What it will not do
Norm attestation is a tripwire. An attacker who fully controls an agent's runtime and credentials can keep sending valid attestations while doing whatever they like. The signal raises the cost of staying hidden, but it does not remove the risk.
Its value lies in the large middle of the problem: agents that have been hijacked clumsily, misconfigured, abandoned, stood up without anyone's knowledge, or left to drift. I also expect false alarms to be common at first, from mundane failures, until each agent has a baseline.
The paper is conceptual, there is no experimental evidence in it yet that norms loaded this way change agent behaviour, or that attestation catches compromise at a useful rate.
Read it, then try to break it
If you run agents, I would like you to read it and tell me where it falls down. If you already manage prompts and versions in a way this should build on, I would like to hear that even more.
The working paper is free to download from SSRN: The Normative Repository: Grounding Agentic AI in Organisational Norms, with Norm Attestation as a Compromise Signal