Observability for a swarm of AI agents: what they are doing, who allowed it, and how to stop them
OpenAI is building a dedicated mechanism to watch and control large numbers of concurrently running agents. Why agent observability differs from classic metrics, logs and traces, why it sits closer to security than to operations, and what to put in place while you still only have a handful of agents.
At Dev Day, OpenAI said it is building a specialized mechanism for watching and controlling large numbers of AI agents running at the same time. Little is public about how it works, so I will talk about it in general terms. What matters is not the implementation but the signal: a major model vendor considers the problem serious enough to build a separate layer for it.
The problem is easy to state and awkward to solve. While you run one agent, a person can keep an eye on it. Once an organization runs hundreds or thousands of autonomous agents, following every step with ordinary observability tooling stops working. Next to AgentOps, a separate market is taking shape: AI agent security and AI agent observability as disciplines of their own.
Here is my angle on it. Familiar observability answers the question "is the system alive and how fast does it respond". For a swarm of agents the question is different: "what are they doing right now, who allowed it, and how do I stop it". The second question leads to a different architecture, and it is far cheaper to put that architecture in place while you have ten agents than to retrofit it at three hundred.
Classic observability measures the health of a system
Metrics, logs and traces were built for a specific job: tell whether a service works and where it slows down. I covered that in detail in a piece on the three pillars of observability for product teams. Everything there revolves around the request: a request arrived, crossed five services, returned a status code in so many milliseconds. Metrics aggregate, logs capture detail, traces stitch one path together.
That apparatus still works for agents, but it answers only half the questions. It will tell you the agent finished in forty seconds and did not crash. It will not tell you why the agent reached into the store of customer contracts, who asked it to, or whether a human was at the end of that chain at all.
The difference comes down to this. A conventional service performs a predefined action: the endpoint is fixed and its behavior is determined by code. An agent receives a goal and picks its own tools to reach it. Two runs from the same phrasing produce different sequences of calls. System health stays green in both cases, even though in one run the agent read a reference table and in the other it rewrote thirty rows in a production database.
The unit of observation shifts from the request to the intent
A swarm of agents calls for watching a different entity. The useful unit here is the pair "intent and action": the goal the agent set itself, and the tool call it made as a result. A request is too small a unit, and a whole task is too large.
Several practical things follow from that.
-
You need an end-to-end trail from a human instruction to a tool call. A person asks for a report on overdue receivables. The agent decides it needs the payments table, calls the export tool, gets a permission denial, reframes the task, goes to another source. That entire path has to form one chain under one identifier running from the human to the last call. If the agent spawned three sub-agents, they inherit that identifier; otherwise the link between the instruction and its consequences is lost.
-
Record attempts, not only successes. A permission denial in an ordinary service is a line in the log at severity four. For an agent, a denial tells you where it tried to reach. A run of denials from one agent is more interesting than any green metric.
-
Store the tool's input and output, not just the fact of the call. Otherwise an incident review a week later runs into the line "agent called
update_records" with no idea what it actually updated. -
Watch the population, not the instance. One agent making forty calls in an hour is normal. Forty agents hitting the same source in lockstep after a prompt update is an event, and it is visible only at the swarm level.
Why this sits closer to security than to operations
This is the shift I most want an owner or a CTO to internalize. An agent moves through your systems with a valid token. To any access control system it is a legitimate client: permissions were granted, the session is live, the requests are correctly signed. The classic perimeter waves that traffic through, because there is nothing to hold against it.
I worked through an adjacent story in a piece on AI assistants and identity as an attack surface. With a swarm, that story scales up. What separates an agent from a compromised employee account is that an agent works at machine speed and never sleeps, so a flaw in its goal-setting plays out in minutes.
The anomaly here is measured in behavior. It has no malware signature and no suspicious IP. What looks wrong is the shape of the activity:
- an agent that normally reads three reference tables starts enumerating the contents of a store;
- the share of write operations in its behavior triples within a day;
- the call chain no longer traces back to a live human instruction and has closed on itself;
- the agent systematically probes tools it has no rights to.
None of those events break anything. All four are behavioral deviations, and you can catch them only if you hold a baseline of normal behavior for that particular agent. Which leads to the conclusion I consider the main one: agent observability is built as access control with history, and the operational metrics ride along as a secondary concern.
Alerts arrive late, so you need a circuit breaker
A classic alert is a notification: a threshold was crossed, the on-call engineer got a message, the on-call engineer started looking. Minutes pass between firing and response, and for human-paced systems that is acceptable.
An agent makes hundreds of calls in those minutes. So instead of a notification you need a mechanism that stops the action by itself, on a rule set in advance. I call these budgets and circuit breakers.
A budget is a limit the agent has to stay inside: number of tool calls per period, tokens spent, how many records it may change in one run, total cost of the task. On the money side I wrote a separate piece on token budgets and controlling LLM costs, and the logic is identical: a limit applied automatically works, developer carefulness does not.
A circuit breaker is the rule that halts execution when a budget is exceeded or an action lands on a red list. The stop has to happen before the tool call, not after it. The gap between "the agent emailed a thousand customers and we saw it happen" and "the agent tried to email a thousand customers and the call was refused" is the gap between an incident and a log entry.
And the third piece, the one almost everyone forgets: you need a mass stop. A switch that kills a whole class of agents at once rather than a single process, selected by attribute: every agent on this prompt version, every agent with access to the payments domain, every agent owned by one team. Test that switch before you need it, and test it regularly. A switch nobody has ever pressed is an assumption rather than a mechanism.
The practical minimum while you still have ten agents
Retrofitting at three hundred agents is expensive. Here is what I suggest putting in place now, even if today you have three scenarios and one enthusiast maintaining them.
-
Keep a registry of agents. A list: which agent, who owns it on the business side, what its goal is, which systems it can reach, which prompt and model version it runs. That is ten rows in a spreadsheet today and the only way to answer "how many of these do we actually have" a year from now. Agents created around the registry are shadow IT, just faster.
-
Give every agent its own identity. A separate technical account and a separate token per agent rather than one shared service key for all of them. Without that you physically cannot build a behavioral baseline or revoke one agent's access while leaving the rest running.
-
Grant minimal rights and review them. An agent inherits the access lifecycle problem in its purest form: rights were granted for one task, the task changed, the rights stayed. The logic matches what I described for people in a piece on least privilege in practice. The difference is that an agent uses everything it was given, and uses it fast.
-
Carry one identifier from the human to the tool call. A human instruction mints the identifier, sub-agents inherit it, and every call records it. This is the cheapest item on the list at the start and the most expensive to add later.
-
Log intent, call, input, output and denial. Five fields are enough to reconstruct the picture a month later. Decide in advance what you are allowed to retain: those calls will contain personal data, and the retention period belongs in a conversation with your lawyers now rather than after a regulator asks.
-
Put budgets on every agent from day one. Call limits, write limits, cost limits. Start generous on purpose: the job of the first month is learning what normal looks like, not squeezing the agent.
-
Build and test the mass stop. One mechanism that halts agents by attribute. Run a drill, and write down the time from decision to every process actually being down.
-
Name an owner. A swarm of agents needs a person accountable for the registry, the budgets and the review of deviations. At ten agents that is half a role. At three hundred it is a function, and it is better grown out of an existing role than created in response to an incident.
Honest caveats
-
None of these mechanisms cancels out a badly stated task. An agent with perfect tracing and careful budgets will still do exactly what it was told, including the part the person did not mean. Observability shortens time to detection and caps the damage; the quality of the instructions stays with people.
-
A behavioral baseline takes time and data. For the first weeks you will see deviations where there are none and miss real ones. That is a normal learning period, and it belongs in the plan rather than coming as a surprise.
-
This tooling market is young and the standards have not settled. Model vendors are building their own control mechanisms, independent companies are building theirs, and how any of it will interoperate is something nobody honestly knows yet. So I suggest keeping your own record of what your agents did on your side, instead of relying on one vendor's console.
-
The deep part of the work stays with specialists. I help map the risk, build the control loop around the agents and put the right questions to vendors. Detailed tuning of detection systems and formal verification of behavior are specialist jobs, and I bring those people in rather than pretending to be an expert in everything.
In short
- Classic observability answers a question about system health, while a swarm of agents needs an answer to what they are doing, who allowed it, and how to stop it.
- The unit of observation becomes the pair "intent and action", and you need an end-to-end trail from a human instruction to a tool call, denials and sub-agent activity included.
- This is closer to a security problem: an agent with a valid token looks like legitimate traffic, and the anomaly is measured in behavior rather than in a signature.
- Alerts arrive too late for machine speed, so you need budgets, circuit breakers that stop a call before it executes, and a mass stop you have actually tested.
- The starting minimum: an agent registry, a distinct identity per agent, minimal rights, one end-to-end identifier, five fields in the log, budgets from day one, a mass-stop drill and a named owner.
If you already run a few agents and you are sizing up what thirty of them will look like, the cheapest time to work this through is now. Discuss your case through the form on the main page; the first conversation commits you to nothing.