mksim.pro
Back to all posts
AI 11 min read

The FTC turns to autonomous agents: what you need to be able to prove about yours

A US regulator is asking whether existing law is enough to hold companies accountable for what autonomous AI systems do. While the legal frame is argued over, the engineering requirement is already clear: audit trail, permissions, sandboxing and human approval stop being optional.

The US Federal Trade Commission has opened a broad inquiry into risks around the technologies of OpenAI and Anthropic. Among other things, the regulator is reported to be looking at cases where autonomous AI systems went beyond the task they were originally given. Nobody is talking about banning the technology. The question on the table is a different one: is existing law sufficient to hold companies accountable for what autonomous systems do.

Little about the inquiry is public, and I am not going to fill in the blanks. What matters is that the question is now being asked at all. Until now, accountability for AI lived in declarations and ethics frameworks. A regulator is now fitting ordinary legal machinery, the kind that already works for people and companies, onto agents. And every such mechanism has a practical consequence that arrives long before any rule does.

The legal frame will be contested, the engineering requirement already is not

The legal argument will take years. Who answers for an agent's action: the model developer, the integrator, the operating company, the employee who set the task. National regimes will diverge, definitions will be rewritten. All of that is noisy and far away.

Underneath any formulation of liability sits the same technical requirement, and it does not depend on how the argument ends. You will have to prove three things:

  • What the agent actually did. Not "it accessed the system", but the specific sequence of calls, the parameters, the results, the timestamps.
  • On whose instruction it did this. Which task, from which person or process, under which grant of permission.
  • Inside which boundaries it was operating. What was reachable at that moment, and where the wall stood that it could not cross.

This is exactly what a mature organization can already say about an employee: what they did, under what instruction, within what authority. The difference is that an employee leaves traces in email, tickets and systems, while an agent by default leaves nothing but a final result and a line in the application log.

Which brings me to the point worth taking from this news. Four things that have been filed under "nice to add in version two" are becoming mandatory parts of the system. Here is each one.

Audit trail: what the agent actually did

An application log answers the question "why did it crash". An agent audit trail answers the question "what was done and on what basis". These are different artifacts, and the second one almost never reconstructs from the first after the fact.

What belongs in it:

  • the originating task and its source: a user, a schedule, an external event;
  • every tool call: name, arguments, which system it went to, what came back;
  • branching decisions: why the agent chose this path and what data it had at that point;
  • model version, prompt version, version of the permission config;
  • everything the agent changed in the outside world, with identifiers of the records it touched.

The key property is that this record must be immutable and live separately from the agent system itself. If the agent writes into the same store it is allowed to edit, the entry carries little evidential weight. The same principle holds here as in ordinary data architecture, where the event log becomes the source of truth: history is written once and only appended to.

The second property gets forgotten: the trail has to be readable by a person who did not build the system. A lawyer, an auditor or an investigator gets an export, not your dashboard. If that export cannot reconstruct an evening of work, the trail exists on paper only.

Permissions and scope: what the agent can reach at all

The most common setup I see in pilots is an agent running under a service account with broad rights, because that is the fastest way to get a demo going. Then the demo becomes production and the account stays.

For accountability this is the worst available option. If the agent could technically reach the entire customer database, then investigating any incident means proving a negative: that it never went there. Proving a negative is expensive and rarely convincing.

The right framing is different: what is the minimum set of rights this task needs, and for how long. None of this differs from how least privilege works for people, except that agents add two wrinkles.

First, an agent moves faster than a person and will use an excessive right before anyone notices. Second, an agent has no instinct telling it that a particular folder is better left alone. It uses everything it was given if that helps finish the task.

So agent permissions get described by task rather than by role. A separate identity per class of task, token lifetime bounded by the task, data scope narrowed to the segment that is actually needed. And that permission config should be a versioned artifact, not a setting someone changed by hand in a console six months ago.

Sandboxing: where the wall runs

The trail records, the permissions constrain, but neither of them stops an agent at the moment it does something destructive. That is what isolation is for.

In practice it means the agent executes in an environment from which it physically cannot reach what it has no business reaching. Code runs in a disposable container with network access limited to an allowlist. File work happens in a dedicated space rather than a shared filesystem. Calls to external systems go through a layer of your own that decides whether to let each one through.

Isolation has a side effect more valuable than isolation itself: the boundary becomes the natural place to log and to enforce policy. If every outbound action passes through one layer, you get a complete trail and a single control point for free, instead of assembling them out of ten separate integrations.

Mature agent tooling is moving the same way. Execution isolation is becoming a standard part of the platform rather than something you bolt on, which I wrote about when looking at long tasks and sandboxing in agent SDKs.

Human approval on irreversible operations

The fourth element is the simplest technically and the hardest organizationally. You have to split agent operations into reversible and irreversible ones, and require explicit human approval for the irreversible ones.

The test is simple: can you undo the result within reasonable time using your own hands. Drafting an email is reversible. Sending it to a customer is not. Preparing a payment order is reversible. Executing the payment is not. Assembling an export is reversible. Shipping it outside is not.

The difficulty is making the approval mean something. If an agent asks a person to click "ok" a hundred times a day, the person clicks without looking, and on the hundred and first click approves something they never read. That loop exists on paper and will not hold up in court. On the balance between full autonomy and full manual control I have a separate note on why the higher the risk, the more a human belongs in the loop.

The version that works looks like this: approvals are rare, they cover genuinely irreversible actions, the person is shown the substance of the decision rather than the fact of a request, and their answer lands in the same audit trail as a distinct event with a name and a timestamp.

Why building it now is the cheaper option

A regulatory frame arrives on its own schedule, and it tends to arrive with a deadline attached. Retrofitting an audit trail onto a running agent system means rebuilding how it talks to every external system, because the trail has to sit on that boundary. Narrowing permissions after the fact means discovering which rights are load-bearing by breaking things in production. Adding isolation later means changing the execution model.

Built into the design, all four are a modest tax on the first release. Added afterwards, they are a project with its own budget, its own risk and its own outage. The cost difference is not a matter of opinion, it is a matter of how many existing integrations have to be touched.

There is a second reason that has nothing to do with regulators. The same four elements are what you need the day an agent does something expensive and someone inside the company asks what happened. That day comes earlier than any deadline from a regulator.

Where to start

  1. List the agents already running. Including the one an enthusiast in a department set up on their own. Without that list, the accountability conversation has no subject.
  2. Answer the three questions above for each one. What it did yesterday, on whose instruction, within what boundaries. Wherever you have no answer, you have a gap in provability.
  3. Create an audit trail separate from application logs. Immutable, carrying identifiers of the objects touched, readable by an outsider.
  4. Give every agent its own identity. One shared service account removes any ability to work out who did what.
  5. Narrow permissions to the task and put a lifetime on access. Start with the agents that reach customer and financial data.
  6. Route every outbound action through a single layer. That gives you the trail and the policy in one place.
  7. Classify operations as reversible or irreversible and gate the second group. Keep the number of approvals low, or they lose their meaning.
  8. Write down who inside the company answers for an agent's behavior. This is an organizational decision, and it has to be made before any rule forces it.

Honest caveats

  • I do not know the contents of the inquiry. Little is public, and the conclusions here are my own. I am working through the engineering consequence of the question being asked, not predicting the outcome.
  • This is not legal advice. I am describing what has to exist in the system to answer "what did the agent do". How that maps onto specific statutes and contracts is for lawyers, and they are worth involving in parallel with the engineering work.
  • All four elements cost money and slow the launch down. That is the honest price. The argument for paying it up front is that fitting an audit trail and permission boundaries onto a live agent system costs more, and when you actually need them there will be no time to build.
  • Full provability is out of reach. The model stays probabilistic, and you will never fully explain why it chose a particular path. What becomes provable is actions and boundaries, not the reasoning, and in most scenarios the former is what anyone asking actually needs.

In short

  • The FTC is asking whether existing law is enough to hold companies accountable for what autonomous systems do. The legal frame will be contested and slow.
  • The engineering requirement is already clear: be able to prove what the agent did, on whose instruction, and inside what boundaries.
  • Four elements stop being architectural improvements: audit trail, permissions and scope, execution isolation, human approval on irreversible operations.
  • Designing them in costs a fraction of retrofitting them, and when the need arrives there will be no time left to build.
  • The same four are what you need the day an agent does something expensive and someone internally asks what happened, which comes earlier than any regulatory deadline.

If you already have agents running and you are not confident you could reconstruct what they did last week, that is a good reason to look at the setup before someone outside asks the same question. Discuss your case through the form on the main page.

Back to all posts
Contact

If this resonated, write to me. I reply personally.

WhatsApp