mksim.pro
Back to all posts
IT 12 min read

The coding agent gets its own environment: rights, review and the bill

OpenAI added reusable cloud development environments to Codex. A coding agent stops being autocomplete in an editor and becomes a participant in the pipeline with its own access. What changes in credentials, secrets, repository history, and who answers for the code it writes.

OpenAI has expanded Codex considerably. According to the company, the agent now gets reusable cloud development environments: the environment persists in the cloud and is reachable from different devices. Alongside it came a refreshed CLI, voice control and new code review tooling.

From the outside this reads as a feature list. What actually moves is the unit of development. While the agent lived inside an editor session, it was a function of someone's workstation: close the laptop and the context goes with it. A persistent environment turns the agent into a long-lived participant in the pipeline. That participant has its own state, its own dependencies, its own keys and its own history of actions. Your perimeter has gained another actor that takes actions on its own.

What a persistent environment actually changes

Strip away the marketing and persistent environments give you three effects.

  • Continuity. A task outlives the laptop. The agent can run a long chain: run tests, fix, run again, build, roll back, try another way. That chain used to be bounded by the length of your working day.
  • Portability. The same environment is reachable from another device. An engineer starts a task from a workstation and checks the result from a phone.
  • Accumulated state. The environment remembers installed packages, caches, configs, environment variables and whatever tokens were placed in it.

The third effect is the one people undervalue, and it is the most interesting. Accumulated state is exactly what infrastructure people call environment drift. The build passes inside the agent's environment because somebody hand-installed a library there six months ago, and it fails in clean CI. The classic disease of developer machines moves to the cloud and takes up permanent residence there. The cure is the same one it always was: the environment is described by a file in the repository, and the environment gets recreated from that description instead of being patched by hand.

I wrote about how a coding agent shifts the assembly point of development back when Copilot made the move: the coding agent and the new assembly point. That piece was about the worker. This one adds infrastructure underneath the worker, and with it a set of questions teams usually ask too late.

Whose credentials does the agent use

This is the first question, and it is almost always deferred. In practice there are three answers.

A developer's personal token. The fastest route and the ugliest one downstream. In history and in logs, the agent's actions are indistinguishable from the human's. An incident review turns into guesswork: did the engineer decide to rewrite that module, or did the agent decide it on its own. Then the person leaves, their token is revoked, and half of an automation nobody documented stops working with it.

A service account. Better: a separate identity, separate rights, a separate trail in logs. The known failure mode of service accounts is that they accumulate rights. Someone grants production repository access for one urgent task, and it stays there forever. That is the same mechanic I described in the piece on the lifecycle of access: the dangerous part is not the moment of granting, it is that nobody ever shrinks what was granted.

Short-lived rights scoped to a task. The agent gets access to the repositories it needs for the duration of the run, and the rights expire with the task. This is the right answer and the most expensive one to set up. Start there if the agent will come anywhere near a production repository.

Then there are the secrets inside the environment itself. A persistent environment holds whatever you put into it: environment files, staging keys, database credentials, tokens for private package registries. That material used to sit on a laptop and fall under endpoint policy. Now it sits with a vendor, in an environment that outlives the session. This is not an argument against the tool. It is an argument for knowing the exact inventory of what lives in there, and for keeping only material you are ready to rotate with one command. Keys to models and services became a perimeter of their own a while ago, and I covered that separately.

There is a third layer people remember last: outbound access. The agent's environment talks to the network, to package registries, to documentation, to third-party APIs. An agent allowed to reach everything will one day pull in a package with a convincing lookalike name. Network boundaries deserve to be written down as explicitly as repository rights.

Repository history stops being human

When the agent commits under the same name as people, you lose a distinction you will need in six months. The question "who wrote this and why" becomes "did a human write this while holding the context, or did a machine land on it during its third self-correction loop". The answer has to live in the data, not in someone's memory.

The minimum marking looks like this:

  • a separate author or committer for the agent, so history can be filtered by it;
  • a link in every commit and pull request to the task and the original instruction the agent worked from;
  • direct pushes to the main branch closed for the agent identity, so everything arrives as a pull request.

That is half a day of configuration and it pays for itself at the first incident review. It also gives you honest numbers: how much code in the project actually came from the agent, and what share of it survived to production.

Review becomes the bottleneck

This is the biggest shift in process, and it goes largely undiscussed because everyone is busy discussing the speed of writing code.

Writing code is getting cheap fast. Reviewing it is not getting cheaper at all. An engineer can read a bounded amount of change carefully per day, and that is physiology rather than a process setting. An agent that worked through the night hands you several tidy pull requests in the morning, each of them convincing on the surface. Three things can happen next.

  1. Review degrades into a rubber stamp. The most dangerous outcome. The diffs are large, they look reasonable, tests are green, approval takes a minute. Responsibility is formally distributed and practically held by nobody.
  2. Review becomes a queue. Pull requests pile up, go stale, conflict with each other, get reworked. The agent produced a lot of work and the team received a lot of work in progress.
  3. The team changes how it reviews. The only outcome that holds.

What changing it means in practice. Line-by-line reading stops being the primary instrument and moves to where machines are useless. Machines check properties: tests, types, API contracts, migrations, performance budgets, dependency scanners, linters. Humans check what property checks never catch:

  • intent, whether the problem that was posed is the problem that got solved;
  • module boundaries and ownership, whether logic leaked somewhere it does not belong;
  • domain and security, what happens to these rights, this data and this user;
  • reversibility, what it takes to roll this back after it ships.

From that follows a plain organizational requirement: the agent must produce small changes. A large autonomous pull request is unreadable by construction, and the discipline of small changes stops being a matter of taste. Shipping behind a flag helps for the same reason, because it detaches risk from the moment of deployment.

What it costs and who counts it

The bill for a long autonomous run has three parts: model calls, environment runtime, and the CI load the agent generates with its own test runs. The third usually surfaces first, because the build bill grows before anyone looks at the model bill.

Measure it as one number: cost per accepted change. Runs on their own tell you nothing. What matters is what you paid for the pull requests that reached the main branch, and what share of the work was discarded. That requires attributing spend to a task and a team, and putting a ceiling on each task: time, attempts, budget. An agent allowed to spin indefinitely will spin indefinitely, and will report its effort beautifully.

What to set up before the agent touches a production repository

  1. A separate identity for the agent. Its own account, its own trail in logs, its own filter on history.
  2. Rights that expire by themselves. Access granted to specific repositories for the length of a task, gone when the task ends.
  3. An inventory of environment secrets. A list of what sits inside, an owner per secret, and a rotation procedure somebody has actually tested.
  4. Network boundaries. An explicit list of where the environment may reach, and a default of no for everything else.
  5. No direct pushes. Every change arrives as a pull request with mandatory checks.
  6. Green CI as the price of entry. Tests, types, lint, dependency scanning and migrations run before human eyes, not after them.
  7. A ceiling per task. Limits on time, attempts and budget, after which the agent stops and hands over what it has.
  8. A readable action log. What the agent ran, what it changed, which commands it executed. You will need it for a post-mortem rather than a report.
  9. A rollback procedure. Who removes a change that passed review and turned out wrong anyway.
  10. A named owner. Every agent task has a human attached to it. They set it, they review it, they roll it back.

The first six items are a day or two of work. The rest settles in as you go, but the decision about the owner has to be made before the first run, or responsibility will dissolve at exactly the moment you need it.

Honest caveats

  • Long autonomous runs work where the result is checkable by tests. Dependency upgrades, migration to a new API, refactoring under coverage, adding types, generating glue code, repetitive edits across a project: here the agent has an honest success criterion and can grind until it converges. Where the quality criterion is subjective, such as an architectural decision, the wording in an interface, or a trade-off between speed and simplicity, long autonomy does poorly. The agent will deliver a confident result and you will have no cheap way to check it, and the savings go straight into argument on the review.

  • The news describes one vendor's tool. I am not discussing pricing, limits or dates, because the source announcement does not carry them and the field moves faster than any such write-up stays valid. The direction is what matters: agents are acquiring infrastructure of their own, and everyone playing this game is doing it. On tooling maturity for long-running tasks I wrote separately.

  • An environment does not repair a process. If review in your team is already ceremonial, tests are thin and module owners are unnamed, a persistent environment will simply speed up the production of unverified code. Tooling amplifies what is already there, disorder included.

  • The deep identity and secrets work belongs to specialists. I map the risk, assemble the loop and put names against ownership; configuring the actual access policies is security engineering. The adjacent story of how assistants with system access create a new attack surface is here.

In short

  • A persistent environment turns a coding agent from a function of the editor into a long-lived participant in the pipeline, with its own state, its own keys and its own history of actions.
  • The first question is whose identity it uses. A personal token breaks the audit trail, a service account accumulates rights, and the working answer is short-lived rights scoped to a task.
  • Environment secrets and outbound network access need an inventory before the first run, not after the first incident.
  • The bottleneck moves to review rather than to writing code. The fix is property checks on machines, small changes, and human reading reserved for what machines cannot see.
  • Autonomy pays off where success is verified by tests and stalls where the quality criterion is subjective.

If you are deciding right now whether to let an agent into a production repository, the cheapest moment to settle rights, secrets and review is before the first run rather than after the first rollback. Discuss your case through the form on the home page, and the first conversation commits you to nothing.

Back to all posts
Contact

If this resonated, write to me. I reply personally.

WhatsApp