Kubernetes for AI agents: when always-on agents need a control plane
OpenClaw Enterprise is being built as an open-source control plane for always-on AI agents, with OpenAI, NVIDIA and Red Hat involved. What the AgentOps layer actually solves, who needs it today, who does not, and the signals that tell you it is time.
OpenClaw Enterprise is being developed as an open-source control plane for managing always-on AI agents. According to the project, OpenAI, NVIDIA and Red Hat are involved, and OpenAI and Red Hat are already testing the technology inside their own infrastructure. The architectural idea reads almost like Kubernetes: a single control plane manages a fleet of agents running across different environments. The control plane itself is deployed on top of Kubernetes.
For an owner or a CTO, the specific system matters less than the signal. A separate infrastructure layer for agents is taking shape, and it already has a name: AgentOps. It is arriving the same way container orchestrators arrived.
A story that repeats for the third time
The beginning is always the same. Someone starts a process by hand, then there are ten of them, and a layer of homemade glue grows around them: a restart script, a spreadsheet of who holds which token, a chat where people agree on who shuts everything down tonight. For a while this beats any platform, because it costs nothing and takes an evening to build.
Then quantity turns into quality. The glue stops holding, nobody except its author understands it any more, and an orchestration layer appears that takes over lifecycle, scheduling, credentials and observability. What was one engineer's personal practice becomes company infrastructure. That happened with virtualization. It happened with containers: I wrote about the moment orchestration became the operational normal in the piece on Kubernetes 1.0. It happened with data, where hand-rolled cron turned into pipeline schedulers. The same story is running with agents now, and it runs faster, because the engineering patterns are already known and taken straight from the Kubernetes world.
What the layer actually solves
A control plane for agents is best read as a list of jobs that someone does anyway. The only open question is whether a person does them by hand or a system does them.
- Lifecycle and scheduling. When an agent starts, how long it lives, what happens when it crashes, how many instances run at once. For an always-on agent this is a state to be watched rather than a one-off launch.
- Granting and revoking access. Agents reach into mail, CRM, repositories, payment gateways. While there are three of them, keys in environment variables are fine. At forty, "who has access to what today" becomes a security question, and revoking one agent's access should take a minute rather than an evening.
- Spend limits. An agent that runs all the time spends tokens all the time. A buggy loop hammering the model shows up on the month-end invoice unless a hard ceiling lives in the system. I covered how to set token budgets in the note on controlling LLM costs.
- Isolation. What an agent can see, where it can write, what happens to the rest if one goes off the rails. Exactly the question that produced namespaces and quotas in the container world.
- Observability. What the agent did, why it decided that, where it stalled. One agent's logs are read by its author. Forty agents need a single journal, otherwise incident review turns into archaeology.
- One stop button. The most underrated item. When something goes wrong, you need one place where everything halts, and one person who knows where that place is.
These are the same jobs a container orchestrator does for processes and a scheduler does for data pipelines. The difference is that an agent makes decisions and acts in external systems, so the cost of a failure is higher than a crashed worker.
Who it is early for, and who is already there
It is early for anyone running two scheduled scripts. Cron, a task queue and decent logging cover that volume completely, and they are cheaper and clearer for the team. A platform here adds a layer that somebody has to operate, and it eats the time better spent making the agents useful.
It is time when the signals line up. They are measurable, and you can check your own situation in half an hour:
- The count runs into the dozens. There are more agents than people who remember them. While the whole list fits in one engineer's head, a separate layer is not needed yet.
- Someone is on agent duty. A team member spends hours every week on restarts, crash triage and handing out credentials. That is operations already, just manual.
- Revoking access takes more than an hour. If "which agent holds the production database key" has no answer within a minute, access management is already broken.
- Spend is unpredictable. The inference bill jumps, and the jump cannot always be explained after the fact.
- Incident review takes longer than the incident. Logs live in several places, and the agent's decision is reconstructed from indirect evidence.
- Agents overlap on data and rights. Two agents reach into the same system with different intentions, and the boundary between them rests on a verbal agreement between people.
Three signals out of six justify putting the layer into next quarter's plan. Five justify dealing with it now, because it only gets more expensive.
The classic mistake: a platform before the load
The most expensive way to burn a year is to install an orchestration platform before there is anything to orchestrate.
The pattern is familiar. A team reads about the new layer, deploys it, spends a quarter on integration, and ends up with a system that manages three agents. The platform needs maintenance, upgrades and its own on-call. The payoff it was bought for would have arrived at forty agents, but forty never happened, because the capacity went into the platform. I have watched this play out with Kubernetes at companies that were fine with a couple of virtual machines, and wrote it up in the note on when orchestration is overkill.
The cause is always the same: the platform is sized for the expected scale rather than the current one. Expected scale is a hypothesis. You pay for it immediately and collect only if the hypothesis holds.
Where to start if the signals line up
- Build an agent register. What it is, whose it is, what it does, what it can reach, what it spends. One spreadsheet, one day of work. It routinely exposes half the problems before any platform is involved.
- Fix access before orchestration. Keys issued from a secret store, short lifetimes, revocation in a single action. That value stays with you whether a platform ever shows up or not.
- Put a spend ceiling on every agent. A hard cap on tokens and calls that stops the agent rather than emailing a warning.
- Pull the journals into one place. A single journal shows the real picture of runs, failures and spend instead of impressions, and it also tells you whether orchestration is due.
- Build one stop button. Even if it is a script that kills everything at once, it has to exist and be tested before it is needed.
- Choose the platform last. Once the first five steps are done you hold real requirements, and the choice stops being a bet on fashion.
Honest caveats
- I say little about the specific project on purpose. What is known: it is being built as an open-source control plane, OpenAI, NVIDIA and Red Hat are involved, and two of them are testing it internally. Drawing conclusions about maturity, functionality or fit for your environment from that is premature. The category of solutions matters more here than the name.
- The category is young. A layer this new changes its interfaces and deployment model faster than a company can wire it into its processes. If you adopt now, budget for a migration and keep your agents loosely coupled to the platform.
- Orchestration does not fix a bad agent. An agent that makes wrong decisions will make them more often and at greater volume under a control plane. The platform solves operations; the quality of the decision stays with you. On the gap between a working demo and something you can trust in production, there is a separate note: from demo to platform.
- A home-grown control plane is still a layer. The temptation to write your own is strong and usually ends with the team maintaining a platform instead of using one. When the jobs are standard, a standard solution comes out cheaper.
The short version
- OpenClaw Enterprise is being built as an open-source control plane for always-on AI agents, with OpenAI, NVIDIA and Red Hat involved; architecturally it mirrors Kubernetes and runs on top of it.
- The meaningful part is that an AgentOps layer is forming at all: the industry is walking the road containers and data pipelines already walked.
- The layer covers six things: lifecycle, credentials, spend limits, isolation, observability and one stop button. Those jobs exist before any platform, they are just done by hand.
- It pays off for teams running dozens of always-on agents with external consequences and a visible inference bill. Two scheduled scripts are well served by cron.
- The common mistake is installing the platform before the load. Start with the agent register, access, limits and journals, and pick the platform last.
If you are already running more than a dozen always-on agents and weighing whether to build a layer under them, discuss your case through the form on the home page: we will walk your situation through the signals above, and the first conversation commits you to nothing.