Writing

Let the Agent Propose the Automation, Not Enable It

Let the Agent Propose the Automation, Not Enable It - abstract illustration

The feature request was one sentence: let the AI set up automations itself. I've been running a scheduled agent fleet for a long time and every job in it was authored by hand, so the appeal was obvious. I wrote down a safe subset of actions the agent would be allowed to create, started building, and discovered in the planning for phase two that the subset I'd written down was not safe.

The interesting part isn't the feature. It's how many separate things had to be true before it was survivable.

Start with what an automation actually is

An automation in my scheduler is a stored record with a trigger and an action. The available actions include running a shell command and running a Claude Code prompt.

So a tool that creates automations converts any agent that can call it into an agent that can run arbitrary commands on this machine, on a schedule, forever, whether or not it holds any other grant. That's the sentence I kept coming back to. Not "an agent that can do a bit more," but a full escalation with a timer attached, and one that outlives the conversation that created it.

That last clause is what separates this from every other tool in the catalog. If a model does something wrong with a file tool, the damage is bounded by the conversation. A wrong automation keeps being wrong at 3am next March.

The trap in reusing an existing grant

My first instinct was to gate automation creation behind the existing scheduler grant. That's the domain the agent-queue tools already live in, and it reads naturally.

It's also the exact wrong answer, for a reason specific to my own permission model. Simple mode, which is how a non-technical user gets an API key by ticking boxes like "read my email," emits scheduler read and write grants unconditionally, with no checkbox of their own, because the agent-queue tools need them for automations to work at all.

Fold automation creation into scheduler write and every Simple-mode key ever issued gains the ability to schedule arbitrary shell commands. Someone who ticked one box about email now holds command execution on a timer.

There was a second, quieter reason too. The scheduler domain's resource scoping is permanently by-name over a fixed set of queue item types. An automation is an id-keyed record with a name a user can change. It doesn't fit the scoping model that domain already committed to. Two different reasons, same conclusion: this needs its own grant domain, off by default for every existing key.

The safe subset that wasn't

Having ruled out the two obviously dangerous actions, I wrote down that the agent could create the other two: queue an item for an agent, or run an AI prompt. Running a prompt seemed clearly bounded. It isn't, for two independent reasons, and either alone is an escalation.

The prompt action names an API key by id. Any key, including the owner key. The tool loop that runs the prompt authorizes every call it dispatches through that named key. An agent holding only a narrow automation grant could create a job whose loop runs with the owner key's full grants. It doesn't need to escalate its own key; it just names a better one.

The prompt action's output file is validated against the wrong set of roots. The validation used all licensed and enabled filesystem roots, not the calling key's own write grants. Same shape as a bug I'd already found and fixed in the image generation tool, inherited here through a different code path.

So the actual permitted subset is: queue an item for an agent, plus run an AI prompt with no tools and no output file. Which is a much smaller feature than the one I set out to build, and is the honest version of it.

I did consider narrowing the output file to the calling key's roots at creation time. I rejected it, and the reason generalizes: the check does not survive revocation. The job outlives the key that authorized it. Validating against a key's grants at creation gives you a job that keeps running with permissions the key no longer has, six months after someone revoked it. A restriction that's checked once, at creation time, against mutable state is not a restriction.

Restrict the triggers too, not just the actions

I nearly stopped at the action list. The triggers turned out to matter just as much.

Event triggers fire on new mail or new feed items. Letting an agent choose one means attacker-controlled content becomes a trigger the agent selected. Watched-folder triggers hand the agent a filesystem watcher over a root. Neither is acceptable when the agent picked the trigger, even though both are perfectly fine when a human did.

Agent-created automations get simple, cron, and interval triggers only. Time-based, content-independent, and boring.

Created disabled, which is the whole design

Every automation an agent creates arrives disabled, with a notification carrying a Review action. A human enables it or it never runs.

This is the difference between a suggestion and a fait accompli, and it's the single decision that makes the rest of the feature defensible. An automation outlives its conversation and the cost of a wrong one is unbounded and delayed, which is exactly the profile where you want a human in the loop once, at creation, rather than never.

Two consequences fell out of it that I liked more than I expected. A pending automation doesn't count against the free tier's job cap, since it isn't running. And the tool's response has to say plainly that the automation was created disabled, so the model relays that to the user instead of reporting a bare success and leaving them to believe it's live.

The staging decision also determined the tool surface. There's a create tool and a list tool, and nothing else. No set-enabled tool, because that would let the agent approve its own automation and defeat the entire point. No update or delete, because an agent that can delete automations can silently switch off the ones you rely on. The list tool discloses summary only: name, id, trigger summary, action kind, enabled state, and origin. Prompt and command text are withheld, since they carry credentials and paths that were never meant for an agent to read.

An automation may not create automations

A scheduled AI tool loop is categorically denied the automation domain at dispatch time, regardless of what its key is granted. An automation that creates automations is the runaway shape, and this feature's real use case is an interactive conversation with a person present.

There's an honest limit here that I wrote into the docs rather than leaving implied. The denial covers the in-process tool loop. A job that spawns the Claude CLI reaches this server as an ordinary HTTP client with whatever key its config holds, and there is no dispatch-source signal that distinguishes it from a human session. For that path, the key's grants are the entire control. I'd rather have that written down as a known boundary than have someone infer a guarantee that doesn't exist.

Where the rule lives matters

The restrictions on what an agent may create are enforced in the draft validation, keyed on the record's origin, not in the tool handler.

I built a single rule set as its own phase precisely so there would be one place these decisions live. A check in the tool handler would bind that tool and nothing else, and would not bind a future caller written by someone who reasonably assumed the validation layer was doing its job. This is the same lesson as filtering a tool out of the menu not being enforcement: the guarantee belongs at the chokepoint everything passes through, not at the door you happened to be standing at.

What I'd take from this

If you're giving an agent the ability to create something that persists, work through four questions before the API design: what grant does it need and does that grant already mean something else, what subset of the thing may it create, what may trigger it, and does a human see it before it takes effect.

The fourth question is the cheapest and does the most. Almost everything on my list got safer the moment the answer to "does this run without a person seeing it" became no.

Building something like this?

This is the kind of work I do for clients. Tell me what you're building and I'll give you a straight read on the approach.

Book a 30-minute call