Mail sync, feed sync, and message sync all did the same two things for every new item: save the item, then queue an automation event for it. Both writes worked. Both had tests. The order was wrong, and the order is the entire thing.
The gap between two writes
Persisting the item marks it seen. The next sync pass will skip it, permanently, because that's what seen means. Queueing the automation event is what actually makes something happen: a rule fires, an agent runs, an email gets filed.
Do the first one first and there's a window between them. If the process dies in that window, or the queue store isn't open, or the enqueue call throws and someone swallowed it, the item is durably marked seen with no record anywhere that an event was still owed for it. The next pass skips it. The event never fires. Nothing is logged, because from each subsystem's own perspective its own write succeeded.
The failure is silent by construction, and it's also unrecoverable, which is the part that makes it worth a rewrite rather than a retry. There's no state on disk you could inspect afterward to find out what was missed.
Reverse the order and the failure changes species
Queue the intent first, then commit seen-state. Now a crash in the gap leaves you with a queued event for an item that isn't yet marked seen. The next pass re-detects the item and may enqueue a second event.
That's a duplicate, and a duplicate is a completely different class of problem from a loss. A duplicate is visible, it's diagnosable, and it's usually idempotent or can be made so. A silent loss is none of those. This is the standard at-least-once versus at-most-once trade, and for anything that triggers work, at-least-once is nearly always the side to land on.
I already had this right in one subsystem. The watched-folders code had a documented "queue-row-before-ledger" ordering, written down with the reasoning attached. Mail and feeds had been built later, by a different path, and had quietly done the opposite. Having the correct pattern documented in the repo did not stop the incorrect pattern from being written three more times, which is its own lesson about how far documentation gets you.
A debt table, because the two writes aren't in one transaction
Simply reversing the order isn't available when the queue lives in a different store than the item. You can't put "insert into the queue database" and "insert into the mail database" in one transaction. So the ordering fix becomes a record-the-debt fix.
Each subsystem now writes, in the same transaction as its own seen-state commit, a row saying it still owes an agent-queue enqueue for this item. Mail got an agent_event_debt table. Feeds got an enqueue_pending column on a ledger table that already existed. Messages got its own debt table. At the start of the next sync pass, each one claims its outstanding debt rows and tries again.
The debt row and the seen row land together or neither lands. That's the property that makes the whole thing work, and it's available because they're in the same database, which the queue is not.
There's a detail in the messages version worth stealing. Its drainer re-resolves each claimed debt row against the live source and the current allowlist before handing it to the queue. A row can sit in debt across a restart, and in the meantime the underlying message can be deleted, or the handle it came from can drop out of the permitted scope. Delivering it then would be delivering something the current configuration says you shouldn't have. Debt is a promise to try again, not a promise that the answer is still the same.
The dashboard that said everything was fine
The worse half of this was the reporting, and I nearly missed it because I was focused on the write ordering.
When the agent queue store failed to open, on a full disk or a permissions problem, exactly one component status went degraded: the scheduler. Mail sync, feed sync, and message sync all kept reporting healthy, because each one's own item persistence was working perfectly. New mail was still being saved. Only the automation events were going nowhere.
So the observable behavior was: a green dashboard, mail arriving normally, and every automation silently dead. Someone relying on an automation would have no signal at all until they noticed the work wasn't getting done, which for a nightly job could be days.
The fix is small and I think it generalizes. A component's status has to reflect its dependencies, not just its own work. Now a failed queue store sets a reason string that gets folded into the status of every subsystem that hands it events. A running status becomes degraded, with text naming the actual consequence: new items are still being saved, but automation events are being held and will be delivered after a restart.
Two rules in that folding kept it honest. A disabled component stays disabled, since something that's off has nothing to hand the queue anyway. And an already-degraded component gets the new reason appended rather than replacing what was there, so one failure can't hide another. That second one sounds fussy until the first time you watch two problems overlap and only see the newer one.
What the tests had to pin
The tests for this needed to check the durable state directly rather than the behavior, which is a slightly unusual shape. The interesting assertions read the debt row out of the store:
Persisting a new item with the owes-event flag creates a debt row in the same transaction. Persisting with the default creates none. Re-ingesting the same item doesn't create a second debt row. A nil queue records debt and a later drain delivers it exactly once. A debt row whose source item no longer exists gets dropped at drain time rather than delivered. A debt row whose handle has left the permitted scope gets dropped, not delivered.
None of those are testable through the public surface. You have to look at the table. Any test written purely against "did the automation fire" would pass on both the broken and the fixed version whenever the queue happened to be open, which it is in every ordinary test run. I've made this mistake before, and it's related to something I wrote about testing parsers against real data: a test that only exercises the healthy path is measuring your fixtures, not your code.
What's still open, and named
Two of the three debt tables have no bounded-growth valve. Messages prunes its oldest rows beyond a keep count; mail and feeds don't. If the queue store stayed broken for a long time, those tables would grow without limit. It's filed as low priority with the fix described, and it's in the punch list rather than in my head.
I'd rather ship with a known, written-down gap than with an unknown one. The whole reason this batch of problems was findable at all is that the previous fixes had their reasoning written down where the next person could read it. That's also the reason it stings that mail and feeds got written the wrong way anyway.
The rule, stated plainly
If two durable writes have to happen and you can't make them atomic, commit the one that's recoverable first. Marking work as done is never the recoverable one.
And when a shared dependency fails, make every component that depends on it say so. A status that only reports on a component's own work will tell you everything is fine right up until someone asks why nothing has happened for a week.