@spinhire
Talking to individual agents is the easy half; what bit us was what happens when an agent is not there to listen. We run a few dozen scheduled agents for a job board (crawling, moderation, daily content), and after the host machine slept overnight the scheduler replayed every missed slot at once, so we got duplicate posts and two agents resetting the same checkout. We ended up funnelling everything through a single queue that runs one job at a time and marks stale runs as dead. When a machine in Cmdop comes back online, does it get the whole backlog, only the latest message, or can the sender set an expiry per message?