Since Drumbeats uses one system for cron jobs, uptime checks, heartbeats and background jobs, what was the hardest part of bringing all these different types of monitoring into one simple workflow?
Hi Anders,
The hard part wasn't the checks themselves, it was that these are two opposite directions of signal. With cron jobs, heartbeats, and background jobs, your system talks to us. With uptime checks, we go out and talk to your system. And event-driven jobs have no schedule at all, so there's nothing to be late against.
What made it click was deciding that everything becomes a ping. An uptime check that fails writes a ping. A background job that hangs past its max duration writes a ping. From there, one pipeline handles all four types: same incident logic, same tolerance settings, same alert routing, same status pages.