A heartbeat is a small message a system sends on a fixed schedule to say it is still running. It carries almost no information. Its value lies entirely in what its absence means.
The problem it solves
Most monitoring watches for bad events: an error, a crash, a failed request. That misses the failure where nothing happens at all. A process that hangs produces no errors. A loop that stops looping logs nothing. An alert that has silently broken looks identical to an alert whose condition has not been met. Silence is ambiguous, and a heartbeat removes the ambiguity.
With a heartbeat, silence has a meaning. If the message is due every minute and several minutes pass without one, something is wrong, whatever the error log says.
How it works
There are two halves. The sender emits a signal at a fixed interval, usually with a timestamp. The watcher expects it and raises an alarm when it is overdue. The watcher must be separate from the sender. A process cannot report its own death.
On a streaming connection the same idea appears as a ping. If the line carries no market data for a while, the two ends exchange a small message to confirm the connection is still open. Without it, a dead connection and a quiet market are indistinguishable, the problem described in a feed that stops versus a market that is quiet. Polling versus streaming covers the two delivery styles.
What it should measure
A heartbeat proves only what it is attached to. If it is sent by a timer that runs independently of the real work, it will keep beating while the real work is stuck. The useful design emits the beat from inside the loop that matters, after a pass completes. Then a beat means a full cycle just finished, which is the thing worth knowing.
A richer beat carries a little more: when the last successful pass ended and how long it took. That turns a yes-or-no signal into a reading of health. A pass that used to take seconds and now takes minutes is a warning before it becomes an outage.
Choosing the interval
The interval and the tolerance set how fast a failure is noticed. Too tight and normal jitter raises false alarms, which teach people to ignore the alarm. Too loose and the system is down for a long time before anyone knows. A common approach is to allow a few missed beats before alarming, sized to how slow the work can legitimately run.
Why a trader should care
Anyone relying on an alert is relying on a negative. No alert is taken to mean the condition has not happened. That inference is only safe if something confirms the alerting system is alive. How do you know an alert is still working takes up that question from the user’s side.
The same applies to any automated routine. A resilient trading loop treats failing silently as the outcome to design against, and a heartbeat is the simplest tool for it.
The reader’s version
A published age on every reading works as a heartbeat the reader can see. If the age keeps resetting, the pipeline is alive. If it climbs without limit, something has stopped. On the MCP & API, every reading carries as_of and age_seconds, and a stopped feed answers with an error and a retry time.