Alerts
Nine monitors that ship on, and a delivery path that tells you when one changes — with the limits of an in-process notifier stated rather than hidden.
/monitors has always computed nine checks. Until now it only ever showed them, which made them a
dashboard rather than an alert: they existed for the length of one render, and the moment you
needed them was the moment nobody was looking.
This page is the other half — the part that speaks first.
Turn it on
One variable. Set it to a Slack incoming webhook, a Discord webhook, or any HTTPS endpoint you control:
EVESTACK_ALERT_WEBHOOK_URL=https://hooks.slack.com/services/T0/B0/xxxxThe payload shape is picked from the hostname — hooks.slack.com gets {"text": …}, a Discord
webhook path gets {"content": …}, anything else gets a structured JSON body. Override it with
EVESTACK_ALERT_WEBHOOK_FORMAT=slack|discord|webhook if you post through a proxy.
Which file it goes in
Worth one table, because the variable has to reach the dashboard's process and the three ways of running it read three different files:
| How you run the dashboard | Put it in | How it arrives |
|---|---|---|
A project from npm create evestack | .env.local | The generated docker-compose.yml gives its dashboard service env_file: .env.local, so every name in that file is set inside the container. .env.example lists all six, commented out. |
This repository's own docker-compose.yml | .env beside it | That file has no env_file:, so a name reaches the container only by being named in the dashboard service's environment: block. All six are. |
npm run dev in packages/dashboard | packages/dashboard/.env.local | Next reads it directly. |
Until 2026-08-19 the repository's own docker-compose.yml named none of the six, and that is
the failure mode this page is least able to help with. Setting EVESTACK_ALERT_WEBHOOK_URL in
the .env beside it did nothing at all: Compose read the file, matched the name against no
${…} anywhere, and dropped it — while /monitors went on saying no delivery target was
configured, to someone who had just configured one. Scaffolded projects were never affected,
because their compose has env_file: .env.local — but their .env.example did not mention
these variables either, so there was nowhere a user could have discovered them.
Restart the dashboard and it says which way it went, once, at boot:
[evestack:alerts] delivering transitions — 1 sink(s), every 60sUnset, it says the other thing, and /monitors says it too rather than leaving you to assume.
That is deliberate: a page showing nine checks under the word firing, with nothing stated about
delivery, invites exactly one conclusion.
Press Send a test on /monitors to prove the channel. It sends one obviously-synthetic message
and does not touch any monitor's state — a test that consumed a real transition could swallow the
alert it was meant to prove.
What gets sent, and what does not
A message goes out when a monitor changes, never merely because it is still bad. An integration that re-sends every minute is muted within a day, at which point it is worth less than nothing: it has trained you to ignore the channel the real one will arrive in.
| From | To | Sent |
|---|---|---|
| never seen | firing | yes — a fresh install that is already broken is worth one first message |
| never seen | ok / not checked | no — nine "all fine" messages on first boot is how a channel gets muted |
| ok | firing | yes |
| not checked | firing | yes |
| firing | ok | yes, the all-clear |
| firing | not checked | yes, as not checked — see below |
| ok | not checked | only for page severity |
| not checked | ok | no |
firing → not checked is never reported as resolved. The monitor stopped being answerable
while it was firing, so the problem is un-observed, not over. Sending "resolved" there would stand
an operator down in the middle of a live incident, and it is the single most dangerous thing this
path could do.
ok → not checked pages only for page-severity checks. For the two that would wake someone,
losing the ability to evaluate is the incident. Below that it is silent, because an unmounted
Docker socket is a configuration choice, not news.
To repeat a still-firing alert, set EVESTACK_ALERT_RENOTIFY_MINUTES=60. It is off by default.
The spend monitor, and which cap it uses
EVESTACK_ALERT_DAILY_SPEND_USD is the install-wide daily spend the alert is judged against.
With it unset, the alert falls back to @evestack/budget's EVESTACK_BUDGET_DAILY_USD (default
10) so that it works out of the box — and it says so in the alert text, because the two are
not the same measure. The budget cap is per principal; this alert sums every priced turn on the
install. On the single-user install that self-hosting usually means they are the same number; with
two users the install crosses $10 while neither person is near their own limit.
This monitor spent its entire existence stuck at not checked. It read
EVESTACK_DAILY_BUDGET_USD — the same four words in a different order from the variable that
exists — so on every install ever made it reported "no cap is configured" while @evestack/budget
enforced its $10/day default a process away. A check that cannot fire is worse than a missing
one, because the page counts it among the ones that passed.
Set it to false to switch spend alerting off without touching your budget caps.
Signing, for the generic sink
Slack and Discord authenticate by secret URL. For your own endpoint, set a shared secret and every POST is signed:
EVESTACK_ALERT_WEBHOOK_SECRET=$(openssl rand -hex 32)x-evestack-alert-timestamp: 1786232997598
x-evestack-alert-signature: sha256=d0b4cec6bb0de…The signature is HMAC-SHA256(secret, "<timestamp>.<raw body>"). The timestamp is inside the
signed material on purpose — signing the body alone authenticates the content and nothing else, so
a captured POST replays forever. Reject anything older than your own tolerance.
Honest limits
-
If the dashboard is down, nothing is delivered, and nothing here can tell you so. That is a genuine limit of any in-process notifier, not an oversight. The payload carries
sentAtso a receiver that cares can alert on silence — the one check that has to live outside the thing it watches. -
The dashboard is an optional compose profile (
profiles: ["full", "dashboard"]).docker compose upwithout a profile starts Postgres and the agent and no notifier at all. -
It is deliberately not delivered through your agent's channel. evestack already has a path that reaches a human — the heartbeat — and it is the wrong one for this. Three of the nine monitors (
wedged,no_spans_while_active,turn_failure_rate) fire exactly when the agent is unwell, so routing them through the agent means the turn that wedges is the turn that would have told you. The dashboard is a separate process and stays up when the agent does not. -
Delivery is at-least-once, not exactly-once. A transition is only marked delivered once a sink returns 2xx, so a receiver that is down replays it on the next tick rather than losing it. A receiver that accepts the POST and then drops it is indistinguishable from one that delivered it — that is the correct place to give up.
-
Delivery is tracked per sink, so one broken channel does not spam the others. Each sink has its own record of what it has been told. A sink that is down is retried until it accepts; a sink that is healthy is told once and then left alone. The first version of this required every sink to accept before a transition counted as delivered — which sounds careful and meant that a single permanently-broken webhook re-sent the identical alert to every healthy channel on every tick, forever. That is the muted channel this whole page is about, arrived at by way of being careful about the other thing.
-
Two dashboards will not normally double-send, and the exception is named. A short lease row in
evestack.alert_leasegives one instance the tick; it expires rather than being released, so a crash mid-send is picked up by the other one interval later. Within a single dashboard, deliveries are serialised outright, so a timer tick and aPOST /api/alertscannot overlap.What the lease does not cover: it is granted on elapsed time and does not check who holds it, so if one instance's sends take longer than the whole interval — several sinks, each slow but healthy — the other can take over while the first is still posting, and one transition goes out twice. Configure a sink list whose total latency fits inside
EVESTACK_ALERT_INTERVAL_SECONDS, or run one dashboard. A duplicate page is the worst case here; nothing is lost.
What it writes
Three tables, all in the evestack schema, created on first delivery and never before — a feature
that is off costs nothing, including a table.
| Table | What it holds |
|---|---|
evestack.alert_state | one row per monitor: what was last seen, what was last delivered, and what each sink has acknowledged |
evestack.alert_deliveries | every POST, whether it worked, the status, and how long it took. Pruned at 30 days |
evestack.alert_lease | which instance is sending |
The two state columns are separate on purpose. The decision to send compares the live state against what was last delivered, so a webhook returning 500 does not consume the transition. If they were one column, a single failed POST would mark the alert handled and the page that started failing at 02:00 would never be mentioned again.
Stored URLs are truncated to their origin. Slack and Discord both put a working credential in the path, so the full URL is never written to a table the dashboard renders.
Reading it from somewhere else
GET /api/alerts returns all nine with their state, plus whether delivery is configured and when
it last ran. POST /api/alerts runs a tick now; POST /api/alerts?test=1 sends the test message.
All three take the dashboard's ordinary credentials, including HTTP Basic:
curl -u "$EVESTACK_AUTH_USER:$EVESTACK_AUTH_PASSWORD" http://localhost:4000/api/alerts