Skip to content
▚ evestack docs

The dashboard

The open replacement for Agent Runs — and it drives the agent, not just watches it.

Observe

  • Sessions — every run on your machine, with title, status, and trigger
  • Run tree — turns and subagents nested by $eve.parent/$eve.root, with duration, input and output tokens, cached-read tokens, and tool count per turn
  • Computed cost — see Architecture for why it's computed rather than read from a span
  • Integrations — the live Composio catalog, connected accounts, one-click connect
  • Monitors — latency percentiles and failure rates over a rolling window, see below
  • Alerts — nine checks that ship on, and a webhook that speaks when one changes. See Alerts

Monitors

/monitors reports p50/p75/p95/p99 over a window you pick (1h, 6h, 12h, 24h, 7d), computed by Postgres with percentile_cont over the same workflow.workflow_runs the session list reads. Nothing is sampled and nothing is estimated.

Two things it does deliberately, because the obvious version of each is wrong:

Turn latency leads, session duration is reported separately. A $eve.type = 'session' row stays running for as long as the conversation is open, so a session someone left open overnight is a nine-hour "duration" that measures the human. Turns start and finish around one model exchange, so turn latency is the number worth alerting on. Session duration is still shown, over sessions that actually ended, and the two are never averaged together.

A failure is not just status = 'failed'. A turn killed by a provider rate limit emits turn.failed on the stream while its workflow row still reads status = 'completed' — the workflow handled the error, so nothing failed as far as it is concerned. eve writes $eve.model only once a model call reports usage, so a finished turn without it never reached the provider. Monitors counts those as failures under no model call and shows them separately from error_code failures, because counting only the latter is the direction that flatters us.

Unfinished turns are excluded from the distribution rather than counted as zero. That one is verified against a real server by contract/runtime/probes/05-monitor-percentiles.probe.mjs, with fixtures chosen so the wrong answer is numerically obvious — otherwise a busier agent would report as a faster one.

Control

This is where evestack diverges from Agent Runs, which is read-only:

  • Start a session and stream the reply straight from the dashboard
  • Send a follow-up to an existing session
  • Resolve a pending approval — see Architecture for the actual protocol underneath
  • Cancel a run — read the cooperative-cancellation note before building anything that assumes a cancel is instant
  • Promote a session to an eval — turn any run, especially one that went wrong, into a real evals/*.eval.ts replaying its actual messages (how to run one)
  • Replay a session into a new one — re-send its messages with one turn rewritten. Read Replaying a session first; it re-runs the original's tool calls

Running the evals

Save a promoted file into evals/ in the agent project, then:

npm run eval                    # the whole suite
npm run eval -- my-eval-name    # just the one you promoted

Not npx eve eval. eve boots its own dev server for a run, and it refuses to boot a second one for a project that already has one — so with npm run dev up, which is where the quickstart leaves you, a bare npx eve eval exits 1 with A dev server is already running for this eve agent and runs nothing. npm run eval hands eve the port this project recorded in EVESTACK_AGENT_PORT instead, and runs the same Postgres, schema and model checks that npm run dev runs first. A --url you pass yourself is always left alone.

The evals drive a real agent through a real Docker sandbox, so Docker has to be running. They need no paid key: on the $0 Ollama path the whole suite passes, and evals/memory.eval.ts needs the second, separate ollama pull nomic-embed-text — npm run eval checks for it by name rather than letting remember fail mid-run.

Replaying a session

Replay into a new session, on a session's page, re-sends that conversation's user messages into a fresh session, optionally rewriting one of them. It answers the question you always have after a bad run — would it have worked if I had said it differently? — without retyping the conversation and hoping you reproduced it.

The durable event log holds every user message, so the dashboard rebuilds turns 1..N and sends them in order: the first creates the fork, and each later one waits for the fork to park before it is sent. You pick where to stop, and may rewrite that last turn. Turns after it are dropped — they answered a conversation that is no longer happening.

It re-executes the tool calls, against the real world

Replaying turn 5 means running turns 1 through 4 again, for real. Nothing is stubbed or simulated: an email the original sent is sent again, money it spent is spent again, a PR it opened is opened again, a file it deleted is deleted again.

The panel is built around that one fact. It reads the per-turn tool list from the same transcript GET /api/evals/promote/:id reads, marks which turns are in range and which tools each of them ran, and keeps the run button disabled behind a checkbox that names those tools — a generic "this spends money" notice would not tell you that turn 3 called send_email. Changing the range clears the checkbox, because ticking it for turns 1–2 must not authorise turns 1–6.

That list is what the original ran, not a promise about the replay. The model is free to take a different path this time and call something that is not on it.

It is not a checkpoint fork

LangGraph forks from a serialized checkpoint: you branch off saved state, edit it if you want, and the turns before the branch point never execute twice. eve's durable record is an event log, not a resumable state snapshot, and it exposes no checkpoint to branch from — so the only way to reach turn 5 here is to run turns 1 through 4 again. That is strictly weaker than a checkpoint fork and strictly more expensive, and it is what the durable model allows.

What it replaces is not a cheap branch. It is a person pasting the same messages back into the chat box, which re-runs the same tools with no warning, no record of what it was forked from, and no chance to stop short.

What it costs

A fork is an ordinary session: it spends real tokens, and it shows up in Sessions with its own computed cost. It also takes about as long as the original conversation did, because each turn has to finish before the next is sent. Leaving the page stops the remaining turns from being sent; the ones already sent keep running.

A partial fork is a normal outcome

The replay stops at the first turn that does not land — it does not skip it and try the rest. When it stops, the fork already exists and already holds the turns that did land, so the route answers with the new session id, turnsPlanned, turnsDelivered and complete: false rather than an error, and the panel reports Sent 2 of 4 turns with the reason instead of claiming success. There is nothing to roll back: the turns that landed have already run their tools.

It stopped becauseWhat that means
awaiting_humana replayed turn parked on a tool approval or a question. The log records that a tool was denied, never why, so the replay will not answer for you — open the fork and answer it there
session_endedthe replay diverged into a run that finished or failed, so there is no session left to send the next turn to
timeouta turn did not come back within 90s, or the replay used up its 240s budget. Nothing is lost — continue the fork from its own page
session_mismatchthe agent started a different session instead of continuing the fork. That run is live and was not part of the replay; cancel it
agent_errorthe agent rejected the follow-up

The API takes the same view. fromTurn is required and never assumed — an empty body used to mean "replay everything", and the largest blast radius is the wrong thing to get by saying nothing. GET the same URL to see the turns, and the tools each one ran, without running any of them.

Who approved it

A browser button that can approve a shell command invites one fair question: approved by whom? eve cannot answer it — its human-in-the-loop protocol records that a request was answered and with which option, because that is all it needs to resume the turn. Identity is not part of it.

So evestack records every decision itself, in evestack.approvals, and shows them under Approvals: the tool, the decision, the person, the session, and — the column that matters — how the identity was established.

evestack deliberately ships no identity provider. The dashboard is meant to sit behind whatever you already trust, so it reads identity from the request and is explicit about its provenance:

SourceRecorded asWorth
EVESTACK_APPROVER_HEADER (a header you name)headeras trustworthy as your proxy
X-Forwarded-User / X-Forwarded-Emailforwarded-user / forwarded-emailset by oauth2-proxy, Cloudflare Access, and friends
HTTP Basic userbasicone shared credential — identifies a deployment, not a person
nothingunidentifiedrecorded anyway, and flagged

An anonymous decision is still written down. A silent gap in an audit log is worse than a visible one, and the Approvals page counts them at the top so they cannot be missed. Set EVESTACK_REQUIRE_APPROVER=1 to refuse them outright — the API answers 403 approver_required and the turn stays parked.

The audit row is written after eve accepts the answer, so the log never claims a decision that did not take effect. Retention is unbounded: this is the row you want a year later, when someone asks why the agent deleted the thing it deleted.

What the agent can be told to do

eve advertises every skill in agent/skills/ to the model and hands it a load_skill tool. Anything in that directory can put instructions into a live turn without a human seeing them first — which is a feature, and is also the reason the Skills page scans each one before listing it: credential reads, environment dumps, network exfiltration, and the shapes in between. ?selftest=1 runs the scanner against a bundled malicious specimen and expects a critical verdict, so you can tell a scanner that is working from one that is merely quiet.

A clean verdict is not proof of safety and the page says so in as many words. It is pattern matching over files; a skill that fetches its instructions at runtime has nothing for it to read.

Point it at the right directory. The page reads, in order: EVESTACK_SKILLS_DIR, then <cwd>/agent/skills, then the skills bundled with the evestack template. In a container the first two do not exist unless you arrange them, so it falls through to the third — and because the bundled skill has the same name the scaffolder writes into your project, the page looks like it is reading yours when it is reading the image's own copy. It labels the source (bundled template, with the absolute path) and says what to do about it.

Both compose files this project ships now arrange it for you:

    environment:
      EVESTACK_SKILLS_DIR: /agent-skills
    volumes:
      - ./agent/skills:/agent-skills:ro

Measured against the published image, with and without those two lines:

resolvedByreads
withoutbundled-template/repo/templates/default/agent/skills — inside the image
withenv/agent-skills — your project, read-only

If you run the image by hand, pass both or the scanner reports a clean verdict about files nobody is running.

Running it

In the project you already have. evestack create writes a dashboard service into your project's docker-compose.yml, behind a profile, pointing at the published image. So there is nothing to clone and nothing to build:

docker compose --profile dashboard up -d
npm run verify   # prints the URL and the sign-in pair

It reads the same .env.local your agent does, through env_file:, so the Postgres URL and the credential are already there. The port is whichever one the scaffolder published — it picks a free one, so it is not always 4000, and npm run verify is what tells you.

This section used to open with "The dashboard is not part of a create-evestack project — clone the repository for it", followed by git clone and cd evestack/packages/dashboard. Both halves were wrong: the scaffolder has shipped that compose service for a while, and the clone recipe did not work either. Installing from packages/dashboard cannot resolve workspace:*, and @evestack/schedules has no dist/ in a fresh clone — its dist/ is gitignored and it has no prepare script — so /schedules fails to compile with Can't resolve '@evestack/schedules/cron'. Three other pages said the opposite of this one.

From a clone, if you are working on the dashboard itself. Install and build from the repository root, not from packages/dashboard:

git clone https://github.com/SammyTourani/evestack
cd evestack
pnpm install
pnpm -r --if-present run build          # workspace dist/, which a cold clone has none of
cp packages/dashboard/.env.example packages/dashboard/.env.local
pnpm --filter @evestack/dashboard dev   # http://localhost:4000

Both halves of EVESTACK_AUTH_* are required and neither is defaulted: with either missing the dashboard serves nothing usable. 503 on every request, with four exceptions — PUBLIC_PATHS in lib/auth.ts holds three paths and proxy.ts opens that tier only to GET, so GET /signin renders the reason with no sign-in form on it, GET /api/auth/session and GET /api/auth/signout reach POST-only routes and get a bare 405, and GET /api/health reaches its own handler, which answers 503 {"status":"unconfigured"} and so reports the container unhealthy. None of the four exposes anything, and there is still nothing to sign in with. Use the same pair your agent has in .env.local — it is one secret per deployment, which the dashboard signs you in with and then presents to the agent.

Or as part of the full stack: docker compose --profile full up -d brings up Postgres and the dashboard together. The container runs as a non-root user; the process inside it binds 0.0.0.0 and what keeps it off your network is the compose port mapping, 127.0.0.1:4000:4000. See Self-hosting before exposing it beyond your own machine.

It reads your database, nothing else

No API calls to evestack, no telemetry, no external service in the loop. Its required environment variables are WORKFLOW_POSTGRES_URL and the EVESTACK_AUTH_* pair — point the first at your Postgres and everything on the observe side is a SQL read.