Skip to content
▚ evestack docs

Upgrading

Two different jobs under one name — moving a project you scaffolded onto a newer evestack, and moving this repository onto a newer eve.

Two audiences, two halves. Read the one you are.

Upgrading your project — you ran npx evestack create some months ago and want the newer dashboard, the newer template and the newer eve pin in your directory. Nothing about the evestack repository is involved.

Upgrading eve inside this repository — you are working on evestack itself, a contract went red, and you need to decide what that means. This is the half everything else links to: contract/run.mjs prints "See docs/upgrading.mdx" when the suite fails, and it means that half.

This page used to be only the second one, filed under a title that read like the first.

Upgrading your project

What a scaffolded project actually is

Worth stating before any commands, because it determines the whole shape of an upgrade: your project is a copy, not an install. create-evestack copies templates/default into your directory once and then has no further relationship with it. There is no evestack upgrade command — packages/evestack-cli/src/ contains no such module — and nothing in the project checks for a newer version of itself.

So these are the pieces, they move independently, and you move each one yourself:

In your projectWhat it isHow it moves
agent/, lib/, evals/, test/, tsconfig.json, HEARTBEAT.mdYour code. Copied from templates/default at scaffold time, then yoursBy hand, from a diff
scripts/*.mjs — approval-demo, bootstrap, checks, dev, eval, prune, retention, start, ui, verifyevestack's helper scripts, copied the same way. Most people never edit theseBy hand, from a diff
deploy/ — README.md, evestack-agent.service, dev.evestack.agent.plista systemd unit and a launchd plist for running the agent as a service, plus the notes for both. Copied, not generatedBy hand, from a diff
package.json dependencies — eve, @evestack/*, ai, @workflow/world-postgresOrdinary npmnpm install
docker-compose.yml — Postgres, and the dashboard behind a dashboard profileGenerated with the ports your machine had free, and committedEdit the image tag
.env.local and .envYour generated credentials. Both git-ignored, both 0600Never overwrite these from a fresh scaffold

.env.local is read by both the agent on your host and the dashboard container (through env_file: in the compose file). .env is read only by Compose itself, for interpolating ${...} in docker-compose.yml — Compose does not read .env.local. Keep the distinction in mind when an upgrade asks you to set a variable.

Find out where you are

# the eve you are pinned to
node -p "require('./package.json').dependencies.eve"

# the dashboard image tag your compose file names
grep image: docker-compose.yml

# whether the four parts are up, and where your dashboard actually is
npx evestack status
npx evestack open --no-open

Neither of those prints the running dashboard's version — status answers "is it up", open answers "where is it and what is the password". The version comes from /api/health, and the next block is how to ask it without guessing a port.

Do not reach for curl http://127.0.0.1:4000/api/health here. This section's whole job is to tell you which stack you are looking at, and 4000 is not necessarily yours: the scaffolder takes the first free port at or after 4000, so a second project on the same machine is published on 4001 or later and records that in its own .env.local. Measured on a project whose dashboard was on 4001, the line above returned {"ok":true,...} from a different project's dashboard — a healthy answer about somebody else's stack, which is the one wrong answer this section cannot afford. npx evestack status and npx evestack open read the port out of EVESTACK_DASHBOARD_URL in your project, so they answer about your project.

If you want the raw JSON, take the origin out of the project rather than typing a port:

DASH=$(grep -m1 '^EVESTACK_DASHBOARD_URL=' .env.local | cut -d= -f2- | sed 's#/api/.*##')
curl -s "$DASH/api/health"

EVESTACK_DASHBOARD_URL is written by create and by attach and holds the ingest endpoint, so its origin is your dashboard. It is the same line evestack open and evestack verify read.

Compare those against a freshly published create-evestack: scaffold a throwaway project (next section) and read its package.json and docker-compose.yml. That pair — template plus image tag — is the combination that was tested together, which is the reason packages/create-evestack/shared.mjs pins the image to a tag rather than latest.

Upgrade the dashboard

The dashboard is a container, so this is a repull and a restart. Your compose file names it as ${EVESTACK_DASHBOARD_IMAGE:-ghcr.io/sammytourani/evestack-dashboard:<tag>}, so you can override it without editing a committed file — the generated .env already carries the line, commented out:

Pick a tag that exists. The examples below say 0.4.0 because that is what this tree pins, and 0.4.0 is not published yet — a manifest request for it against GHCR returns 404, so docker compose pull fails rather than upgrading anything. The published tags at the time of writing were 0.1.0, 0.2.0, 0.3.0, 0.3.1 and latest. Ask the registry rather than trusting this list, which will age:

curl -s "https://ghcr.io/token?scope=repository:sammytourani/evestack-dashboard:pull&service=ghcr.io" \
  | sed 's/.*"token":"\([^"]*\)".*/\1/' \
  | xargs -I{} curl -s -H "Authorization: Bearer {}" \
      https://ghcr.io/v2/sammytourani/evestack-dashboard/tags/list

See CHANGELOG.md under Unreleased for why the pin runs ahead of the registry.

# in .env (NOT .env.local — Compose only interpolates from .env and the shell)
EVESTACK_DASHBOARD_IMAGE=ghcr.io/sammytourani/evestack-dashboard:0.4.0

Then:

docker compose --profile dashboard pull dashboard
docker compose --profile dashboard up -d dashboard

# confirm it — and read the `version`, not the `ok`
DASH=$(grep -m1 '^EVESTACK_DASHBOARD_URL=' .env.local | cut -d= -f2- | sed 's#/api/.*##')
curl -s "$DASH/api/health"

Editing the tag inside docker-compose.yml works too, and is the better choice if you want the version in version control.

"ok":true is not confirmation that the upgrade landed. That field answers "is this process up and can it reach Postgres", and it answered true before the pull as well. The field that moves is version, which the running image reports out of its own package.json — so the upgrade is confirmed when the number in the response equals the tag you just pulled, and not before:

{"ok":true,"database":"connected","version":"0.4.0"}

If it still reports the old number, the container was not recreated. docker compose --profile dashboard up -d dashboard recreates it only when the image reference it resolves has changed; docker compose --profile dashboard up -d --force-recreate dashboard is the hammer.

There is no migration step, and that is by design. A self-hosted install has no migration runner to hang one off, so the dashboard creates what it needs on first use: evestack.spans on the first trace read, the budget tables on the first budget write, and sql/facts.sql, sql/alerts.sql, sql/approvals.sql, sql/memory-audit.sql and sql/query-indexes.sql applied once per process from disk. Most statements in those files are CREATE … IF NOT EXISTS, CREATE OR REPLACE or ADD COLUMN IF NOT EXISTS — sql/alerts.sql already carries two columns added after the first release, with a backfill beside them — so going forward a newer image migrates itself on the first request that touches the table.

Going backward, a newer database now REFUSES an older image, and that is the most visible behaviour change on this release. Two of those files are not idempotent in the sense the note above describes, and both are versioned:

  • sql/traces.sql carries a schema marker for spans and sql/facts.sql one for facts. Each file opens with a guard that raises SQLSTATE EV001 when the database's marker is ahead of the version that file understands, and the guard is deliberately the first statement so that nothing is applied before it runs.
  • sql/facts.sql then DROP TABLEs all three fact tables whenever the marker is not its own version, and rebuilds them — the fact tables are a cache of a join, so rebuilding is the only strategy that cannot leave a column half-populated. That drop is exactly why the guard exists: without it an older image did not merely fail to upgrade a newer database, it dropped that database's fact tables, rebuilt them in the older shape, and stamped the marker back down to its own version without a word.

So rolling the dashboard image back over a database a newer image has written does not half-downgrade it — it stops. What you see:

  • /api/health answers 503 with {"ok":false,"status":"degraded","reason":"schema-too-new"}, which also marks the container unhealthy in docker ps. The body lists which pages are unavailable, which are degraded and which still work.
  • /traces, /sessions/[id], /costs, the overview and /api/metrics/query fail; /sessions degrades to blank fact columns; /monitors, /approvals, /schedules and /evals still work, because they read only the workflow tables. That is four working pages, not five — /charts is not one of them and is deliberately absent from all three lists: app/charts/page.tsx calls notFound() when NODE_ENV is production, so in any image you can pull it is a 404 whatever the schema says. It used to be listed here as "a static demo and stays up", which was wrong in both halves. The health response carries the same three lists, so read it rather than this paragraph if they ever disagree.
  • POST /api/ingest/v1/traces answers 503 to every batch — and the OTLP exporter reports a rejected batch as a success, so on the agent's side this looks like nothing at all.

The way out is forward: run the newer image again. Dropping the evestack schema also works and costs you every span and every materialized fact in it.

Stated from the guards and the health route rather than from a test run: the EV001 path has no automated coverage yet, because exercising it needs two dashboard images and a live Postgres. If you hit it and it behaves differently, that is worth an issue.

Upgrade eve and the evestack packages

npm install eve@0.30.8
npm install @evestack/budget@latest @evestack/composio@latest @evestack/schedules@latest
npm run typecheck
npm test
npm run verify

Do not npm install eve@latest here. This line said exactly that until 2026-08-09, and latest is 0.31.3, which this stack has not been released against.

0.31.3 stopped returning a continuation token from POST /eve/v1/session — its own docs say "session message and control request/response bodies do not accept or return continuation tokens", where 0.30.8's said the response "returns sessionId and the continuationToken". Dashboard images at 0.3.0 and earlier require that field and answer 502 — "the agent accepted the session but returned no handles" without it. That is every new chat and every fork, on the dashboard's only route for starting a conversation.

Dashboard 0.3.1 and later no longer require it and work against both, so 0.31.3 is safe once you are on 0.3.1 or newer. The pin above is the version the contract suite and the runtime probes are green against; docs/upstream.mdx tracks where the pin is going next.

Three things to know before you run that:

  1. eve below 0.30.0 is a security floor, not a preference. On 0.29.x, localDev() granted a full local-dev principal to anyone who sent a crafted Host header. The four @evestack/* packages that import eve declare peerDependencies.eve as >=0.30.0 <1.0.0, so npm will refuse the combination — but if you have --force in your muscle memory, know what it would be forcing. SECURITY.md has the detail.
  2. @workflow/world-postgres is pinned to an exact version, not a range and not a dist-tag. That is deliberate, and it is the second pin this dependency has had. npm latest is on the 4.x line, which eve rejects outright, so the template used to say "beta". A plain npm install re-resolves a dist-tag, so that dependency moved under people without any version number in their package.json changing — and when upstream raised the World spec version from 5 to 6 inside 5.0.0-beta.*, every fresh install died at boot with This Workflow runtime requires a World with matching spec version 5. The template now declares "5.0.0-beta.32" exactly. If you are upgrading a project scaffolded before that, change your own package.json to the exact version too; ^ and ~ do not help, because both still admit 5.0.0-beta.34. If sessions start failing right after an unrelated install, check this first.
  3. npm run verify is the real gate, not npm test. Its own header names what it walks: "Postgres, Docker, the agent, the dashboard, an embedding probe" — plus the model key and the schema — against the stack as it is actually running, with the fixing command on every red line. It exits 1 if anything required failed, so you can put it in a script.

If a new eve changed the workflow schema, npm run db:bootstrap is what applies it. That script is a thin wrapper around @workflow/world-postgres's own setup script — it checks the connection first so a failure names a host and a port instead of a Drizzle stack trace, then hands over unchanged. The migrations are upstream's.

Pick up template changes

This is the step with no automation, so here is the routine that works:

cd /tmp
npx create-evestack@latest reference-agent

Answer no when it offers to bring the stack up — you only want the files. It generates its own credentials, which you will not be copying anywhere.

diff -ru ~/my-agent/scripts /tmp/reference-agent/scripts
diff -ru ~/my-agent/deploy  /tmp/reference-agent/deploy
diff -u  ~/my-agent/package.json /tmp/reference-agent/package.json

scripts/ is the highest-value diff and the safest to take wholesale: unless you have edited them, those eight files are evestack's, and verify.mjs in particular gains checks as new failure modes are found. deploy/ is the systemd unit and the launchd plist, and is the one to read rather than copy — both carry absolute paths you filled in for your machine. A project scaffolded before those files existed will show them as additions. Take the package.json diff as information rather than as a patch — yours has your project name, and may have dependencies you added.

diff -ru ~/my-agent/agent /tmp/reference-agent/agent
diff -ru ~/my-agent/lib   /tmp/reference-agent/lib

Expect noise here — this is your agent's instructions, tools and channels, and you have presumably changed them. What you are looking for is a shape change: a new file, a changed import path, a rewritten auth chain in agent/channels/eve.ts. Apply those by hand.

diff -u ~/my-agent/docker-compose.yml /tmp/reference-agent/docker-compose.yml

It is generated with the ports your machine had free and a project name derived from your directory, so copying it over will point the stack somewhere else. Read it for a new service, a new environment variable, or a new mount, and transplant just that.

It holds a generated dashboard password, a trace-ingest token and a Postgres password that are now lying around in /tmp.

rm -rf /tmp/reference-agent

Upgrade the CLI itself

npx evestack@latest … pins nothing and resolves the newest published version, so there is nothing to maintain. A global install does not update itself:

npm i -g evestack@latest

The order that matters

The dashboard image and the agent template are versioned separately and released together, and the scaffolder pins one image tag per template version precisely because that is the combination that was tested. Moving one a long way without the other is untested rather than forbidden. If you are several releases behind, move both, from the same release, and run npm run verify afterwards.

Nothing here rewrites your data. The one destructive operation in this area is docker compose down -v, which deletes the Postgres volume — every session, trace and memory with it — and it is never part of an upgrade. It is only ever needed to make a new database password take effect (EVESTACK_DB_PASSWORD in .env, which Compose interpolates into POSTGRES_PASSWORD), because Postgres applies that variable once, when the volume is first created.


Upgrading eve inside this repository

Everything below is for work on evestack itself. If you are running a scaffolded project, you want the half above.

Vercel merged 252 pull requests into eve in fourteen days. eve shipped 0.30.0, 0.30.1 and 0.30.2 on the same day evestack launched. That cadence is the environment evestack lives in, and it has one practical consequence:

Anything evestack builds inside eve's own surface area has a shelf life measured in weeks. Assume every assumption below is temporary and write it down where a machine can check it.

The policy

  1. Every assumption evestack makes about eve is a contract in contract/. If you find yourself writing "eve does X, so we can do Y", that belongs in the suite before it belongs in the code.
  2. eve-watch opens the pull request. A human merges it. There is no auto-merge and there will not be one.
  3. A red contract is a decision, not a bug. Sometimes eve is wrong, sometimes we are, sometimes both are right and the contract has gone stale. Those three outcomes need different fixes.
  4. Deterministic checks decide; a model may advise. See On putting a model in CI.

Why a typecheck is not enough

evestack once shipped a strictLocalDev() wrapper. eve 0.29.x decided "is this request from my own machine" by matching the request's own hostname against an unanchored /^127\./, and a request URL is built from the client's Host header. So 127.evil.com — a name anyone can register and point at your agent — received a full local-dev principal with no credentials at all. We measured it: that host answered 200 where a plain foreign host answered 401.

eve 0.30.0 fixed it properly upstream. localDev() now grants based on the process being an eve dev run, and consults nothing in the request. At that moment our wrapper stopped adding protection and started rejecting legitimate local-dev access over a LAN IP, a tunnel, or a container hostname. It had to be deleted.

It typechecked perfectly on both days. tsc had nothing to say when it was load-bearing and nothing to say when it became harmful, because its types never changed — only eve's meaning did.

That is the entire argument for the contract suite. A typecheck cannot catch semantic drift. A behavioural assertion can:

# green against the eve we ship
node contract/run.mjs --only=auth

# red against the eve that had the bug — naming the exact hostile Host header
EVESTACK_CONTRACT_EVE_DIR=node_modules/.pnpm/eve@0.29.5.../node_modules/eve \
  node contract/run.mjs --only=auth

When the suite goes red

Read the failure first. Every contract prints the assumption it pins and what evestack does with it, so the report tells you the blast radius before you open a single file.

pnpm install
pnpm contract
node contract/run.mjs --only=<the failing id> --verbose

The suite is free and offline. There is never a reason to debug this from CI logs alone.

What you findWhat it meansWhat to do
eve changed behaviour we depend on, deliberatelyThe contract did its jobUpdate evestack's code, then update the contract to describe the new behaviour — same commit
eve changed behaviour and it looks like a regressionThe contract found an upstream bugFile it at vercel/eve, pin the previous version, leave the contract red with a link
eve is unchanged; our contract was over-specifiedThe contract is wrongLoosen it to assert what we actually depend on, not what we happened to observe

A deleted contract is an assumption that silently stopped being checked, and the next person has no way to know it was ever true. Loosen it, or replace it with the narrower thing we genuinely rely on. If evestack really no longer depends on it, delete the dependency in the same commit and say so in the message.

The pull request reports the suite against the current pin as well as the candidate. If that column is red, main was broken before this upgrade and the bump is a distraction — fix main first.

When it stays green and something breaks anyway

This is the failure mode worth planning for. A green suite means every assumption someone wrote down still holds — not that the release is safe. eve can change something we depend on that no contract names, and the suite will report success with total confidence.

When that happens, the fix is not just the patch. Write the contract that would have caught it, in the same pull request. That is the only mechanism that makes the suite better over time, and it costs about fifteen minutes while the failure is still fresh. contract/README.md describes the shape.

What is currently pinned

ContractAssumption
version/…One eve version satisfies every range this repo declares — templates, both published packages' peer ranges, the scaffolder
modules/…Every eve/* subpath evestack imports resolves, and every value it binds is still exported
tools/approval-…approval is the gating field (not the AI SDK's needsApproval), and always() still returns "user-approval"
tools/dynamic-…defineDynamic still produces kind: "eve:dynamic", and step.started is still dispatched
protocol/…The session/cancel/stream routes, the NDJSON content type and the stream resume headers the dashboard drives
attributes/…Every $eve.* run attribute the dashboard reads out of Postgres is still one eve writes
auth/…localDev() grants on process state and never on the request; httpBasic() and routeAuth() fail closed
hooks/…Hook handlers return void — they cannot block or park a turn
sandbox/…The SandboxBackend shape @evestack/sandbox-opensandbox duck-types, which nothing typechecks

The import and attribute lists are derived from evestack's own source at run time, so they widen automatically as the codebase grows.

The eve-watch workflow

.github/workflows/eve-watch.yml runs daily and on demand. It asks npm for the latest eve; if it is newer than the pinned range it creates a branch, bumps every manifest with a caret pin, installs, typechecks, runs the contract suite, and opens a pull request reporting all of it alongside the eve changelog entries that touch something the suite pins.

Two details are deliberate:

  • Peer ranges are not bumped. peerDependencies.eve on the published packages is a compatibility promise to users. Widening it automatically would publish support for a version nothing has been tested against. If a new eve falls outside one, the version contract fails and a human decides.
  • The baseline runs first. The suite runs against the current pin before the bump, so a red result after the bump is unambiguous.

Run it by hand against a specific version when you want to test a release candidate:

Actions → eve-watch → Run workflow → version: 0.31.0

On putting a model in CI

The honest engineering read, since it comes up every time.

Where a model genuinely helps. Reading a hundred-line changelog and telling a maintainer which three entries matter to a self-hosted, non-Vercel deployment is a real task with no deterministic solution, and a model is good at it. So is drafting a diff for a mechanical rename once a human has decided the rename is correct. Both are suggestions to a person who is already reading the PR.

Where it is dangerous. Auto-merging a model's fix to an auth path. The localDev() story is the argument: the correct response to that change was to delete protective code, and deleting security code on a model's say-so, on a schedule, with nobody watching, is how you ship an authentication bypass to everyone who ran npx create-evestack. A model asked "does this patch still help?" would very plausibly have said yes — the patch looked defensive and compiled fine.

There is also a supply-chain problem. A changelog is third-party text that arrives on a schedule. A model reading it inside a job that can write to the repository is a prompt-injection target with commit access. The mitigation is not a better prompt; it is not giving it commit access.

So the rule is:

Deterministic checks are the mechanism. A model is a convenience layer on top — optional, gated on a secret being present, and able only to comment.

The optional step in eve-watch follows exactly that. It runs only if OPENAI_API_KEY is set, it reads results the deterministic steps already produced, and its single capability is gh pr comment. It cannot commit, push, merge, re-run the suite, or change what the PR reports.

The model it uses is the repository variable ADVISOR_MODEL, defaulting to gpt-5-mini — set it under Settings → Secrets and variables → Actions → Variables to change it without touching this workflow. The same key drives the nightly evals, so one secret covers both jobs. Its output is labelled as generated and non-authoritative. If it is prompt-injected by a hostile changelog, the worst outcome is a misleading paragraph next to the real contract table — which a human is required to read anyway.

If you do not set the secret, everything above still works. That is the test of whether a model is a convenience layer or a dependency, and this one passes it.

The two questions it is actually asked

Not "summarise this PR" — the deterministic steps already did that, better. It is asked the two things they structurally cannot answer:

1. Did a check pass or fail without reaching its subject? deny-survives parks a turn and denies a gated tool call. If the model under test answered in prose and never called the tool, the run ends waiting with no pending requests and the denial was never exercised at all. That is a vacuous run, not a regression, and the difference decides whether a maintainer re-runs or investigates. It is also not expressible as an assertion: the eval cannot tell "the bug is gone" from "the test never got there" without judging what the transcript means. Naming vacuity is the single change that makes a nondeterministic gate readable enough to keep.

The output annotates. It can never clear a result — a red stays red no matter what the paragraph says.

2. What did this release change that no contract asserts? The suite covers what somebody thought to write down, which is why every red contract's advice ends with "the fix is a new contract, not just a patch." That instruction has always resolved to when a human happens to notice. The step is now handed a listing of every contract and every assertion it makes, and asked to name surface the release touched that nothing pins — writing the assertion as one sentence of prose, in a fenced block, for a human to accept or ignore.

It does not write code, name a file, or open anything. The ratchet is that the suite has a mechanism for growing that is not "remember to."

Both inputs — the changelog and the release notes — are third-party text on a schedule, so the prompt requires anything that appears to address or instruct the model to be quoted verbatim under SUSPICIOUS TEXT and not acted on. That matters most for exactly the sentence a compromised release would want to write: "downstream repos should drop the Host-header assertions."

Running the suite yourself

pnpm contract                        # all contracts
node contract/run.mjs --verbose      # show passing assertions
node contract/run.mjs --only=auth    # one family
node contract/run.mjs --format=json  # for scripts

# point it at any other eve install
EVESTACK_CONTRACT_EVE_DIR=path/to/eve node contract/run.mjs

pnpm test runs the contract suite before the workspace tests, so it is hard to skip by accident.