# evestack, complete documentation > Every page of https://evestack.vercel.app/docs, concatenated, in reading order. > Generated at build time from the repository's docs/ directory. > > Looking for something smaller? https://evestack.vercel.app/agent.md is a paste-sized setup pack, > and https://evestack.vercel.app/llms.txt is the linked index. # Introduction > eve on your own machine, with a dashboard you can drive. $0 infrastructure. [Vercel's `eve`](https://github.com/vercel/eve) is a genuinely good, Apache-2.0 agent framework, and self-hosting it is a supported path — Vercel says so themselves. What you give up off-platform is the **Agent Runs dashboard**; their own answer is to export OpenTelemetry to a collector you run. That leaves you with real tools and a real gap. Jaeger will show you spans, the Workflow DevKit inspector will show you runs, and eve's own dev TUI is better than it gets credit for: since `0.29.3` its `/traces` command is a full-screen live viewer over the local trace spool that replays a session as expandable conversation cards, and `0.30.7` added token/cost/tool chips plus `--verbose` and `--json` to `eve traces`. What none of them give you is one place that understands *sessions, turns, tokens, cost and approvals*, that keeps history for as long as you keep the rows rather than until a spool sweep, that more than one person can open, and that can **act** on the agent rather than only read it. evestack is to `eve` what Ubuntu is to the Linux kernel — not a fork, a distribution. It packages `eve` to run on your own hardware and fills in the operational layer that self-hosting still needs. evestack is not the first self-hosted eve distribution, and this page used to imply it was. [vercel-labs/steve](https://github.com/vercel-labs/steve) — "Self-hosted eve poc" — was published by a Vercel employee on 2026-06-24, weeks before this project. Nothing here was "kept" by anyone. `npx evestack create my-agent` scaffolds it and offers to bring the rest up for you, dashboard included — it is a compose profile in the project, not a separate clone. How the dashboard reads eve's own Postgres tables directly — no ingest pipeline. ## What you get | | evestack | eve on Vercel | | --- | --- | --- | | Compute, workflows, sandbox | your machine — **$0** | metered | | Durable sessions | your Postgres | Vercel Workflows | | Run-state retention | as long as you keep the rows | purged 1 day (Hobby) / 7 (Pro) / 30 (Enterprise) after a run completes ([docs](https://vercel.com/docs/workflows/pricing)) | | Dashboard | yes, anywhere | Vercel only | | Dashboard can *drive* the agent | yes | read-only | | One-click tool sign-in | 121 managed-OAuth toolkits, plus API-key auth across 1,000+ (Composio, hosted) | 4 Vercel-managed connectors, plus Custom OAuth and API-key connectors you register ([docs](https://vercel.com/docs/connect)) | | Long-term memory | included (pgvector, your database) — needs an embeddings provider, and Anthropic is not one ([why](/docs/memory)) | no first-party store; the gallery's memory options are third-party SaaS | The memory row used to read "not included," which is wrong. eve's own `patterns/multi-tenant-memory.md` opens: "You can add long-term memory from the integration gallery using the Memory filter, or build tenant-aware memory from your own application store." The narrower claim is the true one and still the interesting one — eve ships no first-party memory *store*, deliberately ("The storage implementation is deliberately outside eve"), and the packaged options in the gallery (Mem0, Upstash AgentKit) are other people's hosted services. evestack's is a table in the Postgres you already run. The only real cost is model tokens — and [Ollama](/docs/local-setup#ollama) takes that to zero too. Composio is the exception to "everything on your machine": it is hosted, it holds the OAuth tokens for accounts you connect, and it is off until you set `COMPOSIO_API_KEY`. ## The honest tradeoffs evestack doesn't hide the parts of self-hosting that are genuinely harder than paying Vercel: - **Frictionless auth needs someone else's OAuth app.** Composio's managed auth is one click, but shows a Composio-branded consent screen and shares rate limits across every Composio customer. Read [Composio auth](/docs/composio-auth) before onboarding real users. - **Cancellation is cooperative, not immediate.** We measured a cancelled turn continuing to stream for roughly 90 seconds. See [Architecture](/docs/architecture#cooperative-cancellation). - **You own operations.** Backups, uptime, and scaling are yours — that's the actual trade for $0 infrastructure and full data ownership. --- # Quickstart > One command to a running, durable, dashboard-observed agent. ## Requirements - Node 24+ - Docker, running (for Postgres and the agent sandbox) - An `OPENAI_API_KEY` or an `ANTHROPIC_API_KEY` — or skip straight to [Ollama](/docs/local-setup#ollama) for $0 total ## Create the project ```bash npx evestack create my-agent ``` `npx create-evestack my-agent` does the same thing. They are two published names for one implementation — `evestack` is the single command (`create`, `status`, `tour`, `open`, `verify`, `attach`, `doctor` — see [the CLI reference](/docs/cli)), and `create-evestack` is the name npm's `create-*` convention leads people to. Same code, same prompts, same flags. It asks **four** questions, all of them before any work starts, so the install and the image pull are a wait you can walk away from: 1. **Where** — the project directory. 2. **Model** — OpenAI, Anthropic or Ollama, and a key if you have one. 3. **Tools** — Composio's one-click tool sign-in, on by default. 4. **Bring it up** — whether it should start Postgres, create the schema and pull the dashboard for you. Pass `--yes` to skip all four and fill in `.env.local` yourself afterward — the scaffolder never hangs waiting on input it isn't going to get. **Question 4 decides how much of this page you have to type.** Answer **yes** and the scaffolder does *Bring up the database* and *See it in the dashboard* for you, then offers to run *Run it* as well — so there is nothing left to paste, and you can read those sections as a description of what already happened. Answer **no** — or run under `--yes`, or with Docker not running — and it prints those same four commands when it finishes, which are the ones below. Whichever you pick, `.env.local` gets `EVESTACK_PROVIDER` *and* that provider's own key variable. Both matter: `EVESTACK_PROVIDER` is what `agent/agent.ts` branches on, and setting a model name without it leaves you on the previous provider. See [the provider table](/docs/local-setup#picking-a-provider) if you're switching later by hand. **The provider you pick decides whether long-term memory works.** `remember` and `recall` need an embeddings model, and the three providers differ: OpenAI has one on the same key; Ollama has one but it is a **second pull** (`ollama pull nomic-embed-text`); **Anthropic has no embeddings endpoint at all**, so on that path memory needs either an `OPENAI_API_KEY` as well or `EVESTACK_EMBED_PROVIDER=ollama` to run embeddings locally. Nothing else is affected — the agent, its sandbox, durable sessions and the dashboard all work either way, and the first `remember` call names the variable that fixes it. Full detail in [long-term memory](/docs/memory). This generates a unique `EVESTACK_AUTH_PASSWORD` per project. `eve` fails closed by default, so there's no shipped default password sitting between a stranger and your agent. ## Bring up the database ```bash cd my-agent docker compose up -d postgres npm run db:bootstrap ``` Nothing creates the `workflow` schema for you: `@workflow/world-postgres` runs its migrations only from its own CLI, and `eve` never invokes it. Skip this and `npm run dev` starts against a database with no tables. Run it through the script, not as `npx --package=@workflow/world-postgres bootstrap`. That CLI loads `.env` through dotenv and never looks at `.env.local` — which is the only env file `create-evestack` writes — so it silently falls back to `postgres://world:world@localhost:5432/world` and dies on `ECONNREFUSED`. The `db:bootstrap` script passes `--env-file-if-exists=.env.local` explicitly. Pin `@workflow/world-postgres` to an **exact version**. npm's `latest` is the 4.x line and `eve` rejects it outright, but the `beta` dist-tag is not the fix — upstream moves the World spec version inside `5.0.0-beta.*` with no semver signal, so `beta` resolving one release forward is enough to kill a boot. The scaffolded `package.json` pins `5.0.0-beta.32`; leave it exact if you ever touch that dependency by hand. ## Run it ```bash npm run dev ``` The agent boots on the port the scaffolder wrote into `.env.local` as `EVESTACK_AGENT_PORT` — 2000 unless that was already taken when you scaffolded. `npm run dev` passes it to `eve dev` as `--port`, which means **there is no auto-increment**: eve only scans for a free port when no port is given at all, so if something else has grabbed that port since, the boot fails with a plain `EADDRINUSE` rather than moving. Free the port, or change `EVESTACK_AGENT_PORT` in `.env.local` — and if you do, change the `EVESTACK_AGENT_URL` default in `docker-compose.yml` to match, because the scaffolder wrote that number in at generation time. ## Verify it's durable The whole point is that restarting doesn't lose anything. Prove it to yourself: ```bash curl -X POST http://127.0.0.1:2000/eve/v1/session \ -H 'content-type: application/json' \ -d '{"message":"Remember this exact phrase: quokka-orbit-9."}' ``` Kill the dev server (`Ctrl-C`), restart it (`npm run dev`), then send a follow-up using the `continuationToken` from the first response. The agent recalls it — durable state survives the process, because it never lived in the process to begin with. ## See it in the dashboard One command, in the project you just scaffolded: ```bash docker compose --profile dashboard up -d ``` That is a **pull**. The generated `docker-compose.yml` points at `ghcr.io/sammytourani/evestack-dashboard`, pinned to the version tested with this template, and the service is already wired to this project's database and to the `.env.local` your agent reads — so there is no repository to clone, no image to build, and no credential to copy across. That pull resolves a multi-arch manifest — `linux/amd64` and `linux/arm64`, ~230 MB compressed each — so the same command works on an Apple Silicon laptop and on an x86 server without you choosing a platform. It unpacks to about 1 GB on disk. To run an image of your own — a fork, a private registry, a local build under a different name — set `EVESTACK_DASHBOARD_IMAGE` in a `.env` beside the compose file, or export it in your shell. Nothing else in the generated compose file needs to change. Open the dashboard — **usually** `http://localhost:4000`, but check what the scaffolder printed. It calls `freePort(4000)` and takes the first port at or above 4000 that is actually free, so a second project on the same machine, or anything already holding 4000, moves it. The same is true of the agent's 2000 and Postgres's 5433. The generated `docker-compose.yml` and `.env.local` both carry the numbers this project actually got. Sign in with the `EVESTACK_AUTH_USER` and `EVESTACK_AUTH_PASSWORD` generated into `.env.local` — the scaffolder prints them when it finishes. Every route is behind that credential: the dashboard starts agent runs, approves gated shell commands and deletes memories, so it fails closed rather than serving a viewer to anyone who reaches the port. Scripts can use HTTP Basic instead: ```bash curl -u "$EVESTACK_AUTH_USER:$EVESTACK_AUTH_PASSWORD" localhost:4000/api/fleet ``` Every session you just created is already there, read straight out of the same Postgres — no ingest step required for the session list itself. If the container comes up, `docker ps` shows it unhealthy, and every page answers `503` except `/signin` — which renders an error and no sign-in form — the credential did not reach it. The compose service reads `.env.local`, so that is the file to check — both halves are required, and a blank password counts as unset. To hack on the dashboard rather than run it, install and build from the repository root — not from `packages/dashboard`, where `workspace:*` cannot resolve and `@evestack/schedules` has no `dist/` for Turbopack to find: ```bash cd evestack pnpm install pnpm -r --if-present --filter '@evestack/dashboard^...' run build EVESTACK_AUTH_USER=evestack EVESTACK_AUTH_PASSWORD=dev \ WORKFLOW_POSTGRES_URL=postgres://evestack:evestack@localhost:5433/evestack \ pnpm --filter @evestack/dashboard dev ``` --- # Set up with an agent > Hand evestack to Claude Code, Cursor or a local model and let it do the setup. You do not have to read these docs to use evestack. There is a skill pack written for whatever coding agent you already work in, and handing it over takes one paste. ## The fastest path Press **Set up your agent** on [the landing page](/), then paste into your agent. That is the whole flow. Your agent gets the mental model, the five-command bring-up, the CLI, the dashboard's HTTP API, how to add tools and skills, and the failure modes that present as something else. The pack opens with an instruction to save itself, so an agent that can write files turns that one paste into a skill on disk and never needs the paste again. ## Install it instead ```bash npx evestack skills ``` Writes `SKILL.md` and four reference files where your agent will look for them: | Flag | Effect | | --- | --- | | `--dir=PATH` | Where to write it. Default: `agent/skills/evestack` inside an eve project, otherwise `.claude/skills/evestack`. | | `--print` | Write the pack to stdout and touch no files. | | `--force` | Overwrite files that already exist. Without it, an existing file stops the run and nothing is written. | | `--json` | Report what was written, as JSON. | The default target is chosen rather than guessed. Inside a scaffolded project `agent/skills` is a real runtime location — eve scans it and hands your agent a `load_skill` tool — so a pack written there is loadable by the agent you are *building*, not only readable by the agent helping you build it. Everywhere else it lands in `.claude/skills`, which is created if you have never made one. ### Pointing it somewhere else `EVESTACK_PACK_URL` overrides where the pack is fetched from. It defaults to `https://evestack.vercel.app/agent-pack.json`, and you will only need it if you are testing a fork, running against a local build of this site, or serving the pack from your own mirror: ```bash EVESTACK_PACK_URL=http://localhost:3000/agent-pack.json npx evestack skills ``` The pack is fetched from the site rather than bundled into the CLI. That is deliberate: a copy vendored into a published package drifts from the source, and a stale pack means your agent confidently repeats something that stopped being true a release ago. The cost is that `evestack skills` needs a network connection — which `npx` already needed to run it at all. ## Every URL, and what each is for | URL | What it is | Reach for it when | | --- | --- | --- | | [`/agent.md`](https://evestack.vercel.app/agent.md) | The pack, as one self-contained document (~34 KB) | You want to paste, or your model has no fetch tool | | [`/agent-pack.json`](https://evestack.vercel.app/agent-pack.json) | The pack as a file tree | You are installing it programmatically | | [`/llms.txt`](https://evestack.vercel.app/llms.txt) | The linked index | You want to know where something lives | | [`/llms-full.txt`](https://evestack.vercel.app/llms-full.txt) | Every documentation page, concatenated (~240 KB) | You want the whole corpus in one request | `/agent.md` is the one to reach for when the task is *doing something with evestack*. `/llms-full.txt` is the firehose — correct for a large-context model, wasteful for a chat window. ## What is in the pack One skill and four references, so the references stay out of your agent's context until it asks for one: - `SKILL.md` — what evestack is, the four moving parts, the five commands, and the five things that most often go wrong - `references/cli.md` — the eight `evestack` commands, and which of the three diagnostics answers which question - `references/build-an-agent.md` — tools, skills, memory, schedules, evals, channels, registry - `references/dashboard.md` — pages, the HTTP API, auth, and `@evestack/mcp` - `references/troubleshooting.md` — the failure modes, in the order you hit them The source is [`/skills/evestack`](https://github.com/SammyTourani/evestack/tree/main/skills/evestack) in the repository. It is written once and served three ways — copied, installed, fetched — so there is no second copy to go stale. ## Give your agent live access too The pack teaches your agent about evestack. [`@evestack/mcp`](/docs/mcp) lets it read your running stack — sessions, costs, the approval audit log: ```jsonc { "mcpServers": { "evestack": { "command": "npx", "args": ["-y", "@evestack/mcp"], "env": { "EVESTACK_MCP_DASHBOARD_URL": "http://localhost:4000" } } } } ``` Read-only by default, and the four tools that can act on a live agent are withheld from `tools/list` entirely until you set `EVESTACK_MCP_ALLOW_CONTROL=1`. Turning that on lets a model approve a gated tool call a human was asked to stand at — which is the entire reason eve pauses the turn. See [MCP](/docs/mcp) before you do. ## A note on trusting it The pack is instruction text that reaches a model's context and gets quoted back at you as fact, so it holds itself to the rules the rest of this project does: every claim reproducible, no first-or-only claims, no price comparisons, and "I don't know, here are the docs" preferred over a confident guess. If you find something in it that is wrong, that is a bug worth [filing](https://github.com/SammyTourani/evestack/issues). --- # Local setup & troubleshooting > Docker, Postgres, and the Ollama path to a genuinely $0 stack. ## Picking a provider `EVESTACK_PROVIDER` selects the branch in `agent/agent.ts`. It is the only variable that does, and `EVESTACK_MODEL` on its own never changes providers — it just hands a different model name to whichever provider is already selected. | `EVESTACK_PROVIDER` | Key it reads | Default `EVESTACK_MODEL` | | --- | --- | --- | | unset, or `openai` | `OPENAI_API_KEY` | `gpt-5-mini` | | `anthropic` | `ANTHROPIC_API_KEY` | `claude-sonnet-5` | | `ollama` | none | `qwen3` | `create-evestack` asks which one you want and writes both lines. Anything other than those three values is rejected at boot with the list above rather than silently treated as OpenAI — a misspelled provider is a mistake, not a preference. One thing that table does not cover, and it decides whether long-term memory works: **embeddings**. `openai` has an embeddings endpoint on the same key. `ollama` has one, but it is a **second pull** (`ollama pull nomic-embed-text`). `anthropic` has none at all, so `remember` and `recall` on that path need either an `OPENAI_API_KEY` as well or `EVESTACK_EMBED_PROVIDER=ollama` to run embeddings locally. Nothing else about the Anthropic path is affected. The three `EVESTACK_EMBED_*` variables are in [long-term memory](/docs/memory). ## Ollama For zero cost, total: ```bash ollama pull qwen3 # the chat model ollama pull nomic-embed-text # embeddings — a separate model, and `remember` needs it ``` Then set **both** of these in `.env.local`: ```bash EVESTACK_PROVIDER=ollama EVESTACK_MODEL=qwen3 ``` `EVESTACK_PROVIDER` is the one that selects the local path. The model name on its own is handed to the OpenAI provider, and the agent then refuses to boot: ``` Cannot compile agent compaction because the primary compaction trigger model "openai/qwen3" does not have known AI Gateway context window metadata. ``` `create-evestack` offers this as an option at scaffold time and writes both lines for you. Two more, both optional: ```bash OLLAMA_BASE_URL=http://127.0.0.1:11434 # bare host — no /api suffix EVESTACK_CONTEXT_WINDOW=32768 # qwen3's native window ``` `ai-sdk-ollama` appends the API path itself, so a `OLLAMA_BASE_URL` ending in `/api` gives `OllamaError: 404 page not found`. `EVESTACK_CONTEXT_WINDOW` is what the template passes to eve's `modelContextWindowTokens` — the escape hatch that skips the catalog lookup entirely. Raise it for a model with a bigger window; set it too high and compaction fires too late to save the turn. Both are ignored on the hosted providers, whose models the catalog does know. **Check your free RAM first.** qwen3 is 5.2 GB, and it loads on top of Docker, Postgres, the dashboard and the agent — plus the 274 MB embedding model once memory is used. On a machine without roughly *both model sizes + 4 GB* to spare, this does not degrade gracefully — it can take minutes to answer a one-word prompt and can take the host down. On a laptop already running the rest of the stack, a hosted key is the practical choice. Local models generally have weaker tool-calling than gpt-5-mini or Claude. That's the honest tradeoff for $0 total cost — try it, and switch to a cloud key if tool use feels unreliable. ## What it costs on disk The RAM warning above is the one that bites first, but disk is the one nobody budgets for. Measured on macOS with Colima, one project, clean Docker: | | With an API key | On the $0 Ollama path | | --- | --- | --- | | Docker images | ~2.0 GB | ~2.0 GB | | Postgres volume | ~70 MB | ~70 MB | | Project directory | ~300 MB | ~300 MB | | `qwen3` + `nomic-embed-text` | not pulled | 5.5 GB | | **Total** | **≈ 2.4 GB** | **≈ 7.9 GB** | **Budget 10 GB free on the local-model path, 4 GB with an API key.** That leaves room to grow; the flat totals do not. The image figure is `pgvector/pgvector:pg17` at 646 MB, the dashboard at 1.05 GB unpacked from a 228 MB pull, and a 665 MB sandbox template image eve builds on first run, less the `node:24-slim` layers two of them share. The project directory is almost entirely `node_modules` (262 MB measured); `.eve/` accounts for the rest and grows while `eve dev` is running. Each additional project on the same machine costs about **0.5 GB**. The big images are shared — its own sandbox template image (+157 MB), its own Postgres volume (~70 MB) and its own `node_modules` are not. ### What grows, and the one thing nothing cleans up Spans are about 1.55 KB each and self-prune on a 30-day window (`EVESTACK_TRACE_RETENTION_DAYS`). The `workflow` schema **never prunes itself**; `npm run db:prune` is the manual answer to both, and [Operations](/docs/operations) covers it — including why space is not returned to the filesystem without a `VACUUM FULL`. The exception is the per-project **`eve-sandbox-template:*` image**. eve builds one per project, tagged with a hash of the project id, and at 665 MB it is the third-largest single item on disk. Nothing removes it: no evestack command, no `docker compose down -v`, and not deleting the project directory either. It outlives the project that created it, and a cleanup on the development machine reclaimed several GB from seven stale ones that had accumulated this way. ```bash # what is actually there docker images --filter reference='eve-sandbox-template:*' # remove them — `docker image prune` rejects a reference filter, so list and pipe docker images --filter reference='eve-sandbox-template:*' -q | xargs docker rmi ``` Do that after you delete a project, not while one is running: the tag carries no hint of which project it belongs to, and a live project rebuilds its own on the next turn. ## Docker sandbox The default backend, `docker()`, keeps one long-lived container per durable session and persists `/workspace` across turns with **no idle timeout** — genuinely free to run 24/7 on your own machine, since it's just a container. Swaps, one line in `agent/sandbox/sandbox.ts`: - `microsandbox()` — real VM isolation, domain-level network policies, credential brokering. macOS on Apple Silicon or Linux with KVM only. - `justbash()` — no daemon at all, but no real binaries either. ## Common issues **`eve dev` seems to not be listening on port 2000.** It is listening on `EVESTACK_AGENT_PORT` from `.env.local`, which the scaffolder set to the first free port at or above 2000 *at the time you scaffolded* — so on a machine that already had something on 2000, it is 2001 or higher. `npm run dev` passes that number to `eve dev` as `--port`, and eve only scans for a free port when it is given none, so nothing auto-increments at boot: if that port is busy now, you get `EADDRINUSE` and no server. Read `.env.local` for the number, and free the port rather than waiting for eve to move. **A follow-up request 404s or the agent seems to have forgotten everything.** Check that `WORKFLOW_POSTGRES_URL` is actually set and Postgres is reachable — without it, `eve` silently falls back to its local on-disk world under `.eve/.workflow-data`, which does not survive a container rebuild the way Postgres does. **`@workflow/world-postgres` fails at boot with a protocol or spec version error.** It is not pinned to an exact version somewhere. `eve` needs the `5.0.0-beta` line, so npm's `latest` (the 4.x line) is wrong — and so is the `beta` dist-tag, which now resolves past the release eve 0.30.8 can talk to. Pin `5.0.0-beta.32` exactly; see [Troubleshooting](/docs/troubleshooting) for the spec-version message. **The dashboard can't reach the database.** It reads `WORKFLOW_POSTGRES_URL` directly — same variable, same database, as the agent. If the agent works but the dashboard doesn't, the two processes likely have different `.env.local` files with different values. **Reverse-proxying either the agent or the dashboard.** Forward `/eve/` **and** `/.well-known/workflow/`, unrewritten. Forwarding only `/eve/` lets a session start, then stalls it forever — the workflow callback can't get back in. --- # Architecture > Why the dashboard is a SQL query, not an ingest pipeline — and where OpenTelemetry actually earns its keep. ## The discovery that shaped the dashboard `eve` tags every workflow run with framework-owned `$eve.*` attributes — `$eve.type` (`session` | `turn` | `subagent`), `$eve.parent`, `$eve.root`, `$eve.model`, `$eve.input_tokens`, `$eve.output_tokens`, `$eve.cache_read_tokens`, `$eve.tool_count`. eve's own docs say plainly: *these tags power the Agent Runs tab.* `@workflow/world-postgres` persists them to `workflow.workflow_runs.attributes` as JSONB — in **your own database**, the same one storing durable sessions. We verified this by hand: ```sql SELECT attributes FROM workflow.workflow_runs WHERE id = 'wrun_...'; -- {"$eve.type": "turn", "$eve.model": "openai/gpt-5-mini", -- "$eve.input_tokens": "4550", "$eve.output_tokens": "101", ...} ``` That means the dashboard's core view — session list, run tree, model, token totals — is a **plain SQL query against `packages/dashboard/lib/queries.ts`**, not an ingest pipeline with a second source of truth to keep in sync. ## Where OpenTelemetry still matters SQL gives you *that* a turn happened and its numbers. It doesn't give you prompt bodies or tool-call arguments — those exist only on OpenTelemetry spans. `agent/instrumentation.ts` exports them to the dashboard's `/api/ingest/v1/traces` endpoint as OTLP/HTTP JSON (not protobuf — the ingest route only parses JSON). The mere presence of `agent/instrumentation.ts` disables `eve`'s zero-config local trace spool (`.eve/traces/v1`), so `eve traces` stops working once you wire up the dashboard. Delete the file to get it back. ## Computed cost, not reported cost `eve` only attaches `gen_ai.usage.cost` to a span when the call went through Vercel's AI Gateway. A self-hosted agent calls its provider directly, so that attribute is simply absent — nothing upstream can tell you what a turn cost. `packages/dashboard/lib/pricing.ts` computes it from token counts against a table you can override with `EVESTACK_PRICING`. A model with no configured price is labeled `unpriced`, never silently rendered as free. ## Cooperative cancellation `POST /eve/v1/session/:id/cancel` returns `202` immediately, but the in-flight model call keeps running. We measured a cancelled turn continuing to stream for roughly 90 seconds, with the event order: ``` message.completed → step.completed → turn.completed → session.waiting → turn.cancelled → session.waiting ``` `turn.cancelled` arrives *after* a `session.waiting`, not before. Don't build a stop button that assumes silence follows a 202. ## Approvals ride the ordinary follow-up route There's no dedicated approve/deny endpoint. A pending tool approval resolves the same way an `ask_question` does — through the normal continuation route: ``` POST /eve/v1/session/:sessionId { "continuationToken": "...", "inputResponses": [{ "requestId": "...", "optionId": "approve" }] } ``` --- # The dashboard > The open replacement for Agent Runs — and it drives the agent, not just watches it. ## Observe - **Sessions** — every run on your machine, with title, status, and trigger - **Run tree** — turns and subagents nested by `$eve.parent`/`$eve.root`, with duration, input and output tokens, cached-read tokens, and tool count per turn - **Computed cost** — see [Architecture](/docs/architecture#computed-cost-not-reported-cost) for why it's computed rather than read from a span - **Integrations** — the live Composio catalog, connected accounts, one-click connect - **Monitors** — latency percentiles and failure rates over a rolling window, see below - **Alerts** — nine checks that ship on, and a webhook that speaks when one changes. See [Alerts](/docs/alerts) ## Monitors `/monitors` reports p50/p75/p95/p99 over a window you pick (1h, 6h, 12h, 24h, 7d), computed by Postgres with `percentile_cont` over the same `workflow.workflow_runs` the session list reads. Nothing is sampled and nothing is estimated. Two things it does deliberately, because the obvious version of each is wrong: **Turn latency leads, session duration is reported separately.** A `$eve.type = 'session'` row stays `running` for as long as the conversation is open, so a session someone left open overnight is a nine-hour "duration" that measures the human. Turns start and finish around one model exchange, so turn latency is the number worth alerting on. Session duration is still shown, over sessions that actually ended, and the two are never averaged together. **A failure is not just `status = 'failed'`.** A turn killed by a provider rate limit emits `turn.failed` on the stream while its workflow row still reads `status = 'completed'` — the workflow handled the error, so nothing failed as far as it is concerned. eve writes `$eve.model` only once a model call reports usage, so a finished turn without it never reached the provider. Monitors counts those as failures under **no model call** and shows them separately from `error_code` failures, because counting only the latter is the direction that flatters us. Unfinished turns are excluded from the distribution rather than counted as zero. That one is verified against a real server by `contract/runtime/probes/05-monitor-percentiles.probe.mjs`, with fixtures chosen so the wrong answer is numerically obvious — otherwise a busier agent would report as a faster one. ## Control This is where evestack diverges from Agent Runs, which is read-only: - **Start a session** and stream the reply straight from the dashboard - **Send a follow-up** to an existing session - **Resolve a pending approval** — see [Architecture](/docs/architecture#approvals-ride-the-ordinary-follow-up-route) for the actual protocol underneath - **Cancel a run** — read [the cooperative-cancellation note](/docs/architecture#cooperative-cancellation) before building anything that assumes a cancel is instant - **Promote a session to an eval** — turn any run, especially one that went wrong, into a real `evals/*.eval.ts` replaying its actual messages ([how to run one](#running-the-evals)) - **Replay a session into a new one** — re-send its messages with one turn rewritten. Read [Replaying a session](#replaying-a-session) first; it re-runs the original's tool calls ## Running the evals Save a promoted file into `evals/` in the agent project, then: ```bash npm run eval # the whole suite npm run eval -- my-eval-name # just the one you promoted ``` Not `npx eve eval`. eve boots its own dev server for a run, and it refuses to boot a second one for a project that already has one — so with `npm run dev` up, which is where the quickstart leaves you, a bare `npx eve eval` exits 1 with *A dev server is already running for this eve agent* and runs nothing. `npm run eval` hands eve the port this project recorded in `EVESTACK_AGENT_PORT` instead, and runs the same Postgres, schema and model checks that `npm run dev` runs first. A `--url` you pass yourself is always left alone. The evals drive a real agent through a real Docker sandbox, so Docker has to be running. They need no paid key: on the `$0` Ollama path the whole suite passes, and `evals/memory.eval.ts` needs the second, separate `ollama pull nomic-embed-text` — `npm run eval` checks for it by name rather than letting `remember` fail mid-run. ## Replaying a session **Replay into a new session**, on a session's page, re-sends that conversation's user messages into a fresh session, optionally rewriting one of them. It answers the question you always have after a bad run — *would it have worked if I had said it differently?* — without retyping the conversation and hoping you reproduced it. The durable event log holds every user message, so the dashboard rebuilds turns 1..N and sends them in order: the first creates the fork, and each later one waits for the fork to park before it is sent. You pick where to stop, and may rewrite that last turn. Turns after it are dropped — they answered a conversation that is no longer happening. ### It re-executes the tool calls, against the real world Replaying turn 5 means running turns 1 through 4 again, for real. Nothing is stubbed or simulated: an email the original sent is sent again, money it spent is spent again, a PR it opened is opened again, a file it deleted is deleted again. The panel is built around that one fact. It reads the per-turn tool list from the same transcript `GET /api/evals/promote/:id` reads, marks which turns are in range and which tools each of them ran, and keeps the run button disabled behind a checkbox that names those tools — a generic "this spends money" notice would not tell you that turn 3 called `send_email`. Changing the range clears the checkbox, because ticking it for turns 1–2 must not authorise turns 1–6. That list is what the original ran, not a promise about the replay. The model is free to take a different path this time and call something that is not on it. ### It is not a checkpoint fork LangGraph forks from a serialized checkpoint: you branch off saved state, edit it if you want, and the turns before the branch point never execute twice. eve's durable record is an event log, not a resumable state snapshot, and it exposes no checkpoint to branch from — so the only way to reach turn 5 here is to run turns 1 through 4 again. That is strictly weaker than a checkpoint fork and strictly more expensive, and it is what the durable model allows. What it replaces is not a cheap branch. It is a person pasting the same messages back into the chat box, which re-runs the same tools with no warning, no record of what it was forked from, and no chance to stop short. ### What it costs A fork is an ordinary session: it spends real tokens, and it shows up in Sessions with its own computed cost. It also takes about as long as the original conversation did, because each turn has to finish before the next is sent. Leaving the page stops the remaining turns from being sent; the ones already sent keep running. ### A partial fork is a normal outcome The replay stops at the first turn that does not land — it does not skip it and try the rest. When it stops, the fork already exists and already holds the turns that did land, so the route answers with the new session id, `turnsPlanned`, `turnsDelivered` and `complete: false` rather than an error, and the panel reports *Sent 2 of 4 turns* with the reason instead of claiming success. There is nothing to roll back: the turns that landed have already run their tools. | It stopped because | What that means | | --- | --- | | `awaiting_human` | a replayed turn parked on a tool approval or a question. The log records that a tool was denied, never why, so the replay will not answer for you — open the fork and answer it there | | `session_ended` | the replay diverged into a run that finished or failed, so there is no session left to send the next turn to | | `timeout` | a turn did not come back within 90s, or the replay used up its 240s budget. Nothing is lost — continue the fork from its own page | | `session_mismatch` | the agent started a different session instead of continuing the fork. That run is live and was not part of the replay; cancel it | | `agent_error` | the agent rejected the follow-up | The API takes the same view. `fromTurn` is required and never assumed — an empty body used to mean "replay everything", and the largest blast radius is the wrong thing to get by saying nothing. `GET` the same URL to see the turns, and the tools each one ran, without running any of them. ## Who approved it A browser button that can approve a shell command invites one fair question: approved by whom? eve cannot answer it — its human-in-the-loop protocol records that a request was answered and with which option, because that is all it needs to resume the turn. Identity is not part of it. So evestack records every decision itself, in `evestack.approvals`, and shows them under **Approvals**: the tool, the decision, the person, the session, and — the column that matters — *how* the identity was established. evestack deliberately ships no identity provider. The dashboard is meant to sit behind whatever you already trust, so it reads identity from the request and is explicit about its provenance: | Source | Recorded as | Worth | | --- | --- | --- | | `EVESTACK_APPROVER_HEADER` (a header you name) | `header` | as trustworthy as your proxy | | `X-Forwarded-User` / `X-Forwarded-Email` | `forwarded-user` / `forwarded-email` | set by oauth2-proxy, Cloudflare Access, and friends | | HTTP Basic user | `basic` | one shared credential — identifies a deployment, not a person | | nothing | `unidentified` | recorded anyway, and flagged | An anonymous decision is still written down. A silent gap in an audit log is worse than a visible one, and the Approvals page counts them at the top so they cannot be missed. Set `EVESTACK_REQUIRE_APPROVER=1` to refuse them outright — the API answers `403 approver_required` and the turn stays parked. The audit row is written *after* eve accepts the answer, so the log never claims a decision that did not take effect. Retention is unbounded: this is the row you want a year later, when someone asks why the agent deleted the thing it deleted. ## What the agent can be told to do eve advertises every skill in `agent/skills/` to the model and hands it a `load_skill` tool. Anything in that directory can put instructions into a live turn without a human seeing them first — which is a feature, and is also the reason the **Skills** page scans each one before listing it: credential reads, environment dumps, network exfiltration, and the shapes in between. `?selftest=1` runs the scanner against a bundled malicious specimen and expects a `critical` verdict, so you can tell a scanner that is working from one that is merely quiet. A clean verdict is not proof of safety and the page says so in as many words. It is pattern matching over files; a skill that fetches its instructions at runtime has nothing for it to read. **Point it at the right directory.** The page reads, in order: `EVESTACK_SKILLS_DIR`, then `/agent/skills`, then the skills bundled with the evestack template. In a container the first two do not exist unless you arrange them, so it falls through to the third — and because the bundled skill has the same name the scaffolder writes into your project, the page looks like it is reading yours when it is reading the image's own copy. It labels the source (`bundled template`, with the absolute path) and says what to do about it. Both compose files this project ships now arrange it for you: ```yaml environment: EVESTACK_SKILLS_DIR: /agent-skills volumes: - ./agent/skills:/agent-skills:ro ``` Measured against the published image, with and without those two lines: | | `resolvedBy` | reads | | --- | --- | --- | | without | `bundled-template` | `/repo/templates/default/agent/skills` — inside the image | | with | `env` | `/agent-skills` — your project, read-only | If you run the image by hand, pass both or the scanner reports a clean verdict about files nobody is running. ## Running it **In the project you already have.** `evestack create` writes a `dashboard` service into your project's `docker-compose.yml`, behind a profile, pointing at the published image. So there is nothing to clone and nothing to build: ```bash docker compose --profile dashboard up -d npm run verify # prints the URL and the sign-in pair ``` It reads the same `.env.local` your agent does, through `env_file:`, so the Postgres URL and the credential are already there. The port is whichever one the scaffolder published — it picks a free one, so it is not always 4000, and `npm run verify` is what tells you. This section used to open with "The dashboard is not part of a `create-evestack` project — clone the repository for it", followed by `git clone` and `cd evestack/packages/dashboard`. Both halves were wrong: the scaffolder has shipped that compose service for a while, and the clone recipe did not work either. Installing from `packages/dashboard` cannot resolve `workspace:*`, and `@evestack/schedules` has no `dist/` in a fresh clone — its `dist/` is gitignored and it has no `prepare` script — so `/schedules` fails to compile with `Can't resolve '@evestack/schedules/cron'`. Three other pages said the opposite of this one. **From a clone, if you are working on the dashboard itself.** Install and build from the repository root, not from `packages/dashboard`: ```bash git clone https://github.com/SammyTourani/evestack cd evestack pnpm install pnpm -r --if-present run build # workspace dist/, which a cold clone has none of cp packages/dashboard/.env.example packages/dashboard/.env.local pnpm --filter @evestack/dashboard dev # http://localhost:4000 ``` Both halves of `EVESTACK_AUTH_*` are required and neither is defaulted: with either missing the dashboard serves nothing usable. `503` on every request, with four exceptions — `PUBLIC_PATHS` in `lib/auth.ts` holds three paths and `proxy.ts` opens that tier only to `GET`, so `GET /signin` renders the reason with no sign-in form on it, `GET /api/auth/session` and `GET /api/auth/signout` reach POST-only routes and get a bare `405`, and `GET /api/health` reaches its own handler, which answers `503 {"status":"unconfigured"}` and so reports the container unhealthy. None of the four exposes anything, and there is still nothing to sign in with. Use the same pair your agent has in `.env.local` — it is one secret per deployment, which the dashboard signs you in with and then presents to the agent. Or as part of the full stack: `docker compose --profile full up -d` brings up Postgres and the dashboard together. The container runs as a non-root user; the process inside it binds `0.0.0.0` and what keeps it off your network is the compose port mapping, `127.0.0.1:4000:4000`. See [Self-hosting](/docs/self-hosting) before exposing it beyond your own machine. ## It reads your database, nothing else No API calls to evestack, no telemetry, no external service in the loop. Its required environment variables are `WORKFLOW_POSTGRES_URL` and the `EVESTACK_AUTH_*` pair — point the first at your Postgres and everything on the observe side is a SQL read. --- # Alerts > Nine monitors that ship on, and a delivery path that tells you when one changes — with the limits of an in-process notifier stated rather than hidden. `/monitors` has always computed nine checks. Until now it only ever showed them, which made them a dashboard rather than an alert: they existed for the length of one render, and the moment you needed them was the moment nobody was looking. This page is the other half — the part that speaks first. ## Turn it on One variable. Set it to a Slack incoming webhook, a Discord webhook, or any HTTPS endpoint you control: ```bash EVESTACK_ALERT_WEBHOOK_URL=https://hooks.slack.com/services/T0/B0/xxxx ``` The payload shape is picked from the hostname — `hooks.slack.com` gets `{"text": …}`, a Discord webhook path gets `{"content": …}`, anything else gets a structured JSON body. Override it with `EVESTACK_ALERT_WEBHOOK_FORMAT=slack|discord|webhook` if you post through a proxy. ### Which file it goes in Worth one table, because the variable has to reach the **dashboard's** process and the three ways of running it read three different files: | How you run the dashboard | Put it in | How it arrives | | --- | --- | --- | | A project from `npm create evestack` | `.env.local` | The generated `docker-compose.yml` gives its dashboard service `env_file: .env.local`, so every name in that file is set inside the container. `.env.example` lists all six, commented out. | | This repository's own `docker-compose.yml` | `.env` beside it | That file has **no** `env_file:`, so a name reaches the container only by being named in the dashboard service's `environment:` block. All six are. | | `npm run dev` in `packages/dashboard` | `packages/dashboard/.env.local` | Next reads it directly. | Until 2026-08-19 the repository's own `docker-compose.yml` named none of the six, and that is the failure mode this page is least able to help with. Setting `EVESTACK_ALERT_WEBHOOK_URL` in the `.env` beside it did nothing at all: Compose read the file, matched the name against no `${…}` anywhere, and dropped it — while `/monitors` went on saying no delivery target was configured, to someone who had just configured one. Scaffolded projects were never affected, because their compose has `env_file: .env.local` — but their `.env.example` did not mention these variables either, so there was nowhere a user could have discovered them. Restart the dashboard and it says which way it went, once, at boot: ``` [evestack:alerts] delivering transitions — 1 sink(s), every 60s ``` Unset, it says the other thing, and `/monitors` says it too rather than leaving you to assume. That is deliberate: a page showing nine checks under the word **firing**, with nothing stated about delivery, invites exactly one conclusion. Press **Send a test** on `/monitors` to prove the channel. It sends one obviously-synthetic message and does not touch any monitor's state — a test that consumed a real transition could swallow the alert it was meant to prove. ## What gets sent, and what does not A message goes out when a monitor **changes**, never merely because it is still bad. An integration that re-sends every minute is muted within a day, at which point it is worth less than nothing: it has trained you to ignore the channel the real one will arrive in. | From | To | Sent | | --- | --- | --- | | never seen | firing | **yes** — a fresh install that is already broken is worth one first message | | never seen | ok / not checked | no — nine "all fine" messages on first boot is how a channel gets muted | | ok | firing | **yes** | | not checked | firing | **yes** | | firing | ok | **yes**, the all-clear | | firing | not checked | **yes**, as *not checked* — see below | | ok | not checked | only for `page` severity | | not checked | ok | no | **`firing → not checked` is never reported as resolved.** The monitor stopped being answerable while it was firing, so the problem is un-observed, not over. Sending "resolved" there would stand an operator down in the middle of a live incident, and it is the single most dangerous thing this path could do. **`ok → not checked` pages only for `page`-severity checks.** For the two that would wake someone, losing the ability to evaluate *is* the incident. Below that it is silent, because an unmounted Docker socket is a configuration choice, not news. To repeat a still-firing alert, set `EVESTACK_ALERT_RENOTIFY_MINUTES=60`. It is off by default. ## The spend monitor, and which cap it uses `EVESTACK_ALERT_DAILY_SPEND_USD` is the install-wide daily spend the alert is judged against. With it unset, the alert falls back to `@evestack/budget`'s `EVESTACK_BUDGET_DAILY_USD` (default `10`) so that it works out of the box — and it **says so in the alert text**, because the two are not the same measure. The budget cap is per principal; this alert sums every priced turn on the install. On the single-user install that self-hosting usually means they are the same number; with two users the install crosses $10 while neither person is near their own limit. This monitor spent its entire existence stuck at *not checked*. It read `EVESTACK_DAILY_BUDGET_USD` — the same four words in a different order from the variable that exists — so on every install ever made it reported "no cap is configured" while `@evestack/budget` enforced its $10/day default a process away. A check that cannot fire is worse than a missing one, because the page counts it among the ones that passed. Set it to `false` to switch spend alerting off without touching your budget caps. ## Signing, for the generic sink Slack and Discord authenticate by secret URL. For your own endpoint, set a shared secret and every POST is signed: ```bash EVESTACK_ALERT_WEBHOOK_SECRET=$(openssl rand -hex 32) ``` ``` x-evestack-alert-timestamp: 1786232997598 x-evestack-alert-signature: sha256=d0b4cec6bb0de… ``` The signature is `HMAC-SHA256(secret, ".")`. The timestamp is inside the signed material on purpose — signing the body alone authenticates the content and nothing else, so a captured POST replays forever. Reject anything older than your own tolerance. ## Honest limits - **If the dashboard is down, nothing is delivered, and nothing here can tell you so.** That is a genuine limit of any in-process notifier, not an oversight. The payload carries `sentAt` so a receiver that cares can alert on silence — the one check that has to live outside the thing it watches. - **The dashboard is an optional compose profile** (`profiles: ["full", "dashboard"]`). `docker compose up` without a profile starts Postgres and the agent and no notifier at all. - **It is deliberately not delivered through your agent's channel.** evestack already has a path that reaches a human — the [heartbeat](/docs/proactive) — and it is the wrong one for this. Three of the nine monitors (`wedged`, `no_spans_while_active`, `turn_failure_rate`) fire exactly when the agent is unwell, so routing them through the agent means the turn that wedges is the turn that would have told you. The dashboard is a separate process and stays up when the agent does not. - **Delivery is at-least-once, not exactly-once.** A transition is only marked delivered once a sink returns 2xx, so a receiver that is down replays it on the next tick rather than losing it. A receiver that accepts the POST and then drops it is indistinguishable from one that delivered it — that is the correct place to give up. - **Delivery is tracked per sink, so one broken channel does not spam the others.** Each sink has its own record of what it has been told. A sink that is down is retried until it accepts; a sink that is healthy is told once and then left alone. The first version of this required *every* sink to accept before a transition counted as delivered — which sounds careful and meant that a single permanently-broken webhook re-sent the identical alert to every healthy channel on every tick, forever. That is the muted channel this whole page is about, arrived at by way of being careful about the other thing. - **Two dashboards will not normally double-send, and the exception is named.** A short lease row in `evestack.alert_lease` gives one instance the tick; it expires rather than being released, so a crash mid-send is picked up by the other one interval later. Within a single dashboard, deliveries are serialised outright, so a timer tick and a `POST /api/alerts` cannot overlap. What the lease does *not* cover: it is granted on elapsed time and does not check who holds it, so if one instance's sends take longer than the whole interval — several sinks, each slow but healthy — the other can take over while the first is still posting, and one transition goes out twice. Configure a sink list whose total latency fits inside `EVESTACK_ALERT_INTERVAL_SECONDS`, or run one dashboard. A duplicate page is the worst case here; nothing is lost. ## What it writes Three tables, all in the `evestack` schema, created on first delivery and never before — a feature that is off costs nothing, including a table. | Table | What it holds | | --- | --- | | `evestack.alert_state` | one row per monitor: what was last **seen**, what was last **delivered**, and what each sink has acknowledged | | `evestack.alert_deliveries` | every POST, whether it worked, the status, and how long it took. Pruned at 30 days | | `evestack.alert_lease` | which instance is sending | The two state columns are separate on purpose. The decision to send compares the live state against what was last *delivered*, so a webhook returning 500 does not consume the transition. If they were one column, a single failed POST would mark the alert handled and the page that started failing at 02:00 would never be mentioned again. Stored URLs are truncated to their origin. Slack and Discord both put a working credential in the path, so the full URL is never written to a table the dashboard renders. ## Reading it from somewhere else `GET /api/alerts` returns all nine with their state, plus whether delivery is configured and when it last ran. `POST /api/alerts` runs a tick now; `POST /api/alerts?test=1` sends the test message. All three take the dashboard's ordinary credentials, including HTTP Basic: ```bash curl -u "$EVESTACK_AUTH_USER:$EVESTACK_AUTH_PASSWORD" http://localhost:4000/api/alerts ``` --- # Observability > eve keeps the session tree in workflow run tags, not on OTel spans — so evestack reads it back out of your own Postgres with SQL, and uses OpenTelemetry only for what SQL cannot hold. ## What you lose the moment you leave Vercel `eve` observes an agent through two independent systems, and only one of them travels over OpenTelemetry. The first is **workflow run tags**. eve stamps every session, turn and subagent run with reserved `$eve.*` attributes — the run's type, its parent, its root, its model, its running token counts. They are framework-owned, emitted whether or not you author an `instrumentation.ts`, and authored code cannot set or override the `$eve.` namespace. eve's own instrumentation guide is explicit about where they live: on the workflow run, and *"queryable in the Workflow dashboard, not on OTel spans"*. The second is the **OpenTelemetry export** you configure in `agent/instrumentation.ts`. That is the only mechanism eve offers for getting telemetry off the machine, and it is what eve's docs point at for Braintrust, PostHog, Arize, Honeycomb, Datadog and Jaeger. Those two facts compose badly for a self-hoster. The tags power Vercel's **Agent Runs** tab, which is a Vercel-only surface — the platform detects `eve` as the framework and renders the view under your project's Observability tab, currently gated per team. Export OTLP to Jaeger instead and you get spans: a flat waterfall of model calls and tool calls, correct on timing, with no session, no turn boundary, no subagent lineage and no token rollup. The structure was never on the wire. This is the gap evestack closes, and the scope is worth stating before the mechanism. This is not a provider-neutral answer to local agent observability, and it is not an OpenTelemetry technique at all. It works because you run `@workflow/world-postgres` and therefore own the table eve's tags land in — see [Why this works, and where it stops](#why-this-works-and-where-it-stops). ## The tags are already in your database `@workflow/world-postgres` persists a run's attributes to `workflow.workflow_runs.attributes` as JSONB. Nothing extra is configured; storing them is what the world does. A real turn row from the verified build: ```json {"$eve.root":"wrun_…","$eve.type":"turn","$eve.model":"openai/gpt-5-mini", "$eve.parent":"wrun_…","$eve.tool_count":"11","$eve.input_tokens":"4550", "$eve.output_tokens":"101","$eve.cache_read_tokens":"2176","$eve.cache_write_tokens":"0"} ``` So the data behind Agent Runs is sitting in the same database that holds your durable sessions, in a column Postgres can index and query. The dashboard reads it directly — `packages/dashboard/lib/db.ts` opens one pool against `WORKFLOW_POSTGRES_URL`, and `packages/dashboard/lib/queries.ts` is the whole of the reconstruction. There is no ingest pipeline behind the session list, no second store, and nothing to keep in sync. ### The keys evestack reads Verified against eve 0.30.6's own emitters (`buildSessionAttributes`, `buildTurnAttributes`, `buildSubagentRootAttributes`, and the per-step usage write in the tool loop) and pinned by `contract/contracts/06-run-attributes.contract.mjs`. | Key | On | What the dashboard does with it | | --- | --- | --- | | `$eve.type` | every agent run | `session` \| `turn` \| `subagent`. Every query filters on it | | `$eve.parent` | turn, subagent | Edge to the immediate parent session | | `$eve.root` | turn, subagent | Edge to the root of the chain — groups a whole tree | | `$eve.subagent` | subagent | Compiled graph node id, so a subagent run gets a name | | `$eve.title` | session | Session label, truncated from the first user message | | `$eve.trigger` | session, subagent | Which channel kind started it | | `$eve.model` | turn | Model id, and the input to cost | | `$eve.input_tokens` | turn | Cumulative, last write wins | | `$eve.output_tokens` | turn | Cumulative, last write wins | | `$eve.cache_read_tokens` | turn | Billed at the cache rate, not the input rate | | `$eve.cache_write_tokens` | turn | Shown per turn; not in eve's published tag list | | `$eve.tool_count` | turn | Tools **available** to the turn, not tools invoked | eve 0.30.6 emits four more that the dashboard does not read: `$eve.channel_request_id`, `$eve.parent_call`, `$eve.parent_turn`, and `$eve.cost_usd` — the last of which is written only when a Vercel AI Gateway response reported a cost, so a self-hosted run never has it. See [Cost is computed, never reported](#cost-is-computed-never-reported). ### The query Simplified for reading — the real one in `queries.ts` uses a `LEFT JOIN LATERAL` and aggregates `(model, tokens)` tuples so each turn can be priced against its own model. The shape is faithful: ```sql SELECT s.id, s.status, s.attributes ->> '$eve.title' AS title, s.attributes ->> '$eve.trigger' AS trigger, COUNT(*) FILTER (WHERE t.attributes ->> '$eve.type' = 'turn') AS turns, SUM((t.attributes ->> '$eve.input_tokens')::bigint) AS input_tokens, SUM((t.attributes ->> '$eve.output_tokens')::bigint) AS output_tokens, SUM((t.attributes ->> '$eve.cache_read_tokens')::bigint) AS cache_read_tokens FROM workflow.workflow_runs s LEFT JOIN workflow.workflow_runs t ON t.attributes ->> '$eve.root' = s.id OR t.attributes ->> '$eve.parent' = s.id WHERE s.attributes ->> '$eve.type' = 'session' GROUP BY s.id, s.status, s.attributes, s.created_at ORDER BY s.created_at DESC; ``` The `$eve.type` filter is load-bearing, not tidiness. One user message writes **three rows**: the session (`workflow//eve//workflowEntry`), the turn (`workflow//eve//turnWorkflow`), and an internal `workflow//eve//sessionTimeoutWorkflow` run carrying **no** `$eve.type` at all. Drop the filter and that third row is counted as a user session — inflated counts, phantom rows, and per-session cost divided across runs that never called a model. That same message produced **38 events** in Postgres and **zero** files on disk, which is how we confirmed the Postgres world was live rather than silently falling back to the local file world. ## Why this works, and where it stops This is not a general OpenTelemetry technique. It works for one reason: **you run the Postgres world**, so a table you own holds the attributes. Stated precisely: - It depends on `@workflow/world-postgres` writing `workflow.workflow_runs.attributes`. On the default local file world the same attributes exist, but in `.eve/.workflow-data/runs/*.json` — durable, and not SQL. - It depends on eve's `$eve.*` names, which are framework-owned and can change. They are string literals inside SQL, invisible to `tsc`, and a rename returns `NULL` rather than raising — the dashboard would render a confident zero. That is why contract 06 exists and why it is the contract with the least tolerance for drift. - It is not provider-neutral. It says nothing about observing an agent built on some other framework, or on eve with a third-party world. ## Tier two: OpenTelemetry, for what SQL cannot hold `workflow_runs` records that a turn happened, which model ran, and how many tokens it burned. It records nothing about content. The system prompt, the message history, the arguments the model passed to `bash` and what came back exist only on spans. | | Tier 1 — `workflow.workflow_runs` | Tier 2 — `evestack.spans` | | --- | --- | --- | | Written by | `@workflow/world-postgres` | your agent's OTLP exporter | | Configuration | none — always on | `EVESTACK_DASHBOARD_URL` + `EVESTACK_INGEST_TOKEN` | | Session / turn / subagent tree | yes | partial (see below) | | Model id and token counts | yes | yes | | Tools available to the turn | yes | **no** | | Duration and status | yes, per run | yes, per span | | System prompt, message history | **no** | yes | | Tool arguments and results | **no** | yes | | Safe to drop | no — durable session state | yes, it is replayable telemetry | The two schemas are separated on purpose. `workflow` belongs to world-postgres and evestack only ever reads it; everything evestack writes goes to `evestack`. An eve upgrade cannot collide with these tables, and `DROP SCHEMA evestack CASCADE` costs telemetry and never a durable session. The exporter is one file, and the template ships it: ```ts title="agent/instrumentation.ts" const INGEST_TOKEN_HEADER = "x-evestack-ingest-token"; export default defineInstrumentation({ setup: ({ agentName }) => { if (!endpoint) return; registerOTel({ serviceName: agentName, traceExporter: new OTLPHttpJsonTraceExporter({ url: endpoint, ...(ingestToken ? { headers: { [INGEST_TOKEN_HEADER]: ingestToken } } : {}), }), }); void reportIngestCredential(endpoint, ingestToken); }, recordInputs: process.env.EVESTACK_TRACE_CONTENT !== "off", recordOutputs: process.env.EVESTACK_TRACE_CONTENT !== "off", }); ``` With `EVESTACK_DASHBOARD_URL` unset it registers nothing, so an agent running without the dashboard pays no exporter cost and fails no exports. Set it to the **full path** — `http://localhost:4000/api/ingest/v1/traces`. `@vercel/otel` uses the `url` verbatim and appends nothing, so a bare origin or the conventional `:4318/v1/traces` never arrives. ### The credential, and why a 401 here is invisible The ingest route is the one route a human is not the caller for, so it has its own shared secret: `EVESTACK_INGEST_TOKEN`, sent as `x-evestack-ingest-token` and compared in constant time by `ingestAuthorized()` in `packages/dashboard/lib/auth.ts`. The agent and the dashboard must hold the **same** value. `create-evestack` generates it into `.env.local`, which the dashboard container also reads through `env_file:`, so a scaffolded project needs no extra step. Leaving `EVESTACK_INGEST_TOKEN` unset on **both** sides does not give you an open endpoint — it gives you a broken one. With no token configured the route falls back to the dashboard's ordinary session auth, and an OTLP exporter has no cookie and no password, so every span is refused with `401`. That refusal is silent unless something goes looking for it, which is why the template probes the endpoint once at boot. `@vercel/otel`'s exporter is a `fetch` whose `.then` branch calls the success callback and whose `.catch` branch handles errors — and an HTTP 401 *resolves* a fetch. So a rejected batch is reported to the batch processor as `ExportResultCode.SUCCESS`, the spans are dropped, and nothing retries. The status reaches `diag.debug`, and `@vercel/otel` installs a diag logger at all only when `OTEL_LOG_LEVEL` is set. Without the probe, a wrong token and an idle agent produce exactly the same empty Traces tab. Use `OTLPHttpJsonTraceExporter`, not `OTLPHttpProtoTraceExporter`. eve's own `instrumentation/jaeger` registry item uses the protobuf one, because Jaeger takes protobuf. The ingest route parses JSON only and rejects protobuf by content type with `415` and a message naming the fix, rather than half-decoding a span into the table. Nothing is stored on that path. Spans land in `evestack.spans`, keyed `(trace_id, span_id)` so an at-least-once retry upserts instead of duplicating. Read them through `packages/dashboard/lib/traces.ts` — `listModelCalls`, `listToolCalls`, `getSpanTree` — or with SQL directly: ```sql SELECT attributes ->> 'gen_ai.tool.name' FROM evestack.spans WHERE name LIKE 'execute_tool %'; ``` ## Exporting changes the vocabulary you receive This is the sharpest edge in the whole trace tier, it is not documented upstream, and it cost us a working feature before we found it. eve has **two** telemetry emitters: ``` agent.* / ai.prompt.* eve's own `eve.agent` tracer gen_ai.* / ai.settings.context.eve.* the vendored AI SDK exporter ``` Only the second ever leaves the machine. `createAgentOtelInstrumentation()` — the source of the rich `agent.session` / `agent.turn` / `agent.step` / `agent.action` family — has exactly one caller, `installLocalInstrumentationRuntime()`, and eve's dev host installs that runtime *only* when the app authors no `agent/instrumentation.ts`: ```js compiledArtifacts.instrumentationPluginPath === void 0 && plugins.unshift(… local-tracing-runtime-plugin.ts) ``` So authoring instrumentation to export anywhere silently opts you out of the `agent.*` span family. Measured in our own database: of **5,542** exported spans, three carried a session id — and those three had been posted by hand out of the local spool during an experiment. Teaching the schema's generated columns the AI SDK vocabulary took that to 99, and prompts resolved for the first time. Three corrections to what we believed first, all from re-verifying rather than trusting the first pass: - **Not a 0.30.x regression.** The gate is byte-identical in every tarball from 0.29.5 to 0.30.6. There was never a working "before". - **Prompts are not lost when exporting.** They arrive under the AI SDK's conventions instead — `gen_ai.input.messages` on `chat ` spans, `gen_ai.tool.call.arguments` and `gen_ai.tool.call.result` on `execute_tool `. The dashboard reads both vocabularies, so either shape resolves. - **The root session id has no exported counterpart.** That is the one thing an exporting deployment genuinely cannot reconstruct from spans, so subagent traces cannot be stitched to their parent. Tier 1 has the lineage regardless, via `$eve.root`. `contract/contracts/14-telemetry.contract.mjs` pins that coupling. Read a failure there as good news: it most likely means eve made the agent instrumentation reachable from authored instrumentation, which is the fix we want. ### The local spool's span tree For completeness, this is the shape eve's zero-config spool produces — the one you see from `eve traces`, and the one you get in `.eve/traces/v1` when there is **no** `instrumentation.ts`: ``` agent.session (ROOT) └── agent.turn ├── agent.step │ └── ai.streamText │ └── ai.streamText.doStream └── agent.turn.terminal ``` Do not expect that tree in a collector. It is exactly the family the export gate withholds. Adding `agent/instrumentation.ts` disables eve's zero-config trace spool, so `eve traces` and the dev TUI's `/traces` viewer stop working. eve hands telemetry to authored instrumentation and does not run its own writer alongside it. Delete the file to get them back. The trade is deliberate. The spool's retention is bounded — 7 days, 512 MB, 20 traces, tunable with `EVE_TRACES_MAX_AGE_MS`, `EVE_TRACES_MAX_TOTAL_BYTES` and `EVE_TRACES_RETAIN_COUNT` — and it is local to one machine and one process. `evestack.spans` has no bound and survives `eve dev` exiting. ## Cost is computed, never reported eve attaches `gen_ai.usage.cost` to a span, and `$eve.cost_usd` to a run, only when the call was served by Vercel's AI Gateway and the gateway's own metadata carried a price. A self-hosted agent calls its provider directly, so both are simply absent. Token *counts* are always there, which makes price × tokens the only path to a dollar figure. `packages/dashboard/lib/pricing.ts` does that arithmetic: - Cached reads are billed at the cache rate and subtracted from the input total, because eve reports them *inside* it. Billing them twice is the obvious bug here. - `cacheRead` defaults to 10% of the input rate when a table entry omits it. - The built-in table goes stale — providers reprice, and a wrong table silently reports wrong money. Override it without editing the file: ```bash EVESTACK_PRICING='{"openai/gpt-5-mini":{"input":0.25,"output":2,"cacheRead":0.025}}' ``` - A prefix wildcard covers a family: `ollama/*` is priced at zero, which is the honest number for a model running on your own hardware. A model with no configured price contributes `0` to the total and is labelled **`unpriced`** in the UI, never silently counted as free. An unpriced model must never look cheap. ## Privacy ```bash EVESTACK_TRACE_CONTENT=off ``` Sets `recordInputs` and `recordOutputs` to `false`, so no prompt body, message history or tool result is recorded on a span. You keep span timing, model ids and token counts, and you keep the whole of tier 1 — the session tree, tokens and cost are unaffected, because they never came from spans in the first place. Leaving traces off entirely is also a supported configuration: unset `EVESTACK_DASHBOARD_URL` and the dashboard's sessions, turns, tokens, cost and approvals all still work. Neither tier reports anything to evestack or to any third party. Dropping tier 2 drops `EVESTACK_INGEST_TOKEN` with it; `WORKFLOW_POSTGRES_URL` and the `EVESTACK_AUTH_*` pair are required either way. ## Two failure modes worth knowing before you trust a badge ### A failed turn still records `status = 'completed'` eve's event stream emits `turn.failed`, but the workflow row disagrees: the workflow *handled* the error, so as far as it is concerned nothing failed. Trusting `status` alone paints a green badge on a turn that produced nothing. eve writes `$eve.model` and the token tags only once a model call reports usage, so their absence on a finished turn is the surviving evidence that the call never landed. `queries.ts` exposes that as `noModelCall` — a turn of `$eve.type = 'turn'`, with `completed_at` set, and no `$eve.model`. The run tree renders those turns as *no model call — turn produced nothing* alongside the `completed` status eve reported, rather than letting the status stand alone. ### Cancellation is cooperative `POST /eve/v1/session/:id/cancel` returns `202` immediately, but the in-flight model call keeps streaming. We measured roughly **90 seconds**, and `turn.cancelled` arrives *after* a `session.waiting`: ``` message.completed → step.completed → turn.completed → session.waiting → turn.cancelled → session.waiting ``` Timing and token counts recorded during those 90 seconds are real spend. Don't build a stop button that assumes silence follows a 202. The short version of this page, plus the approval protocol and the cancellation event order. What the run tree actually renders, and the parts that drive the agent rather than watch it. Running the agent and the dashboard off a laptop, including what the reverse proxy must forward. --- # Composio auth > One browser flow, 1,000+ apps — and the third party that holds your OAuth tokens. ## What you get [Vercel Connect](https://vercel.com/docs/connect) ships four **Vercel Managed Connectors** — Slack, GitHub, Snowflake, Salesforce — where Vercel registers the OAuth client and you just authorize it. It also ships two **Customer Managed Connector** types that cover everything else: **Custom OAuth** ("OAuth 2.0 / OIDC against any service URL you provide", both the authorization-code and client-credentials flows) and **API key**. Those need you to register the OAuth client or generate the key yourself and hand Vercel the credentials. eve separately connects to arbitrary [MCP servers](https://eve.dev/docs). This page used to say "four managed connectors" and stop, which understated Connect considerably — the coverage limit is not the four apps, it is who does the OAuth-client registration. evestack wires in [Composio](https://composio.dev) instead: **1,000+ toolkits**, including Gmail (61 tools) and GitHub (871 tools) — both of which connect in a single click. **Why "1,000+" and not an exact figure.** We counted the live catalog against `GET /api/v3/toolkits` and got 1,070 on 2026-08-04; Composio's own public directory listed 1,069 five days later. It moves by ones as toolkits are added and retired, so every surface in evestack — this page, the README, the site, the scaffolder's prompt, the dashboard's Integrations page — says **1,000+**, which stays true. The dashboard shows the live number: its *Apps available* tile reads whatever `countToolkits()` returns from your own key. **The catalog and the one-click subset are two different numbers, and attaching the first to the second is the mistake this page exists to prevent.** Measured against the live catalog on 2026-08-19: **1,326** toolkits, of which **121** publish a Composio-managed OAuth client. Those 121 connect with one click and nothing else. A further 96 support OAuth but require you to register your own OAuth client — the exact registration burden described at the top of this page — and 1,126 authenticate with an API key you go and fetch yourself. So a sentence of the form "one-click OAuth into 1,000+ toolkits" is false: it attaches the catalog size to a capability 121 of them have. Say "1,000+ toolkits" about the catalog, and "one click" only about the managed-OAuth subset. The apps the scaffolder offers by name — Gmail, Slack, Notion, Linear, Google Calendar, GitHub — are all inside the 121, so the path most people take really is one click. Re-measure by counting `composio_managed_auth_schemes` containing `OAUTH2` over `GET /api/v3/toolkits`. **Composio is a hosted third party, and this is the one place evestack's "everything runs on your network" stops being true.** Composio performs the OAuth flow and holds the resulting access and refresh tokens for every account you connect — your Gmail, your GitHub, your Slack. That data lives on Composio's infrastructure, not yours, and is subject to their terms rather than your Postgres backup policy. It is genuinely opt-in: with `COMPOSIO_API_KEY` unset the agent still boots with its sandbox, memory and durable sessions, and simply has no connectors. If token custody is the thing you are self-hosting to avoid, leave it unset. ## How it reaches the model The agent doesn't get ten thousand individual tools. It gets a handful of meta-tools — `COMPOSIO_SEARCH_TOOLS`, `COMPOSIO_GET_TOOL_SCHEMAS`, `COMPOSIO_MANAGE_CONNECTIONS`, `COMPOSIO_MULTI_EXECUTE_TOOL` — and Composio's Tool Router resolves the real set per session. Which ones it returns is the router's decision, not a guarantee: `COMPOSIO_MULTI_EXECUTE_TOOL` covers both single and batch calls, and `COMPOSIO_EXECUTE_TOOL` is a legacy slug the current router no longer returns. `agent/tools/composio.ts` in the template is two lines of code, under a comment block explaining why: ```ts import { composioTools } from "@evestack/composio"; export default composioTools(); ``` With `COMPOSIO_API_KEY` unset, this resolves to no tools and logs one line. The agent still boots, still has its sandbox and memory — you just don't get the connectors. Nothing about Composio being unreachable holds up the rest of the stack. ## The tradeoff, stated plainly Composio's managed OAuth is genuinely one-click, and that convenience has a real cost: - The consent screen reads **"Composio wants to access your account,"** not your app's name - Rate limits are **shared across every Composio customer** on the managed app - A **15-minute minimum polling interval** applies to managed auth That's the right tradeoff for personal use, prototypes, and internal tools. Before you onboard other people's accounts in production, read Composio's docs on registering your own OAuth app — it removes all three limits at the cost of the one-click setup. ## The dashboard's integrations page `http://localhost:4000/integrations` lists the live catalog, what's already connected, and a one-click **Connect** button per app — built directly against the Composio API, not a static list. --- # Long-term memory > Free semantic recall on the Postgres you're already running — and the indexing bug we hit building it. ## Why this is free The Postgres storing your durable sessions is already running, and `docker-compose.yml` uses the `pgvector/pgvector` image — so the storage side of long-term memory costs one `CREATE EXTENSION` and zero new containers. The agent gets `remember` and `recall` tools out of the box (`agent/tools/remember.ts`, `agent/tools/recall.ts`). ``` "Remember this for the future: my production database runs on port 5433 and I always deploy on Fridays." ``` ...saved in one session, comes back correctly in a completely fresh one: ``` "Your production database runs on port 5433, and you deploy on Fridays." ``` ## It needs an embeddings provider, and Anthropic is not one Storage is free; turning text into a vector is not automatic. `lib/memory.ts` resolves the embedding model from your chat provider, and the three providers do not all have one: | `EVESTACK_PROVIDER` | Embeddings | Model, vector width | | --- | --- | --- | | unset, or `openai` | yes, same key | `text-embedding-3-small`, 1536 | | `ollama` | yes, but **a second `ollama pull`** | `nomic-embed-text` (274 MB), 768 | | `anthropic` | **none — Anthropic has no embeddings endpoint** | borrows OpenAI's if `OPENAI_API_KEY` is set | On `anthropic` with no `OPENAI_API_KEY`, `remember` and `recall` do not work. The first call fails with the fix in it: ``` EVESTACK_PROVIDER=anthropic has no embeddings endpoint, so long-term memory needs one from somewhere else. Either set OPENAI_API_KEY, or run embeddings locally with EVESTACK_EMBED_PROVIDER=ollama (then `ollama pull nomic-embed-text`). ``` That is deliberately a tool failure and not a boot failure — the agent, its sandbox and its durable sessions all still work, because a project that cannot do embeddings is an ordinary configuration rather than a broken one. `npm run verify` says the same thing before you find out from a tool call. On `ollama`, the embedding model is a **separate pull from the chat model**: ```bash ollama pull nomic-embed-text ``` Miss it and `remember` fails on a model Ollama does not have. `npm run verify` checks for it by name. ## The three variables Set nothing here and embeddings follow `EVESTACK_PROVIDER`, which is right almost always. The overrides exist for the Anthropic case above, and for a different embedding model: | Variable | Default | What it does | | --- | --- | --- | | `EVESTACK_EMBED_PROVIDER` | follows `EVESTACK_PROVIDER` | `openai` or `ollama`. This is the one that makes memory work on the Anthropic path | | `EVESTACK_EMBED_MODEL` | `text-embedding-3-small` / `nomic-embed-text` | the embedding model to call | | `EVESTACK_EMBED_DIMENSIONS` | 1536 / 768 | the vector width, which must match what that model returns | Model and width go together, and the width is baked into `evestack.memories` at creation. Change either one and the old rows are unusable — vectors from two different models are not comparable — so the table has to be dropped: ```bash docker compose exec postgres psql -U evestack -d evestack \ -c 'DROP TABLE evestack.memories' ``` `lib/memory.ts` checks the existing column width on the first memory call and says exactly this, with both numbers, rather than letting the mismatch surface later as pgvector's `expected 1536 dimensions, not 768` from inside an INSERT. ## HNSW, not IVFFlat — a correctness fix, not a preference We built this with an IVFFlat index first, because it's the more commonly recommended default. It silently broke memory. IVFFlat assigns vectors to `lists` centroids **at index-build time**. Build it on an empty table — which any bootstrap migration must, since memory starts with nothing in it — and the centroids are meaningless. The planner switches to the index once a query looks selective enough, probes a near-empty list, and returns **zero rows** for a query that plainly should match. We reproduced this exactly: the same `recall()` call returned **2 results at `LIMIT 3`** and **0 results at `LIMIT 20`**, purely because the query plan flipped between them. ```sql -- lib/memory.ts uses this, not ivfflat: CREATE INDEX memories_embedding_idx ON evestack.memories USING hnsw (embedding vector_cosine_ops) ``` HNSW builds a navigable graph incrementally and needs no training data — it's correct from the first row, which is what an empty-by-definition memory table needs. ## Tags and filtering `remember` accepts optional tags; `recall` can filter by them. The agent chooses its own tags — in testing, saving a deployment preference produced `["preference", "devops", "deployment"]` without being told to. --- # Proactive agents > An agent that wakes up on its own and messages you first — with budgets, approvals and an audit trail. An agent that only answers when spoken to is a chatbot. An agent that wakes up, checks on things, and messages you when something needs you is a different product — and it is the one that made personal AI assistants the most-installed category of 2026. It is also the one that made them the most notorious. That category shipped with agents holding shell access and long-lived credentials, no spend limit, no record of what was approved, and in one widely-reported case a skills marketplace where the most-downloaded package was an infostealer. The capability was never the problem. The absence of everything around it was. evestack ships the same capability with the governance attached. This page is the recipe. ## The five pieces Everything here already exists in the default template. Turning the heartbeat on is the only step that is off by default. | | What it does | Where | | --- | --- | --- | | **Heartbeat** | wakes the agent on a cron, and stays quiet unless there is news | `agent/schedules/heartbeat.ts` | | **A channel** | where it reaches you — Telegram is the fastest to finish | `agent/channels/` | | **Memory** | so it remembers what it noticed last time | `@evestack/memory` | | **Budget caps** | a hard ceiling on what unattended work can cost | `@evestack/budget` | | **A gated tool** | anything destructive parks for a human, and the decision is recorded | `approval: always()` | ## Turn it on ```bash # 1. a channel it can reach you on (Telegram is a two-minute setup — see /docs/channels/telegram) TELEGRAM_BOT_TOKEN=123456789:AAH... TELEGRAM_WEBHOOK_SECRET_TOKEN=... # 2. the heartbeat itself EVESTACK_HEARTBEAT_CHANNEL=telegram EVESTACK_HEARTBEAT_TARGET={"chatId":123456789} EVESTACK_HEARTBEAT_CRON=0 * * * * ``` Then write what it should check. `HEARTBEAT.md` sits at the root of the project and is read at every fire, so editing it takes effect on the next wake-up — no restart, no redeploy. ```md ## Checks - Look through my recent memories with `recall`. If any two contradict each other, tell me which and ask which one is right. - If any session in the last 24 hours ended in an error, summarise what failed. - Anything I asked you to follow up on that has gone quiet for more than two days. ``` ## `eve dev` does not fire schedules on a timer Worth knowing before you conclude the heartbeat is broken, because everything *looks* right: the schedule compiles, the agent boots clean, and nothing ever happens. Under `eve dev`, schedules are registered but not driven by a clock. They fire when you ask: ```bash curl -X POST http://localhost:2000/eve/v1/dev/schedules/heartbeat # {"scheduleId":"heartbeat","sessionIds":["wrun_…"]} ``` The id is the filename under `agent/schedules/`. A 404 lists the ids that do exist, which is the fastest way to check eve found your file at all. The clock arrives with a built server — `eve build && eve start` runs Nitro's schedule runner, and that is what fires the cron in production. So the loop to expect is: develop against the dev route, deploy for the timer. This is eve's behaviour, not evestack's; `tracked()` records a dispatched fire exactly as it records a scheduled one, so the Schedules page fills in either way. ## Why it does not become spam This is the part that decides whether the feature survives contact with real use. The agent is told to reply with exactly `HEARTBEAT_OK` when nothing needs you. An hourly heartbeat that always sends something is an hourly notification, and you will mute it within a day — at which point you have a worse product than one that never spoke. **Two things this page used to promise that the code does not do.** Both were found by auditing the feature rather than reading it, and both are written up in the note at the top of `agent/schedules/heartbeat.ts`. **The token is not dropped.** Nothing filters it. The handler hands the turn to eve with `receive()`, and eve posts the reply itself, so evestack never sees the text and has nowhere to filter it from. A quiet hour therefore delivers the literal string `HEARTBEAT_OK` to your channel. There used to be an exported `isWorthDelivering(reply)` predicate here, called from nowhere. It has been deleted rather than shipped in a template people read and edit — a function named after what it would do, wired to nothing, reads as a working feature whatever its comment says. The one rule it encoded is recorded in `agent/schedules/heartbeat.ts` for whoever wires it, along with the two places it could actually go. **Wake-ups are not isolated sessions.** This page said each one "runs in an isolated session with a light context". The `isolatedSession` that was credited for it does not exist anywhere in evestack or in eve — so a wake-up costs whatever the session it lands in costs, and the few-thousand-tokens figure describes an intended design, not this code. ## What keeps it safe **It cannot quietly spend your money.** `@evestack/budget` caps at $2 per session and $10 per principal per day by default, and unattended work is exactly where an uncapped agent hurts — nobody is watching the turn that loops. **It cannot quietly do damage.** Anything behind `approval: always()` parks and waits. The template ships `forget` that way as the worked example. A heartbeat that hits an approval will sit there until a human answers, which is why `HEARTBEAT.md` tells the agent to describe what it would do rather than request approval it cannot get at 3am. **Every decision is attributable.** The [Approvals](/docs/dashboard) page records who approved what, when, and how their identity was established. Set `EVESTACK_REQUIRE_APPROVER=1` to refuse decisions that cannot be attributed to a person at all. **Every fire is on the record.** The Schedules page shows each wake-up, what it cost, what failed, and lets you pause one without a redeploy. Self-hosted eve runs schedules in-process and keeps no history, so without this a 3am job that has been failing for a week looks identical to one that has been fine. **Skills are scanned before they load.** eve advertises every skill in `agent/skills/` to the model and hands it a `load_skill` tool, so a skill can put instructions into a live turn before you see them. The [Skills](/docs/dashboard) page scans each one for injection, credential and exfiltration patterns — and says plainly that a clean verdict is not proof of safety. ## Catch-up, and when you do not want it If the machine was asleep at 09:00, should the 09:00 heartbeat run at 09:40? For a digest, yes — you still want it. For anything that acts on the world, no. So catch-up is opt-in per schedule, capped, and bounded by a window: the heartbeat replays at most 3 missed fires from the last 6 hours, so a laptop shut for a week does not wake up and replay a week. ```ts tracked("heartbeat", CRON, handler, { catchUp: true, catchUpLimit: 3, catchUpWindowMs: 6 * 60 * 60 * 1000, }); ``` Replayed fires are labelled as replays in the history, because a replay is not the same event as a live fire. ## Honest limits - **The heartbeat needs the agent running.** It is a cron inside your process, not a hosted scheduler. If the box is off, nothing fires — catch-up is what softens that, not a fix for it. - **Channels need a public HTTPS URL.** None of Slack, Discord or Telegram has a polling mode in eve, so a laptop needs a tunnel. See [Telegram](/docs/channels/telegram), which spells out the tunnel step; [Slack](/docs/channels/slack) and [Discord](/docs/channels/discord) are the same shape. - **A denied approval used to kill the session.** That was a real eve bug, [reported upstream](https://github.com/vercel/eve/issues/1658) and first worked around in the template's model middleware; `@ai-sdk/openai` v4 handles the denied output type natively, which is what the template pins. It matters more here than anywhere: an unattended agent hitting a denial at 3am should still be alive in the morning. --- # Telegram > The fastest non-HTTP way to talk to your self-hosted agent — a BotFather token and one tunnel. Slack wants an app manifest and a workspace admin. Discord wants an application, a public key and registered commands. Telegram wants a chat with [@BotFather](https://t.me/BotFather) and about sixty seconds. No review, no OAuth, no org to belong to. That makes Telegram the channel most self-hosters will actually finish, so it ships in the default template as `agent/channels/telegram.ts`. **There is no polling mode.** eve's Telegram adapter is webhook-only, so Telegram has to be able to reach your machine over public HTTPS. On a laptop that means a tunnel. See [Why you need a tunnel](#why-you-need-a-tunnel) — it is the only genuinely annoying step, and you cannot skip it. ## Why you need a tunnel The Telegram Bot API offers two ways to receive updates: `getUpdates` (long polling, your process calls out) and `setWebhook` (Telegram calls in). Polling needs no public URL at all, which would make this channel trivial to self-host. eve 0.30.6 implements only the second one. We checked the shipped adapter rather than guessing: - `getUpdates` does not appear anywhere in the `eve` package. - The adapter calls exactly these Bot API methods: `sendMessage`, `sendChatAction`, `answerCallbackQuery`, `editMessageReplyMarkup`, `getFile`, plus the raw file download. - `TelegramChannelConfig` exposes `api`, `botUsername`, `credentials`, `events`, `onCallbackQuery`, `onMessage`, `route` and `uploadPolicy`. There is no polling option. - The channel is a route — `POST /eve/v1/telegram` — and nothing in eve ever initiates a call to Telegram to fetch updates. eve also never calls `setWebhook` for you. Registering the URL is a one-time curl you run by hand, [below](#5-point-telegram-at-your-tunnel). If you want polling, it has to be built: a `defineChannel` sidecar that drives `getUpdates` itself and hands each update to `receive()`. That is real work, not a config flag. Until then, a tunnel is the answer, and the free ones take one command. ## 1. Get a token from BotFather 1. Open Telegram and message [@BotFather](https://t.me/BotFather). 2. Send `/newbot`. 3. Give it a display name (anything) and a username that must end in `bot` — e.g. `my_evestack_bot`. 4. BotFather replies with a token shaped like `123456789:AAH...`. That token **is** the bot. Anyone holding it can read and send everything the bot can, so treat it like a password. Optional, but do it now if you ever want the bot in a group: send `/setprivacy`, pick your bot, choose **Disable**. With privacy mode enabled (the default) Telegram only forwards commands and replies to the bot; disabling it lets `@mentions` through too. ## 2. Set the environment Add these to `templates/default/.env.local` (gitignored, never committed): ```bash TELEGRAM_BOT_TOKEN=123456789:AAH... # from BotFather TELEGRAM_WEBHOOK_SECRET_TOKEN=... # you invent this; see below TELEGRAM_BOT_USERNAME=my_evestack_bot # optional, no @ — needed for group @mentions ``` Generate the secret token yourself. Telegram accepts 1–256 characters from `A-Z a-z 0-9 _ -`: ```bash openssl rand -hex 32 ``` This secret is the **only** thing standing between your agent and anyone on the internet who finds the webhook URL. Telegram echoes it back in the `X-Telegram-Bot-Api-Secret-Token` header on every delivery, and eve compares it in constant time before parsing a single byte of body. The Telegram route is **not** behind evestack's HTTP Basic auth. `agent/channels/eve.ts` guards the eve HTTP channel; `POST /eve/v1/telegram` is guarded by the secret token alone. That is correct — Telegram's servers cannot send Basic credentials — but it means a weak or missing secret is a wide-open door. eve fails closed if `TELEGRAM_WEBHOOK_SECRET_TOKEN` is unset: every inbound request gets `401 unauthorized`. With no `TELEGRAM_BOT_TOKEN` at all, the agent still boots. The channel logs one line and sits idle, exactly like `agent/tools/composio.ts` without its key: ``` [evestack:telegram] TELEGRAM_BOT_TOKEN is not set, so the Telegram channel is idle. ``` ## 3. Add the channel It is already in the default template. To add it to an existing eve project: ```bash eve registry add @evestack=https://raw.githubusercontent.com/SammyTourani/evestack/main/registry/r/{name}.json eve add @evestack/channel-telegram ``` Or write the file yourself — the filename is the channel id: ```ts title="agent/channels/telegram.ts" import { telegramChannel } from "eve/channels/telegram"; export default telegramChannel({ botUsername: process.env.TELEGRAM_BOT_USERNAME, uploadPolicy: { allowedMediaTypes: ["image/*", "application/pdf"], maxBytes: 20 * 1024 * 1024, }, }); ``` Confirm eve sees it: ```bash eve channels list # eve # telegram ``` ## 4. Expose your laptop Pick one. Both give you HTTPS on port 443, which Telegram requires (it only accepts webhook URLs on 443, 80, 88 or 8443). **cloudflared** — no account, no signup, one command: ```bash brew install cloudflared cloudflared tunnel --url http://localhost:2000 # ... # https://random-words-here.trycloudflare.com ``` **ngrok** — free account required: ```bash ngrok http 2000 # Forwarding https://random.ngrok-free.app -> http://localhost:2000 ``` Both free tiers hand you a **new random hostname every restart**. When the tunnel restarts you must re-run `setWebhook` with the new URL, or Telegram keeps posting into the void. If the bot goes quiet after a reboot, this is why. A named cloudflared tunnel on a domain you own fixes it permanently and still costs nothing. ## 5. Point Telegram at your tunnel ```bash curl -X POST "https://api.telegram.org/bot$TELEGRAM_BOT_TOKEN/setWebhook" \ -H "Content-Type: application/json" \ -d '{"url":"https://random-words-here.trycloudflare.com/eve/v1/telegram", "secret_token":"'"$TELEGRAM_WEBHOOK_SECRET_TOKEN"'", "allowed_updates":["message","callback_query"]}' ``` `allowed_updates` matters: `message` is the conversation and `callback_query` is how human-in-the-loop buttons report back. Leave `callback_query` out and approval buttons will appear but never resolve. Check that it took, and keep checking it — this endpoint is the single best debugging tool for this channel: ```bash curl -s "https://api.telegram.org/bot$TELEGRAM_BOT_TOKEN/getWebhookInfo" ``` A healthy response has your URL, `"pending_update_count": 0` and no `last_error_message`. If `last_error_message` says `Wrong response from the webhook: 401 Unauthorized`, your secret token does not match what the agent has. If it says connection refused or timed out, the tunnel is down or `eve dev` is not running. ## 6. Talk to it Make sure the agent is up (`npm run dev` in your project, serving `:2000`), then DM your bot in Telegram. It should start typing and answer. ## What actually reaches the agent Group chats are deliberately stricter than DMs. eve's default dispatch rule, read straight out of the adapter: | Where | Wakes the bot | | --- | --- | | Private chat | Any message with text, a caption, or a supported attachment | | Group / supergroup | A `/command`, an `@yourbot` mention, or a reply to a message from *any* bot | | Broadcast channel | Nothing — always ignored | | Any chat, sender is a bot | Nothing — always ignored | Two sharp edges worth knowing: - **`@mentions` need `TELEGRAM_BOT_USERNAME`.** Without it eve has no idea what its own handle is, so in a group only `/commands` and replies work. - **Any slash command wakes it.** `/anything` with no `@target` counts, so an unrelated bot's `/roll` in a shared group will start a turn. Scope it with `/roll@otherbot` or keep the bot out of busy groups. Override `onMessage` if you need stricter rules. - **The reply rule checks "is a bot", not "is *this* bot".** eve gates on `replyToMessage.from.isBot`, so replying to *any* bot in the group also wakes yours. Same mitigation as above. Forum topics carry `message_thread_id` through the continuation token, so each topic keeps its own session. ## Identity, and why it matters for budgets eve derives a principal from every inbound message: - Private chat: `telegram:` - Group or supergroup: `telegram::` — the same person in two groups is two principals That is exactly the key [`@evestack/budget`](/docs/registry) caps spend against, so a per-principal daily limit is a per-Telegram-user daily limit with no extra wiring. ## Attachments The template allows images and PDFs up to 20 MB. eve does not download anything until the policy says yes, then fetches it on demand with `getFile`. The 20 MB is not arbitrary: the Bot API refuses to serve files larger than that through `getFile`, so eve's own 25 MB default only converts a clean policy rejection into a failed download. Widen the types if you need to, but do it knowingly — `allowedMediaTypes: "*"` hands your model whatever a stranger in a group chat decides to attach. ## Formatting The default `message.completed` handler sends plain text with no `parse_mode`, so Markdown from the model shows up literally — `**bold**` renders as `**bold**`. Replies over Telegram's 4096-character limit are split across several messages. If you want real formatting, override the handler and set `parse_mode`. Be careful: Telegram's MarkdownV2 requires escaping a long list of characters, and an unescaped one makes the whole `sendMessage` call fail, which looks like the bot ignoring you. ## Teardown and rotation ```bash # stop delivery without touching the bot curl -X POST "https://api.telegram.org/bot$TELEGRAM_BOT_TOKEN/deleteWebhook" ``` If a token ever leaks, send `/revoke` to BotFather, take the new token, update `.env.local`, and re-run `setWebhook`. The old token dies immediately. ## Troubleshooting | Symptom | Cause | | --- | --- | | `401 unauthorized` from `/eve/v1/telegram` | `TELEGRAM_WEBHOOK_SECRET_TOKEN` unset, or it does not match the value you passed to `setWebhook` | | `getWebhookInfo` shows a rising `pending_update_count` | Telegram cannot reach you: tunnel down, agent down, or a stale hostname after a tunnel restart | | Bot answers in DMs but is silent in a group | Privacy mode is on (`/setprivacy` → Disable) and/or `TELEGRAM_BOT_USERNAME` is unset | | Buttons appear but nothing happens when tapped | `callback_query` missing from `allowed_updates` — re-run `setWebhook` | | Agent boots but the channel does nothing | `TELEGRAM_BOT_TOKEN` unset; look for the `[evestack:telegram]` line at startup | | Replies contain literal `**asterisks**` | Expected — the default handler sends plain text | --- # Slack > Put your self-hosted agent in a Slack workspace with a bot token and a signing secret. No Vercel account, no Connect connector. ## The Vercel Connect question, answered eve's own Slack documentation is written entirely around Vercel Connect and states that credentials "run through Vercel Connect... so there's no `SLACK_BOT_TOKEN` or `SLACK_SIGNING_SECRET` for you to manage." Read quickly, that sounds like a requirement. It is not. In eve 0.30.6 the whole `credentials` object is optional and every field falls back to the environment: | What Connect supplies | What eve falls back to | Where | | --- | --- | --- | | `credentials.botToken` | `process.env.SLACK_BOT_TOKEN` | `resolveSlackBotToken` | | `credentials.webhookVerifier` | `process.env.SLACK_SIGNING_SECRET`, HMAC-verified in-process | `verifyInbound` | Connect is one implementation of the `webhookVerifier` hook — the escape hatch, not the contract. `agent/channels/slack.ts` passes no `credentials` at all and works on two ordinary environment variables. So the comparison table stands: Vercel Connect ships four managed connectors and wants a Vercel account for the managed OAuth path. A self-hosted agent needs a bot token. We proved the handshake without ever registering a Slack app, by signing a real `url_verification` payload with a throwaway secret and calling eve's own route handler: ``` 200 "3eZbrw1aB2rC6hjIsCLNL1lLbyTYs9SC3mHtNjnpNyGqDMWDb5" ← signed url_verification 401 "unauthorized" ← tampered signature 401 "unauthorized" ← replay outside 300s skew 401 "unauthorized" ← no signature headers ``` `VERCEL_USE_EXPERIMENTAL_FRAMEWORKS`, `vercel connect create`, and `@vercel/connect` appear nowhere in this path. ## Before you start Slack only calls **public HTTPS** URLs. There is no polling mode for the Events API. A laptop therefore needs a tunnel: ```bash cloudflared tunnel --url http://localhost:2000 # or: ngrok http 2000 ``` Everything below uses `https://YOUR-HOST` for whatever that prints. The one route you need is: ``` POST https://YOUR-HOST/eve/v1/slack ``` Both Events **and** interactive button clicks go to that single path — eve branches on the content type internally. This route is **not** behind the HTTP Basic policy in `agent/channels/eve.ts`. Route auth there guards the three eve session routes; a channel's routes carry their own verification, which for Slack is the request signature. Do not put Basic auth in front of `/eve/v1/slack` in a reverse proxy — Slack cannot answer a challenge, and you would be trading a signature check for a 401. ## 1. Create the app Go to **[api.slack.com/apps](https://api.slack.com/apps) → Create New App → From an app manifest**, pick your workspace, and paste this: ```yaml display_information: name: evestack features: bot_user: display_name: evestack always_online: false app_home: home_tab_enabled: false # Required for DMs. Without it, users cannot type to the bot at all. messages_tab_enabled: true messages_tab_read_only_enabled: false oauth_config: scopes: bot: - app_mentions:read - chat:write - im:history - im:write - channels:history - groups:history - files:read settings: org_deploy_enabled: false socket_mode_enabled: false token_rotation_enabled: false ``` The manifest deliberately omits the request URLs. Slack verifies a request URL the moment it appears in a manifest, and you do not have the signing secret to verify it with until the app exists. Scopes first, URLs in step 4. Prefer clicking? **From scratch** works too — you then add each scope by hand under **OAuth & Permissions → Scopes → Bot Token Scopes**, and flip **App Home → Show Tabs → Messages Tab** on along with *Allow users to send Slash commands and messages from the messages tab*. ### What each scope buys | Scope | Why | | --- | --- | | `app_mentions:read` | Receive `app_mention`. Without it the bot never hears you. | | `chat:write` | Post replies. Everything the agent says goes through this. | | `im:history` | Read DM content. Required for `message.im` to carry text. | | `im:write` | Open a DM conversation — used by `postDirectMessage`, and by the default HITL flow to deliver a sign-in challenge privately instead of in-channel. | | `channels:history` | Public-channel thread continuation and `threadContext`. Optional. | | `groups:history` | The same in private channels. Optional. | | `files:read` | Download inbound attachments from authenticated Slack URLs so the model can see them. Optional. | | `files:write` | Only if the agent should upload files back. | Not needed: `users:read` (eve attributes speakers by stable user id and never does profile lookups), and `commands` (slash commands do not reach eve's handlers). `assistant:write` is worth knowing about. eve shows progress with `assistant.threads.setStatus` — `Thinking…`, then `Working…`, then a live action label. That method belongs to Slack's assistant surface: it needs `assistant:write` and applies to assistant threads. In an ordinary channel thread the call fails, eve logs and swallows it, and you lose the indicator and nothing else. Replies are unaffected either way. ## 2. Install to the workspace, get the token **OAuth & Permissions → Install to Workspace → Allow.** Copy the **Bot User OAuth Token** (`xoxb-...`). Re-installing is required every time you add a scope. Slack will show a yellow banner; take it seriously, because the old grant keeps working and the new scope silently does not. ## 3. Get the signing secret **Basic Information → App Credentials → Signing Secret → Show.** This is not the App-Level Token (`xapp-`) and not the deprecated Verification Token. Copy the wrong one and every request gets a 401. Put both into `templates/default/.env.local`: ```bash SLACK_BOT_TOKEN=xoxb-... SLACK_SIGNING_SECRET=... ``` Restart the agent so the process actually has them: ```bash npm run dev ``` An unconfigured channel is not a broken one — the route still registers and the agent logs a single line at boot: ``` [evestack:slack] SLACK_BOT_TOKEN and SLACK_SIGNING_SECRET not set, so the Slack channel is registered but idle: unsigned inbound requests get a 401 and outbound Web API calls throw. ``` ## 4. Event Subscriptions and the challenge **Event Subscriptions → Enable Events → On.** Set **Request URL** to: ``` https://YOUR-HOST/eve/v1/slack ``` Slack immediately POSTs a `url_verification` body containing a random `challenge` string, and expects it echoed back. eve does this for you and answers `200 text/plain` with the challenge verbatim. The field turns green and says **Verified**. **Order matters.** Slack signs the challenge request like every other request, so `SLACK_SIGNING_SECRET` must already be in the running process. If it is not, eve answers 401, Slack reports *"Your request URL didn't respond with the value of the challenge parameter"*, and your agent log carries the real reason: ``` [eve:slack.channel] slack inbound verification failed Error: slackChannel: missing signing secret. Pass credentials.signingSecret, set SLACK_SIGNING_SECRET, or supply credentials.webhookVerifier. ``` Then, under **Subscribe to bot events**, add: | Event | Effect | | --- | --- | | `app_mention` | `@evestack ...` in any channel the bot is in. The baseline. | | `message.im` | Direct messages. | | `message.channels` | Lets a thread continue without re-mentioning the bot. Optional. | | `message.groups` | The same in private channels. Optional. | **Save Changes**, and reinstall if prompted. ## 5. Interactivity (needed for approvals) **Interactivity & Shortcuts → Interactivity → On**, and set **Request URL** to the *same* `https://YOUR-HOST/eve/v1/slack`. Skip this and human-in-the-loop breaks in a way that looks like nothing at all: the agent posts its Approve/Deny buttons, you click one, and the turn stays parked forever. Slack delivers `block_actions` as `application/x-www-form-urlencoded` to the interactivity URL, and eve routes that content type to its interaction handler on the same path. ## 6. Invite the bot In each channel it should work in: ``` /invite @evestack ``` `app_mentions:read` grants nothing in a channel the bot has not joined. ## What you get | You do | The agent does | | --- | --- | | `@evestack summarize this thread` | Replies in-thread with a live status indicator. | | DM the bot | Same, in your IM conversation. | | Reply in a thread it is already in | Continues without a mention — needs `message.channels` + `channels:history`. | | Click an approval button | Resumes the parked turn. | | Upload a file with your mention | Staged and passed to the model — needs `files:read`. | Signature verification is HMAC-SHA256 over `v0:{timestamp}:{raw body}`, compared in constant time, with a 300-second skew window. Retries whose `x-slack-retry-reason` is `http_timeout` are dropped, so a slow first turn does not get delivered twice. ## Installing this into an existing eve project ```bash eve registry add @evestack=https://raw.githubusercontent.com/SammyTourani/evestack/main/registry/r/{name}.json eve add @evestack/channel-slack ``` Use that rather than `eve add channel/slack`, which is eve's own scaffold and walks you into the Vercel Connect flow. ## When it does not work | Symptom | Cause | | --- | --- | | "Request URL didn't respond with the value of the challenge parameter" | `SLACK_SIGNING_SECRET` missing from the *running* process, tunnel down, or the agent was not restarted after editing `.env.local`. | | Every event 401s, no log about the secret | Wrong secret. The App-Level Token (`xapp-`) and the legacy Verification Token both look plausible and are both wrong. | | `inbound request timestamp outside allowed skew` | Host clock drift over five minutes. Fix NTP. | | `inbound request signature mismatch` in production only | Something between Slack and eve is rewriting the body. The HMAC covers raw bytes — a proxy that re-serializes JSON, strips whitespace, or decompresses will break it. | | Bot ignores mentions in one channel | Not invited there (`/invite @evestack`). | | Scope added, still unauthorized | Slack requires a reinstall to widen a grant. | | DMs do nothing | App Home messages tab off, or `message.im` / `im:history` missing. | | Approval buttons do nothing | Interactivity request URL not set (step 5). | | `Error: SLACK_BOT_TOKEN is required.` | Inbound verification passed, so events arrive — but the reply cannot be posted. Outbound needs the token; inbound needs the secret. They fail independently. | ## Customizing `templates/default/agent/channels/slack.ts` ships two hooks: `onAppMention`, which answers `@evestack ...`, and `onMessage`, which handles DMs and continues threads the agent already has a session in. Three details are worth reading before you edit it: - **`onAppMention` is not optional here, even though it only reproduces eve's default.** eve resolves an inbound event with `(kind === "app_mention" ? onAppMention : onDirectMessage) ?? onMessage`, so defining `onMessage` alone captures app mentions as well. They then hit the `isSubscribed()` gate — false for a first mention, because no session exists yet — and the turn is dropped. The symptom is a bot that ignores you until it has already answered you once, which it never does. If you delete `onAppMention`, you reintroduce that. - Defining `onMessage` takes DMs away from eve's built-in DM default, because eve resolves `onDirectMessage ?? onMessage`. That is why the file has an explicit DM branch reproducing the default's auth derivation and typing indicator. - eve drops channel messages containing `<@botUserId>` before `onMessage` runs, because the same message already arrived as a separate `app_mention` delivery with its own event id — and the duplicate cache is keyed on event id, so it would not catch it. A consequence: inside `onMessage`, `ctx.isBotMentioned()` is always false for channel messages. eve's own docs show a snippet that gates on it; on the `onMessage` path that branch is unreachable. `ctx.reset()` for a `/new` command, `ctx.cancel()` for debouncing, `threadContext` for injecting earlier thread replies, and `onEvent` for things like `team_join` are all documented in eve's bundled `node_modules/eve/docs/channels/slack.mdx` — and every example there that reads `credentials: connectSlackCredentials(...)` can simply have that line deleted. --- # Discord > Slash commands into a self-hosted eve agent — signature-verified, no Vercel Connect client. Vercel's guided setup (`eve add channel/discord`) creates a **Vercel Connect** client, stores your bot token in Vercel, and configures the endpoint for you. None of that works without a Vercel account. This page is the same result done by hand, on your own machine, for $0. The whole channel is one file: ```ts title="agent/channels/discord.ts" import { defaultDiscordAuth, discordChannel } from "eve/channels/discord"; export default discordChannel({ /* … */ }); ``` It serves `POST /eve/v1/discord`, which handles slash commands, human-in-the-loop button clicks, and modal submissions on one route. ## Who authenticates this route Read this before you expose anything. Discord signs every interaction with **Ed25519** over `X-Signature-Timestamp` + the exact raw request body, and sends the signature in `X-Signature-Ed25519`. An interactions endpoint that does not check that signature is not "unauthenticated" in the mild sense — it is a public endpoint where anyone can forge a message from any Discord user in any server. **eve verifies it for you.** `verifyDiscordRequest` runs before the body is even parsed: it rebuilds the signed payload, checks it against `DISCORD_PUBLIC_KEY`, rejects a timestamp more than five minutes from now, and returns a bare `401` on any failure. You write no crypto, and you should not try to — there is no hook that asks you to. What *is* your responsibility is the one env var: With `DISCORD_PUBLIC_KEY` unset, eve has no key to compare against, so **every** interaction fails verification — including Discord's own endpoint-verification PING. The channel is registered, the route answers 401, and nothing reaches the model. That is the intended fail-closed state, not a bug, and the agent logs one line at boot saying so. Note also what does *not* guard this route: the HTTP Basic policy in `agent/channels/eve.ts` covers `/eve/v1/session*` only. Discord cannot send an `Authorization` header, so do not put your tunnel or reverse proxy behind Basic auth — Discord will simply fail to reach you. The signature is the auth here, and it is a stronger one. ## The ordering trap Discord will not save an Interactions Endpoint URL until that URL answers its PING challenge correctly. During validation Discord sends both a properly signed request (which must return `{"type":1}`) and a deliberately **mis**-signed one (which must be rejected). So the sequence is not "configure Discord, then run the agent" — it is the reverse: 1. Create the application and copy its **Public Key**. 2. Put `DISCORD_PUBLIC_KEY` in `.env.local` and restart the agent. 3. Get the agent on a **public HTTPS** URL. 4. *Then* paste the Interactions Endpoint URL into the Portal and save. Do it in the other order and the Portal rejects the URL with an unhelpful error, and people spend an hour assuming their key is wrong. ## Click-path through the Developer Portal ### 1. Create the application 1. Go to [discord.com/developers/applications](https://discord.com/developers/applications) and sign in. 2. **New Application** → give it a name → accept the developer ToS → **Create**. 3. You land on **General Information**. Two values live here: - **Application ID** → `DISCORD_APPLICATION_ID` - **Public Key** → `DISCORD_PUBLIC_KEY` Leave the **Interactions Endpoint URL** field alone for now. That is step 5. ### 2. Get the bot token 1. Left sidebar → **Bot**. 2. **Reset Token** → confirm → copy it. Discord shows a bot token exactly once; if you lose it you reset it again and the old one dies. 3. Turn **Public Bot** off unless you want strangers installing your agent. 4. **Privileged Gateway Intents — leave all three off.** Presence, Server Members, and Message Content are gateway features. This channel is HTTP Interactions only: eve never opens a gateway socket, so enabling them adds risk and buys nothing. If a guide tells you to switch on Message Content, that guide is about a gateway bot, not this one. `DISCORD_BOT_TOKEN` is **not** required for slash commands to work. Replies ride the interaction token, and eve reads the application id off the inbound payload. The token buys three things: the typing indicator (failures are swallowed silently), proactive sessions started from a schedule, and the fallback to a normal channel message after the interaction token expires 15 minutes into a long session. Skip it and commands still answer. ### 3. Install the bot into a server Left sidebar → **Installation** (or **OAuth2 → URL Generator** on older Portal builds). - **Scopes**: `applications.commands` is the only one slash commands need. Add `bot` as well if you set a bot token and want the channel-message fallback or proactive posts. - **Bot permissions**: with the `bot` scope, **Send Messages** (and **Read Message History** if you want it to see thread context). Nothing else. Copy the generated install URL, open it, pick a server, authorize. ### 4. Register the command The Portal has no UI for creating slash commands — it is an API call. eve extracts the prompt from a string option literally named `message`, so use that name: ```bash # Guild-scoped: appears instantly. Best for testing. curl -X PUT "https://discord.com/api/v10/applications/$DISCORD_APPLICATION_ID/guilds/$GUILD_ID/commands" \ -H "Authorization: Bot $DISCORD_BOT_TOKEN" -H "Content-Type: application/json" \ -d '[{"name":"ask","description":"Ask the eve agent","type":1, "options":[{"name":"message","description":"What should the agent do?","type":3,"required":true}]}]' ``` Drop `/guilds/$GUILD_ID` for a global command, which works everywhere but can take up to an hour to propagate. This is the one step that genuinely needs the bot token. ### 5. Go public, then save the endpoint URL Discord only calls public HTTPS, so `localhost:2000` is not reachable. Tunnel it: ```bash cloudflared tunnel --url http://localhost:2000 ``` Put the credentials in `templates/default/.env.local` (gitignored — never commit them): ```bash DISCORD_PUBLIC_KEY=... # required; without it every interaction 401s DISCORD_APPLICATION_ID=... # needed for step 4; optional at runtime DISCORD_BOT_TOKEN=... # optional; typing, proactive posts, expired-token fallback DISCORD_ALLOWED_GUILD_IDS=... # optional; see below ``` Restart the agent so it picks them up. **Now** go back to **General Information**, set **Interactions Endpoint URL** to: ``` https:///eve/v1/discord ``` and **Save Changes**. Discord runs its PING check against a live agent and the save succeeds. ### 6. Use it Type `/ask` in a channel where the bot is installed. Discord enforces a three-second acknowledgement deadline; eve defers immediately and edits the deferred reply when the turn finishes, so a two-minute agent run is fine. ## Locking down who can spend your budget An interactions endpoint is world-reachable by design, and every accepted command runs on *your* API key. eve's default is to dispatch for any user in any server the bot is in. ```bash DISCORD_ALLOWED_GUILD_IDS=123456789012345678,987654321098765432 ``` Commands from anywhere else get an ephemeral "Command ignored." and never start a turn. Unset means "any guild", which is exactly eve's stock behavior. Direct messages to the bot have no guild id, so they are rejected once the list is non-empty. This is a complement to, not a replacement for, `@evestack/budget` — the allow-list caps *who* can start a turn, the budget caps *how much* a turn costs. Discord principals arrive as `discord::`, which is what the per-principal daily cap meters on. ## What this channel gives the agent - **Slash commands** with the `message` option as the prompt. - **Human-in-the-loop**: eve renders approval requests as Discord components. Confirmations and option lists become buttons, `display: "select"` becomes a string select, and freeform input becomes a button that opens a modal. Answering resumes the parked session — which means `forget`'s approval gate works from Discord with no extra wiring. - **Long replies** split at Discord's 2000-character limit, with `allowed_mentions` cleared so a generated message can never mass-ping a server. Inbound file attachments are not supported by eve's Discord channel today. ## Verified without a bot Everything below was checked with no Discord account, by generating a throwaway Ed25519 keypair, pointing `DISCORD_PUBLIC_KEY` at it, and driving the route handler directly: | Request | Result | | -------------------------------------- | ----------------------------------------- | | PING with no signature headers | `401 unauthorized` | | PING with a body altered after signing | `401 unauthorized` | | Correctly signed PING | `200 {"type":1}` — the PONG Discord wants | | Correctly signed `/ask` | `200 {"type":5}`, one turn dispatched | The dispatched turn carried the option text as the message, `principalId` `discord::`, and authenticator `discord-interaction`. What still needs a real bot: that Discord's own endpoint validation accepts the URL end to end, that command registration returns 201, that deferred-reply edits and followups land in the channel, and that HITL buttons round-trip. --- # Self-hosting > The production runbook behind eve's self-hosting spec — schema, proxy routes, auth off Vercel, and what each one does when it's wrong. [eve's deployment guide](https://eve.dev/docs/guides/deployment/self-hosting) is short and correct. `eve build` writes a Nitro server to `.output/`, `eve start` serves it, a workflow world must be built against the same `@workflow/*` line as your eve release, `vercelOidc()` is not a production authenticator off Vercel, and your proxy must forward both `/eve/` and `/.well-known/workflow/`. Every one of those statements is true. It is a specification, not a runbook. It never names a database: no schema, no migration command, no connection-string variable, no reference topology, and no note about which of these failures are loud and which are silent. This page is that half — written from the stack in this repository, running. [Quickstart](/docs/quickstart) is the happy path on your laptop. This is what changes when the process is built rather than `eve dev`, and something other than you can reach the port. ## Reference topology | Component | Port | Bound to | Notes | | --- | --- | --- | --- | | Agent, `eve dev` | 2000 | **`127.0.0.1`** | A bare `eve dev` auto-increments if 2000 is taken. A scaffolded project does not: `npm run dev` passes `--port`, so a busy port is `EADDRINUSE` | | Agent, `eve start` | `$PORT`, else **3000** | `--host`, default all interfaces | Not 2000 — see below | | Postgres | host **5433** → container 5432 | `127.0.0.1` in compose | 5433 so it never collides with a local Postgres on 5432. Override with `POSTGRES_BIND` | | Dashboard | 4000 | `127.0.0.1` in compose | A control plane; see [Expose the dashboard last](#expose-the-dashboard-last) | `eve dev` listens on 2000; `eve start` defaults to `$PORT` and then 3000. The dashboard's default `EVESTACK_AGENT_URL` is `http://127.0.0.1:2000`, so run the built server as `PORT=2000 eve start` and the rest of the stack needs no reconfiguration. Otherwise set `EVESTACK_AGENT_URL` to wherever it actually landed. **Both cells in that first row were wrong until 2026-08-09, and the binding one matters on Linux.** They said "all interfaces" and "auto-increments if 2000 is taken". Neither is true of a scaffolded project: - `eve dev` falls back to `DEFAULT_DEVELOPMENT_SERVER_HOST`, which is `127.0.0.1` (`dist/src/internal/nitro/host/dev-server-url.js`). Loopback, not all interfaces. - `retryOnAddressInUse` is set only when **no** port is passed, and `scripts/dev.mjs` always passes `--port` from `EVESTACK_AGENT_PORT`. So a collision throws rather than moving. (`eve dev` with no port *does* auto-increment — which is why the template's own comment describing that behaviour is correct and this table was not.) **The Linux consequence.** The generated compose reaches the agent at `http://host.docker.internal:2000` with a `host-gateway` mapping. On macOS Docker Desktop proxies that to loopback and it works. On Linux `host-gateway` is the bridge IP, which a loopback-bound `eve dev` is not listening on — so Postgres reads keep working while everything the dashboard *drives* fails, which reads like a dashboard bug rather than a networking one. Run the agent as `eve dev --host 0.0.0.0` (or use the built server, which honours `--host`) if you need the container to reach it. Not reproduced here — this laptop is macOS — so treat it as derived from the source above rather than measured. Both mappings in the compose file are on loopback: the dashboard's is written `127.0.0.1:${DASHBOARD_PORT:-4000}:4000` and Postgres's is `${POSTGRES_BIND:-127.0.0.1}:${POSTGRES_PORT:-5433}:5432`. Set `POSTGRES_BIND` if you deliberately need the database reachable from another machine — and set `POSTGRES_PASSWORD` in the same breath, because the repo's compose still defaults it to `evestack`, which is only defensible while the port is on loopback. This paragraph used to describe an asymmetry — the dashboard on `127.0.0.1` and Postgres "published on every interface" — and told you to prefix the mapping yourself. That was true when it was written and is not now: both compose files bind loopback, and the one `create-evestack` generates also writes a per-project password rather than a default. Left visible rather than quietly deleted, because a runbook that overstates an exposure spends the reader's trust on the paragraphs that are still true. There is no OTLP collector in this topology and nothing listens on 4318. Trace export goes to the dashboard's own ingest route (`http://localhost:4000/api/ingest/v1/traces`), because `@vercel/otel` uses the configured URL verbatim — the conventional collector address will not reach it. See [Observability](/docs/observability). ## Deploy `eve`, the `@workflow/*` line, and the world package are one compatibility unit, and only one of the three is checked for you. The shipped `package.json` pins `eve` at `^0.30.8` and `@workflow/world-postgres` at exactly `5.0.0-beta.32` — an exact version, not a range and not a dist-tag. `>=0.30.0` is the floor for any deployment reachable from a network — see [Auth off Vercel](#auth-off-vercel). ```bash docker compose up -d postgres ``` The compose file uses `pgvector/pgvector:pg17` rather than plain `postgres`, because the same database also backs [agent memory](/docs/memory) — one container, two jobs. Data lives in the named volume `evestack-pgdata`. ```bash npm run db:bootstrap ``` Nothing creates it for you. `@workflow/world-postgres` runs its migrations only from its own CLI and eve never invokes it, so a server started against a fresh database starts against a database with no tables. Run this before the first boot and again after upgrading the world package. `placeholderAuth()` and `vercelOidc()` — what stock `eve init` scaffolds — are both wrong off Vercel. `agent/channels/eve.ts` in this template ships `httpBasic` reading `EVESTACK_AUTH_USER` / `EVESTACK_AUTH_PASSWORD`, which `create-evestack` generates per project. JWT (HMAC or ECDSA), generic OIDC, and custom verifiers are equally valid; eve owns that surface and documents it in [auth and route protection](https://eve.dev/docs/guides/auth-and-route-protection). ```bash npm run build # eve build -> .output/ npm start # eve start, on EVESTACK_AGENT_PORT ``` Every `eve` command loads `.env`/`.env.local` from the app root first, `eve start` included, so the same file that drives `eve dev` drives the built server. No `PORT=` prefix is needed and this page used to print one: `scripts/start.mjs` passes `EVESTACK_AGENT_PORT` through as `--port`, which beats `$PORT` and eve's own default of 3000. That is the number `npm run verify` probes and the number the generated compose file points the dashboard at, so there is one answer to "where is the agent" rather than three. eve does not supervise itself. A unit and a plist ship in the project at `deploy/` — see [Operations](/docs/operations#run-the-agent-as-a-service), which also covers why the agent is not a compose service beside Postgres and the dashboard. Forward `/eve/` **and** `/.well-known/workflow/`, unrewritten. Config for nginx and Caddy is [below](#reverse-proxy). This is the step that silently half-works if you get it wrong. ```bash curl https://agent.example.com/eve/v1/health curl -u "$EVESTACK_AUTH_USER:$EVESTACK_AUTH_PASSWORD" https://agent.example.com/eve/v1/info ``` `/eve/v1/health` is a Nitro-level route outside the channel's auth policy, so a 200 there proves the proxy path and nothing about your credentials. `/eve/v1/info` is registered and authenticated with the same `auth` input as the session routes, so it is the one that proves both. Then drive a real turn: `eve dev https://agent.example.com` attaches the TUI to a deployed server. ## The Postgres world Pin `@workflow/world-postgres` to an **exact version** — `5.0.0-beta.32`, which is what `templates/default` declares. Never `latest`, and no longer the `beta` dist-tag either. `latest` is a whole major behind the protocol eve speaks and always has been. `beta` was the documented workaround for that, and it stopped being safe: upstream ships World **spec** changes inside the `5.0.0-beta.*` line with no semver signal, so the same tag hands out incompatible worlds on different days. Measured against eve 0.30.8: | `world-postgres` | pulls `@workflow/world` | World spec | eve 0.30.8 | | --- | --- | --- | --- | | `5.0.0-beta.32` | `5.0.0-beta.25` | 5 | boots | | `5.0.0-beta.34` | `5.0.0-beta.27` | 6 | `worker init failed` at startup | On 2026-08-19 the `beta` tag moved from `.34` to `.35` inside a single working session, so "check the tag yourself" is not a defence — the answer expires. `^5.0.0-beta.32` and `~5.0.0-beta.32` are not fixes either: both still resolve to `.34` and `.35`. Move the pin only after installing a candidate and watching it boot. Selecting the world is one field in `agent/agent.ts`, and it is conditional on the connection string being present: ```ts const workflow = process.env.WORKFLOW_POSTGRES_URL ? { world: "@workflow/world-postgres" } : undefined; ``` That conditional is the silent failure to know about. With `WORKFLOW_POSTGRES_URL` unset or unreachable, eve falls back to its on-disk world under `.eve/.workflow-data` and keeps working — the agent answers, sessions resume, and nothing in the logs reads like an error. You only find out when the container is replaced and the history is gone. If you deliberately run the on-disk world, mount that directory on persistent storage. `npm run db:bootstrap` creates three schemas — `workflow`, `workflow_drizzle`, and `graphile_worker` — and, inside `workflow`, the tables `workflow_runs`, `workflow_events`, `workflow_steps`, `workflow_hooks`, `workflow_waits`, and `workflow_stream_chunks`. `workflow_runs` is the one you will query: its `attributes` JSONB column holds eve's `$eve.*` run tags, which is why the dashboard needs no ingest pipeline. See [Architecture](/docs/architecture). Run bootstrap through the `db:bootstrap` script, not as `npx --package=@workflow/world-postgres bootstrap`. That CLI loads `.env` through dotenv and never reads `.env.local` — the only env file `create-evestack` writes — so it falls back to `postgres://world:world@localhost:5432/world` and dies on `ECONNREFUSED`. The script passes `--env-file-if-exists=.env.local` explicitly. Two more variables matter under load. `world-postgres` defaults to a worker concurrency of 50 against a pool of 10 and warns about it on every boot; `WORKFLOW_POSTGRES_MAX_POOL_SIZE` and `WORKFLOW_POSTGRES_WORKER_CONCURRENCY` (both 20 in the shipped `.env.example`) silence the warning and stop workers queueing on connections. ## Auth off Vercel eve fails closed: when no authenticator in the chain grants, `routeAuth()` answers 401. That is the correct posture and it produces one result that reliably reads as a bug. **On a built server, `127.0.0.1` gets a 401 too — and that is correct.** From eve 0.30, `localDev()` grants on the process being an `eve dev` / `vercel dev` run (`EVE_DEV=1`) and consults nothing in the request. `eve build && eve start` is not that process, so nothing is granted implicitly and every request needs the Basic credentials, loopback included. Measured against this stack: all hosts 401, correct credentials 200. The inverse used to be true, and it was exploitable. On eve 0.29.x, `localDev()` decided "is this my machine" from the request URL's hostname — derived from the client's `Host` header — matched against an unanchored `/^127\./` plus `endsWith(".localhost")`. So `127.evil.com`, a name anyone can register and point at your agent, received a full local-dev principal with no credentials: session creation and tool execution, unauthenticated, from the internet. We measured it (that host answered 200 where a plain foreign host answered 401) and shipped a `strictLocalDev()` wrapper. Vercel fixed it properly upstream in 0.30.0 and `isLoopbackRequest` is gone, so the wrapper was deleted rather than kept — on 0.30 it could add no protection and would have rejected legitimate local-dev access over a LAN IP, a tunnel, or a container hostname. **Pin `eve` `>=0.30.0`.** That is the release the fix landed in, and it is the same floor SECURITY.md and every published peer range state; the template pins `^0.30.8` because that is the release evestack tests against, which is a different question. `contract/contracts/07-auth.contract.mjs` asserts both halves of the fixed behaviour — that no hostile `Host` makes `localDev()` grant, and that it still grants inside `EVE_DEV=1` — and it goes red against 0.29.5. See The contract suite in `contract/` records which eve versions have been run against it. ### The route-auth policy does not cover everything under `/eve/` Two categories of framework route are outside it by design, and putting a blanket credential check in front of `/eve/` in your proxy breaks both: - **Callbacks.** `/eve/v1/callback/:token` and `/eve/v1/connections/:name/callback/:token` are unauthenticated by design. An OAuth IdP arrives by 3xx redirect from a user's browser with no eve credentials attached; the unguessable token is the capability that authorizes the resume. - **Chat channels.** `/eve/v1/slack`, `/eve/v1/telegram`, and `/eve/v1/discord` each carry their own inbound verification — an HMAC signature, a secret token, an Ed25519 signature. No chat provider can answer a Basic challenge, so fronting these paths with one trades a signature check for nothing. See [Channels](/docs/channels/slack). ## Reverse proxy Two prefixes, both forwarded, neither rewritten: - `/eve/` — health, sessions, streams, channels, tools, subagents - `/.well-known/workflow/` — workflow callbacks Forwarding only `/eve/` is the most common self-hosting failure and it is not a clean one: the session starts, returns `202` with a continuation token, and then stalls forever because the run's callback cannot get back in. Nothing errors. It just never finishes. ```nginx server { listen 443 ssl; server_name agent.example.com; # No URI part on proxy_pass — that is what keeps the path unrewritten. # `proxy_pass http://127.0.0.1:2000/;` (trailing slash) strips the location # prefix and breaks both route families. location /eve/ { proxy_pass http://127.0.0.1:2000; proxy_http_version 1.1; proxy_set_header Host $host; proxy_set_header X-Forwarded-Proto $scheme; proxy_read_timeout 600s; # a turn can run for minutes } location /.well-known/workflow/ { proxy_pass http://127.0.0.1:2000; proxy_http_version 1.1; proxy_set_header Host $host; proxy_set_header X-Forwarded-Proto $scheme; } } ``` The equivalent Caddyfile — `reverse_proxy` passes the path through as received: ``` agent.example.com { reverse_proxy /eve/* 127.0.0.1:2000 reverse_proxy /.well-known/workflow/* 127.0.0.1:2000 } ``` You do not need to disable proxy buffering for nginx specifically: eve's message-stream route (`GET /eve/v1/session/:sessionId/stream`) sets `x-accel-buffering: no` and `cache-control: no-store, no-transform` on its own response. Any proxy that does not honour that header has to be told not to buffer, or streamed turns arrive all at once at the end. ## Sandbox backend `docker()` is the default, set in `agent/sandbox/sandbox.ts`. Do not use `vercel()` off Vercel — it creates hosted Vercel sandboxes from your self-hosted process, which is the one thing this stack exists to avoid. eve keys sandboxes to durable sessions and keeps one long-lived container per session, persisting `/workspace` across turns with no idle timeout. Verified on this stack: a container named `eve-sbx-ses-docker-…--__root__` came up per session, with working `bash`, `uname -s` reporting `Linux`, and a working directory of `/workspace`. It needs a Docker daemon the agent process can reach, and nothing else. **The sandbox ships with egress denied.** `EVESTACK_SANDBOX_NETWORK` defaults to `deny-all`, so a scaffolded agent's shell has no network until you say otherwise — which is the right default and also the answer to "why can't the agent `curl` anything". `resolveNetworkPolicy()` in `templates/default/agent/sandbox/sandbox.ts` treats an unrecognised value as a hard error rather than falling back to either side, because `allow_all` and `none` are typos and guessing wrong gives you either a broken shell or an open one. One further limit worth knowing before you plan network policy: the Docker backend honours only `allow-all` and `deny-all`. A domain allow-list needs a different backend. `microsandbox()` (macOS on Apple Silicon, or Linux with KVM) is one. The other is [`@evestack/sandbox-opensandbox`](https://github.com/SammyTourani/evestack/tree/main/packages/sandbox-opensandbox), which runs the sandbox on an OpenSandbox server you host instead of on your Docker daemon, and which takes a domain allow-list from 0.4.0 on — but only as `opensandbox({ networkPolicy })`, fixed for that sandbox's whole life. OpenSandbox sets a sandbox's default action when the sandbox is created and its run-time egress API preserves it, so on that backend `session.setNetworkPolicy()` rejects instead of half-applying a restriction, `subnets` (IP/CIDR) and per-domain `transform` / `forwardURL` rules throw when the backend is constructed rather than being dropped, and changing the policy afterwards means a new session rather than a re-egressed sandbox. Be precise about what that buys you, because isolation is the reason people reach for it. An OpenSandbox **server** can be configured on a gVisor, Kata or Firecracker runtime — but which runtime a sandbox lands on is the server's decision, not the adapter's: the client SDK has no runtime selector, so this backend neither requests one nor can verify it got one. On the plain Docker runtime, which is what it was tested against, you get container namespaces — the same isolation `docker()` already gives you, with a server in front of it. Do not choose it as kernel-boundary isolation unless you have independently confirmed your own server runs a secure runtime. Use Docker unless you have a specific reason not to. The one-line swap is also a [registry](/docs/registry) item, `@evestack/docker-sandbox`, if you are adding it to an existing eve project rather than starting from the template. ## The dashboard container The dashboard ships as an image, not as an npm package — npm has no way to run a Next.js application, so `@evestack/dashboard` is `private: true` and the deliverable is `ghcr.io/sammytourani/evestack-dashboard`. It is public and needs no registry login to pull. A scaffolded project's `docker-compose.yml` already points at it, pinned to the version tested with that template: ```bash docker compose --profile dashboard up -d ``` Published since `dashboard-v0.1.0` as a multi-arch manifest — `linux/amd64` and `linux/arm64`, ~230 MB compressed per platform, about 1 GB unpacked. `.github/workflows/publish-dashboard.yml` builds each arch on its own native runner and merges the digests, so neither is emulated. Check any tag yourself with `docker manifest inspect ghcr.io/sammytourani/evestack-dashboard:` and sum the layer sizes. ### Building it yourself You need this for two reasons and no others: you are changing the dashboard, or you want an image in a registry you control. Otherwise pull it. The context is the repository **root**, not `packages/dashboard`: ```bash docker build -t ghcr.io/sammytourani/evestack-dashboard:0.4.0 -f packages/dashboard/Dockerfile . ``` That is not a style preference. `packages/dashboard/package.json` declares `"@evestack/schedules": "workspace:*"`, and `workspace:` is a pnpm protocol npm does not implement; a build scoped to the package directory ends at `EUNSUPPORTEDPROTOCOL` before it installs anything. The Dockerfile therefore runs `pnpm install --filter @evestack/dashboard...` at the root, against the real lockfile, and builds `@evestack/schedules` first because its `dist/` is gitignored and a cold clone has none. Tagging it with the **published** name rather than something like `evestack-dashboard:local` is what makes a scaffolded project find it with no further configuration: Docker uses a local image when one exists under that name and only reaches for the registry when it does not. From this repository the same build is one command, because the root compose file declares both `image:` and `build:`: ```bash docker compose --profile dashboard up -d # `--profile full` is the same profile ``` To run something else entirely — a fork, a private registry, a differently-named local build — set `EVESTACK_DASHBOARD_IMAGE`. Both compose files read `${EVESTACK_DASHBOARD_IMAGE:-ghcr.io/sammytourani/evestack-dashboard:}`, so it is a one-variable change and not a topology change. The version tag tracks `packages/dashboard/package.json`. The publish workflow refuses to release when the git tag, that `version`, and the tag the scaffolder writes into generated compose files disagree — the failure it prevents is silent, since a release that publishes `:0.1.1` while every scaffold asks for `:0.1.0` looks entirely successful and is a 404 for every user. ## Expose the dashboard last The dashboard is not a viewer: it can start sessions, send follow-ups, resolve pending tool approvals, and cancel runs. Two independent controls stand in front of that, and it is worth being precise about which one does what. **The port mapping.** `docker-compose.yml` publishes `127.0.0.1:4000:4000`, so the container is reachable from the host and from nowhere else. The process *inside* the container binds `0.0.0.0` — it has to, or nothing outside the container's own network namespace could reach it, including Docker's own `HEALTHCHECK`. Exposure is decided by the mapping, not by the bind address, and `docker run -p 4000:4000` on the image by hand puts the control plane on every interface the host has. **The credential.** Every route requires `EVESTACK_AUTH_USER` and `EVESTACK_AUTH_PASSWORD`, and with either missing the dashboard serves nothing usable: `503` on every request except `GET /signin`, which renders the reason and no sign-in form, and `GET /api/health`, whose own handler answers `503 {"status":"unconfigured"}` and so reports the container unhealthy. There is no bypass flag. Browsers get a signed `HttpOnly` session cookie from `/signin`; scripts send HTTP Basic. Details in [`packages/dashboard/README.md`](https://github.com/SammyTourani/evestack/tree/main/packages/dashboard#auth). The one exception is trace ingest, whose caller is a program. `/api/ingest/v1/traces` takes `EVESTACK_INGEST_TOKEN` in an `x-evestack-ingest-token` header, and the agent must hold the same value — `create-evestack` generates it into the `.env.local` that both the host agent and the dashboard container read. Split them across hosts and you have to copy it yourself; get it wrong and every span is refused with a `401` that the agent's OTLP exporter records as a successful export, so the symptom is an empty Traces tab rather than an error. See [Observability](/docs/observability). **What the credential does not buy.** It is one shared secret per deployment, so `approver` in the audit log names an *installation*, not a person. There is no lockout, no rate limit and no second factor in front of it. It is sent as Basic over whatever transport you provide, so without TLS it is on the wire in base64. And it protects HTTP routes, not the database: the credential does nothing for Postgres, which is on loopback in both compose files and carries a per-project password only in the generated one. The repo's own compose still defaults to `${POSTGRES_PASSWORD:-evestack}`, so widening `POSTGRES_BIND` without also setting a password hands the session store to anyone who can route to the host. So: still put the dashboard behind something you already trust before it leaves loopback — a reverse proxy doing OAuth, Cloudflare Access, Tailscale, a VPN. The difference from before is that a slip in that layer is no longer immediately an unauthenticated button that can approve a shell command. For per-person attribution rather than per-installation, put a proxy that authenticates people in front and set `EVESTACK_TRUSTED_PROXY`. Until that is set, `X-Forwarded-User`, `X-Forwarded-Email` and `EVESTACK_APPROVER_HEADER` are **not read at all** — they are three words of `curl` away from anyone who can reach the port, and an audit log that can be dictated to is worse than one that admits it knows nothing. Serve it over TLS if it is reachable from anywhere but your own machine. The session cookie is `Secure` when the request is https, when `EVESTACK_PUBLIC_URL` is https, or when a trusted proxy reports `X-Forwarded-Proto: https`. If the dashboard is not on the same host as the agent, give it `EVESTACK_AGENT_URL` plus the agent's own `EVESTACK_AUTH_USER` / `EVESTACK_AUTH_PASSWORD` — it attaches Basic credentials to agent calls only when both are set, which is the same pair it signs you in with. Its required variables are now `WORKFLOW_POSTGRES_URL` *and* that credential; everything on the observe side is still a SQL read. See [The dashboard](/docs/dashboard). ## Operations Everything below is about the deployment you already have. Three things that only matter once you leave it running — supervising the agent, capping container logs, and pruning a `workflow` schema that nothing prunes for you — have a page of their own: [Operations](/docs/operations). ### Restarts State never lives in the process, so restarting is not an event. On boot against Postgres, world-postgres reclaims what was in flight and says so: ``` [world-postgres] Re-enqueued 2 active run(s) on startup ``` We proved this the hard way rather than by reading it: killed the dev server, stopped and started the Postgres container, restarted the agent, and a follow-up on a pre-restart session recalled the first message verbatim. Ordering follows from that — Postgres has to be accepting connections before the agent starts, which is what the compose healthcheck and its `service_healthy` condition enforce for the dashboard. The agent has no such condition because it is not a compose service. Under a unit it exits fast and the restart policy retries, which is the same guarantee reached differently; see [Operations](/docs/operations#run-the-agent-as-a-service) for why the restart limiter has to be disabled for that to work. ### Backups It is your database, which is the whole trade. Everything durable is in the one Postgres: | Schema | Owner | Holds | | --- | --- | --- | | `workflow`, `workflow_drizzle`, `graphile_worker` | `@workflow/world-postgres` | Durable sessions, runs, events, steps, hooks, waits, stream chunks | | `evestack` | evestack | `memories` (pgvector), `approvals`, the memory-deletion audit log, ingested spans | So a database-level dump is the backup — `pg_dump` of the whole database, not of a schema, and not a copy of the container. The `evestack-pgdata` volume is the other thing to know exists: `docker compose down -v` removes it and takes every session with it. Two schema-specific notes. `evestack.approvals` is retained forever by design — it is the row someone wants a year later, when they ask why the agent deleted the thing it deleted. And the agent's `evestack.memories` table is created lazily on the first `remember` call, so a restore into an empty database is fine; the table comes back on use, HNSW index and all. A third that is easy to assume the wrong way round: **the `workflow` schema has no retention at all.** `evestack.spans` expires on a 30-day window and prunes itself, which makes it natural to assume the rest of the database does something similar. It does not — world-postgres deletes nothing, so every run, event and step is still there. Roughly 11 MB per month at 700 sessions, measured. [Operations](/docs/operations#retention-the-workflow-schema-is-never-pruned) has the opt-in pruning procedure and, more importantly, what is safe to delete and what eve still needs. A backup you have never restored is a hypothesis. Restoring into a scratch database and pointing a throwaway agent at it costs one container and answers the question. ### The trace spool expires; your database does not eve's zero-config trace spool under `.eve/traces/v1` is written only by local dev and is bounded on purpose: `EVE_TRACES_MAX_AGE_MS` defaults to 7 days, `EVE_TRACES_MAX_TOTAL_BYTES` to 512 MB, and `EVE_TRACES_RETAIN_COUNT` keeps the newest 20 regardless. eve sweeps when a session finishes and when the dev server starts. It is a debugging buffer, not a record. Two consequences for a self-hosted deployment. The spool is not a system of record — anything you need next quarter has to be in Postgres or in a backend you operate. And authoring `agent/instrumentation.ts` at all disables that spool, so `eve traces` stops working the moment you wire up trace export; delete the file to get it back. Details in [Observability](/docs/observability). Supervising the agent, bounding the logs, and pruning a schema nothing prunes for you. Docker, Ollama, and the setup-time failures that are not deployment failures. --- # Operations > Leaving it running — supervising the agent, bounding the logs, and pruning a schema that nothing prunes for you. [Self-hosting](/docs/self-hosting) gets the stack deployed. This page is about the week after: the agent surviving a closed laptop, the logs not filling the disk, and a `workflow` schema that grows forever because nothing in eve, in `@workflow/world-postgres`, or in evestack ever deletes from it. Three things, in the order they bite. ## Run the agent as a service `docker-compose.yml` gives Postgres and the dashboard `restart: unless-stopped`. Nothing gives the agent anything: it is `npm run dev` in a terminal, or `npm start` in a terminal, and the one component that does the work is the only one that does not come back. The scaffolded project carries the fix as real files, not as snippets in a doc that drifts: | File | Host | | --- | --- | | `deploy/evestack-agent.service` | Linux, systemd | | `deploy/dev.evestack.agent.plist` | macOS, launchd | | `deploy/README.md` | the short version of this section, inside the project | ### Why not a compose service The obvious answer is an `agent:` service beside the other two. It is the wrong one here, and the reasons are specific to this stack rather than general. **The sandbox is the host's Docker daemon.** `agent/sandbox/sandbox.ts` selects eve's `docker()` backend, and that backend shells out to the `docker` CLI — `process.env.EVE_DOCKER_PATH ?? "docker"`, spawned with no shell, in eve's `dist/src/execution/sandbox/bindings/docker-cli.js` — to create one long-lived `eve-sbx-…` container per session. An agent inside a container can only do that with `/var/run/docker.sock` bind-mounted, and access to that socket is equivalent to root on the host. It would be granted to the single process in the system whose shell commands are written by a language model, which is the thing the sandbox exists to prevent. **There is no agent image.** Nothing publishes one. A compose service would therefore need `build:`, and every deploy would build your project into an image before it could start. **A supervisor needs nothing installed.** systemd and launchd are already on the host, the agent stays a normal Node process with the Docker CLI on its PATH, and `scripts/start.mjs` was already written for this: it is deliberately thinner than `scripts/dev.mjs`, runs no preflight, and exits with eve's own status code so a restart policy sees what eve saw. The compose file says the same thing, in a comment where somebody about to add an `agent:` service will read it. If you have a reason to containerise anyway — an image you build in CI, a remote `DOCKER_HOST` for the sandbox so the socket is not the local one — those reasons are real. The default should not assume them. ### systemd ```bash npm run build # eve build -> .output/ npm start # exactly what the unit runs ``` A unit that fails at boot and a build that never happened are the same line in `systemctl status`. Prove the second before you enable the first. ```bash sudo cp deploy/evestack-agent.service /etc/systemd/system/ sudo $EDITOR /etc/systemd/system/evestack-agent.service ``` Four things to edit and no more: `User`, `Group`, `WorkingDirectory` (the project root — `.env.local`, `.output/` and `.eve/` all resolve against it), and the absolute path to `node` in `ExecStart`. `command -v node` tells you the last one; a version manager's shim is a bad thing to depend on from a unit that runs at boot. ```bash sudo usermod -aG docker evestack ``` The sandbox needs it. That group is root-equivalent on the host, which is why the unit runs as a user of its own rather than as one that owns anything else. ```bash sudo systemctl daemon-reload sudo systemctl enable --now evestack-agent systemctl status evestack-agent journalctl -u evestack-agent -f ``` ### launchd Same shape, one structural difference: it is a **LaunchAgent** in `~/Library/LaunchAgents`, not a LaunchDaemon. Docker Desktop only runs while a user is logged in, so a boot-time daemon would start before the daemon it depends on and fail every sandbox command. The cost is honest — log out and the agent stops. A Mac that must serve while nobody is logged in wants the Linux unit. ```bash mkdir -p logs # launchd will not create it, and the job fails if it is missing cp deploy/dev.evestack.agent.plist ~/Library/LaunchAgents/ $EDITOR ~/Library/LaunchAgents/dev.evestack.agent.plist # every /ABSOLUTE/PATH, and node's launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/dev.evestack.agent.plist launchctl print gui/$(id -u)/dev.evestack.agent | head -20 ``` `launchctl bootout gui/$(id -u)/dev.evestack.agent` stops it, and is also how you reload after an edit. ### The two traps that only appear under a supervisor **PATH is not what you think it is.** npm prepends `node_modules/.bin` to PATH for the command it runs. systemd and launchd do not — systemd's default is `/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin`, launchd's is shorter still. Half of this is fixed for you. `scripts/start.mjs` used to spawn the bare name `eve`, so under a unit it printed *"eve is not installed in this project — run `npm install`"* on a project where eve was installed perfectly well. Measured with the real script and a stub binary: with `node_modules/.bin` on PATH it ran `eve start --port 2000`; without it, that message. It now resolves `node_modules/.bin/eve` by absolute path (`eveBinary()` in `scripts/checks.mjs`). The other half is not automatic. eve spawns **`docker`** by name for every sandbox, and on a Mac the CLI lives in `/usr/local/bin` (Docker Desktop) or `/opt/homebrew/bin` (Homebrew) — neither of which is on launchd's default PATH. The shipped plist sets PATH for exactly this; `EVE_DOCKER_PATH` names the binary outright if you prefer. Get it wrong and the failure is *partial*: the agent boots, serves, answers model calls, and fails every bash tool call with `DockerUnavailableError`. The second trap is the restart limiter. systemd's default is five starts in ten seconds, then `failed` and a wait for a human — sensible for a service that fails deterministically, wrong for one whose two boot-time failures (Postgres not accepting connections yet, a provider timing out) clear by themselves. The shipped unit sets `StartLimitIntervalSec=0`, in `[Unit]`, where it has belonged since systemd 230. Both files were written against this stack's actual requirements — each non-obvious line cites the file it comes from — and the PATH behaviour above was measured. Neither has been booted by a real systemd or launchd from this repository: the plist is checked with `plutil -lint`, the unit is not machine-checked at all. Run `systemd-analyze verify` on yours after editing it, and expect to adjust `ReadWritePaths` if your layout differs. ## Bound the logs Docker's default `json-file` driver has no `max-size` and no `max-file`. Two containers with `restart: unless-stopped` therefore append to `/var/lib/docker/containers//-json.log` until the disk is full, with no warning first — and a full disk stops Postgres, which stops everything. `docker-compose.yml` now caps both at 10 MB × 3 files through a shared `x-logging` anchor: ```yaml x-logging: &container-logs driver: json-file options: max-size: "${LOG_MAX_SIZE:-10m}" max-file: "${LOG_MAX_FILE:-3}" ``` 30 MB per container, 60 MB for the pair. The number is a choice rather than a measurement of your traffic: a mostly idle Postgres in this stack wrote 8.5 KB of log in its first eight hours, so 30 MB is weeks of quiet operation — while a container in a crash loop, which is the case that actually fills disks, is capped in minutes instead of never. `driver: json-file` is stated rather than inherited on purpose. On a host whose daemon defaults to `journald` or `local`, `max-size` and `max-file` would be silently ignored, which is the exact failure the block exists to prevent. Resize the ceiling from the `.env` beside the compose file; changing the *driver* means editing the block, because a different driver takes different options. Those two names carry no `EVESTACK_` prefix, deliberately, and neither do `POSTGRES_BIND`, `POSTGRES_PORT` or `DASHBOARD_PORT` beside them: Compose interpolates them on the host before it parses the file, so unlike the `EVESTACK_*` variables they never reach a container and no code reads them. The agent's own logs are the supervisor's problem, and the two platforms differ: - **systemd** writes to the journal, which rotates itself (`SystemMaxUse` in `journald.conf`). Nothing to do. - **launchd** writes to the file you name in `StandardOutPath` and rotates **nothing**. The shipped plist's header carries a ready-made `/etc/newsyslog.d/evestack.conf` line — 5 generations of 10 MB, gzipped — and `sudo newsyslog -nvv` checks the syntax without waiting a day. eve's own trace spool under `.eve/traces/v1` is already bounded (7 days, 512 MB, newest 20 kept) and is written only by local dev. See [Self-hosting](/docs/self-hosting#the-trace-spool-expires-your-database-does-not). ## Retention: the `workflow` schema is never pruned Everything in this section deletes conversation history permanently. There is no undo, no tombstone and no soft delete. Take a `pg_dump` first, and read [Backups](/docs/self-hosting#backups) — a backup you have never restored is a hypothesis. Half of this database expires and half does not, which is the part worth knowing before you assume either way. | Table | Retention | Driven by | | --- | --- | --- | | `evestack.spans` | 30 days by default | `EVESTACK_TRACE_RETENTION_DAYS`, applied hourly on ingest via `evestack.prune_spans` | | `evestack.alert_deliveries` | 30 days | the alert dispatcher | | `evestack.fact_turn`, `fact_tool_call` | none — but derived, and rebuildable | — | | `evestack.approvals` | **forever, by design** | it is the row someone wants a year later | | `workflow.*` | **none at all** | — | `@workflow/world-postgres` ships no retention: its only use of the word is a check that a hook's `tokenRetentionUntil` is not too *far* in the future. So `workflow_runs`, `workflow_events`, `workflow_steps`, `workflow_hooks`, `workflow_waits` and `workflow_stream_chunks` keep everything they have ever been given. Measured against a database holding one realistic month — 3,322 runs across 700 sessions: ``` workflow_events 32,994 rows 6,280 kB workflow_steps 10,046 rows 2,864 kB workflow_runs 3,322 rows 2,216 kB ───────── ~11.2 MB / month, 4.75 runs per session ``` Small, and monotonic. The point is not that it is urgent; it is that nothing ever makes it go down, and the same database has a table that does expire, so "retention is handled" is an easy and wrong conclusion to draw. ### What is safe to delete, and what eve still needs **Prune whole sessions or nothing.** Deleting old *turns* out of a session leaves eve a history with holes in it, and there is no supported way to ask whether it minds. A session is the unit a person thinks in anyway. Three facts shape the query, and each one is a way to get this wrong: Zero. Verified against a live database (`pg_constraint` has no `contype = 'f'` row in the `workflow` namespace) and in world-postgres's own migrations, which declare none. A `DELETE FROM workflow.workflow_runs` therefore orphans that run's events, steps, hooks, waits and stream chunks **silently** — they are joined by `run_id` with nothing enforcing it. Every table has to be named. A run in `pending` or `running` is work world-postgres will re-enqueue on the next boot ("Re-enqueued 2 active run(s) on startup"). It is also what an **open session** looks like: eve's session run is `workflow//eve//workflowEntry` and it stays `running` for as long as the session is open. One live run anywhere in a session protects the whole session. `$rootRunId` is set by the workflow runtime on any run started from inside another run, so it is present on turns, on subagents, **and** on the untagged `sessionTimeoutWorkflow` companion eve starts once per session — which carries no `$eve.*` attributes at all. Group on `$eve.parent` instead and you leave one orphan run, plus its events, behind for every session you prune. On the same 3,322-row database, grouping by `COALESCE($rootRunId, id)` produced exactly 700 groups, all 700 rooted in an `$eve.type = 'session'` run, with every row accounted for. Not deleted, and not by accident: - `evestack.approvals` — the audit trail. Kept forever on purpose. - `evestack.memories` — long-term memory is not session state. `forget` is the tool for it. - `graphile_worker.*` — the job queue. world-postgres and graphile-worker clean up after themselves. - `workflow_drizzle.*` — the migration ledger. Six rows; deleting them re-runs migrations. ### The procedure ```bash npm run db:prune -- --older-than=90d # dry run. Always start here. npm run db:prune -- --older-than=90d --apply # asks, then deletes npm run db:prune -- --older-than=90d --apply --yes # no prompt, for a maintenance window ``` Nothing runs it for you, and nothing schedules it. The dry run is the default, `--older-than` has no default, and `--apply` refuses off a TTY without `--yes` — a destructive command that reads stdin from a cron job either hangs forever or answers itself. A dry run reports what would go: ``` Sessions with no activity for 90 days postgres://evestack:***@127.0.0.1:5433/evestack sessions 1 runs 5 events 4 steps 2 hooks 1 waits 1 stream chunks 1 oldest 2026-01-21 15:50:07 · newest 2026-01-21 15:50:07 (UTC) ``` Other flags: `--batch=` (sessions per transaction, default 200 — the command loops, so each batch commits on its own rather than holding locks over a year of backlog) and `--keep-facts`. **"Nothing to prune" with plenty of old sessions** almost always means a run stuck in `running`. One is enough to protect its whole session, which is the intended behaviour, and a row whose status disagrees with the payload it carries is a known failure with a known repair — see [Troubleshooting](/docs/troubleshooting#the-agent-will-not-start-any-more-with-invalid-input-expected-undefined). ### The query itself `scripts/retention.mjs` builds it, so what you inspect is what runs. Every clause is load-bearing: ```sql WITH prunable AS ( SELECT COALESCE(r.attributes ->> '$rootRunId', r.id) AS session_id FROM workflow.workflow_runs r GROUP BY 1 HAVING bool_or(r.attributes ->> '$eve.type' = 'session' AND r.attributes ->> '$rootRunId' IS NULL) AND count(*) FILTER (WHERE r.status IN ('pending', 'running')) = 0 AND max(r.updated_at) < (now() AT TIME ZONE 'utc') - $1::interval ORDER BY max(r.updated_at) LIMIT $2 ), doomed AS ( SELECT r.id FROM workflow.workflow_runs r JOIN prunable p ON p.session_id = COALESCE(r.attributes ->> '$rootRunId', r.id) ), del_events AS (DELETE FROM workflow.workflow_events e WHERE e.run_id IN (SELECT id FROM doomed) RETURNING 1), del_steps AS (DELETE FROM workflow.workflow_steps s WHERE s.run_id IN (SELECT id FROM doomed) RETURNING 1), del_hooks AS (DELETE FROM workflow.workflow_hooks h WHERE h.run_id IN (SELECT id FROM doomed) RETURNING 1), del_waits AS (DELETE FROM workflow.workflow_waits w WHERE w.run_id IN (SELECT id FROM doomed) RETURNING 1), del_chunks AS (DELETE FROM workflow.workflow_stream_chunks c WHERE c.run_id IN (SELECT id FROM doomed) RETURNING 1), del_runs AS (DELETE FROM workflow.workflow_runs r WHERE r.id IN (SELECT id FROM doomed) RETURNING 1) SELECT (SELECT count(*) FROM del_runs) AS runs; ``` `(now() AT TIME ZONE 'utc')`, never a bare `now()`. These columns are `timestamp without time zone` holding UTC, so comparing one to a `timestamptz` makes Postgres reinterpret it in the **server's** zone — measured at four hours under `America/New_York`, in the direction that makes a recent session look old enough to delete. That bug shipped once already, in the dashboard's wedged-turn monitor, and `contract/contracts/21-naive-timestamps.contract.mjs` exists so it cannot again. And `--older-than=90` is not ninety days. Postgres reads a bare number as **seconds**. The command rejects it rather than guessing, and rejects `6m` too, because that reads as minutes to Postgres and as months to most people — a factor of 43,200, in the direction that deletes everything. One statement, therefore one transaction: six deletes across six unrelated tables cannot half-apply. Verified against 3,322 real rows inside a rolled-back transaction — 50 sessions selected, 245 runs, 2,442 events and 740 steps removed, and a follow-up count of events whose `run_id` no longer resolved returned zero. ### Afterwards **Derived rows.** `evestack.fact_turn` has no delete path anywhere in the dashboard — `evestack.refresh_facts` is an upsert keyed on `run_id` — so pruned runs would keep feeding every chart while vanishing from the session list. `db:prune` clears the orphans by default (they are pure derivations of `workflow_runs` and `evestack.spans`, which is what makes it safe); `--keep-facts` opts out. **Spans.** `evestack.spans` prunes on its own 30-day window. If you keep sessions for 90 days, their traces are already gone at day 31, and the dashboard labels those turns `span_coverage = 'none'` rather than pretending. Raise `EVESTACK_TRACE_RETENTION_DAYS` to keep the two aligned, and see [Observability](/docs/observability). **Disk.** Postgres does not return the space to the filesystem. Autovacuum makes it reusable, which is enough if the deployment keeps running; `VACUUM FULL` is what actually shrinks the files, and it takes an exclusive lock and needs room for a second copy of the table. Deploying it in the first place: schema, proxy, auth off Vercel, backups. Symptoms, including the stuck run that stops a prune finding anything. --- # Upgrading > Two different jobs under one name — moving a project you scaffolded onto a newer evestack, and moving this repository onto a newer eve. **Two audiences, two halves. Read the one you are.** **[Upgrading your project](#upgrading-your-project)** — you ran `npx evestack create` some months ago and want the newer dashboard, the newer template and the newer eve pin in *your* directory. Nothing about the evestack repository is involved. **[Upgrading eve inside this repository](#upgrading-eve-inside-this-repository)** — you are working on evestack itself, a contract went red, and you need to decide what that means. This is the half everything else links to: `contract/run.mjs` prints "See docs/upgrading.mdx" when the suite fails, and it means that half. This page used to be only the second one, filed under a title that read like the first. ## Upgrading your project ### What a scaffolded project actually is Worth stating before any commands, because it determines the whole shape of an upgrade: your project is a **copy**, not an install. `create-evestack` copies `templates/default` into your directory once and then has no further relationship with it. There is no `evestack upgrade` command — `packages/evestack-cli/src/` contains no such module — and nothing in the project checks for a newer version of itself. So these are the pieces, they move independently, and you move each one yourself: | In your project | What it is | How it moves | | --- | --- | --- | | `agent/`, `lib/`, `evals/`, `test/`, `tsconfig.json`, `HEARTBEAT.md` | Your code. Copied from `templates/default` at scaffold time, then yours | By hand, from a diff | | `scripts/*.mjs` — `approval-demo`, `bootstrap`, `checks`, `dev`, `eval`, `prune`, `retention`, `start`, `ui`, `verify` | evestack's helper scripts, copied the same way. Most people never edit these | By hand, from a diff | | `deploy/` — `README.md`, `evestack-agent.service`, `dev.evestack.agent.plist` | a systemd unit and a launchd plist for running the agent as a service, plus the notes for both. Copied, not generated | By hand, from a diff | | `package.json` dependencies — `eve`, `@evestack/*`, `ai`, `@workflow/world-postgres` | Ordinary npm | `npm install` | | `docker-compose.yml` — Postgres, and the dashboard behind a `dashboard` profile | Generated with the ports your machine had free, and committed | Edit the image tag | | `.env.local` and `.env` | Your generated credentials. Both git-ignored, both `0600` | Never overwrite these from a fresh scaffold | `.env.local` is read by *both* the agent on your host and the dashboard container (through `env_file:` in the compose file). `.env` is read only by Compose itself, for interpolating `${...}` in `docker-compose.yml` — Compose does not read `.env.local`. Keep the distinction in mind when an upgrade asks you to set a variable. ### Find out where you are ```bash # the eve you are pinned to node -p "require('./package.json').dependencies.eve" # the dashboard image tag your compose file names grep image: docker-compose.yml # whether the four parts are up, and where your dashboard actually is npx evestack status npx evestack open --no-open ``` Neither of those prints the running dashboard's **version** — `status` answers "is it up", `open` answers "where is it and what is the password". The version comes from `/api/health`, and the next block is how to ask it without guessing a port. **Do not reach for `curl http://127.0.0.1:4000/api/health` here.** This section's whole job is to tell you which stack you are looking at, and 4000 is not necessarily yours: the scaffolder takes the first free port at or after 4000, so a second project on the same machine is published on 4001 or later and records that in its own `.env.local`. Measured on a project whose dashboard was on 4001, the line above returned `{"ok":true,...}` from a *different* project's dashboard — a healthy answer about somebody else's stack, which is the one wrong answer this section cannot afford. `npx evestack status` and `npx evestack open` read the port out of `EVESTACK_DASHBOARD_URL` in your project, so they answer about your project. If you want the raw JSON, take the origin out of the project rather than typing a port: ```bash DASH=$(grep -m1 '^EVESTACK_DASHBOARD_URL=' .env.local | cut -d= -f2- | sed 's#/api/.*##') curl -s "$DASH/api/health" ``` `EVESTACK_DASHBOARD_URL` is written by `create` and by `attach` and holds the ingest endpoint, so its origin is your dashboard. It is the same line `evestack open` and `evestack verify` read. Compare those against a freshly published `create-evestack`: scaffold a throwaway project (next section) and read its `package.json` and `docker-compose.yml`. That pair — template plus image tag — is the combination that was tested together, which is the reason `packages/create-evestack/shared.mjs` pins the image to a tag rather than `latest`. ### Upgrade the dashboard The dashboard is a container, so this is a repull and a restart. Your compose file names it as `${EVESTACK_DASHBOARD_IMAGE:-ghcr.io/sammytourani/evestack-dashboard:}`, so you can override it without editing a committed file — the generated `.env` already carries the line, commented out: **Pick a tag that exists.** The examples below say `0.4.0` because that is what this tree pins, and `0.4.0` **is not published yet** — a manifest request for it against GHCR returns 404, so `docker compose pull` fails rather than upgrading anything. The published tags at the time of writing were `0.1.0`, `0.2.0`, `0.3.0`, `0.3.1` and `latest`. Ask the registry rather than trusting this list, which will age: ```bash curl -s "https://ghcr.io/token?scope=repository:sammytourani/evestack-dashboard:pull&service=ghcr.io" \ | sed 's/.*"token":"\([^"]*\)".*/\1/' \ | xargs -I{} curl -s -H "Authorization: Bearer {}" \ https://ghcr.io/v2/sammytourani/evestack-dashboard/tags/list ``` See [CHANGELOG.md](https://github.com/SammyTourani/evestack/blob/main/CHANGELOG.md) under *Unreleased* for why the pin runs ahead of the registry. ```bash # in .env (NOT .env.local — Compose only interpolates from .env and the shell) EVESTACK_DASHBOARD_IMAGE=ghcr.io/sammytourani/evestack-dashboard:0.4.0 ``` Then: ```bash docker compose --profile dashboard pull dashboard docker compose --profile dashboard up -d dashboard # confirm it — and read the `version`, not the `ok` DASH=$(grep -m1 '^EVESTACK_DASHBOARD_URL=' .env.local | cut -d= -f2- | sed 's#/api/.*##') curl -s "$DASH/api/health" ``` Editing the tag inside `docker-compose.yml` works too, and is the better choice if you want the version in version control. **`"ok":true` is not confirmation that the upgrade landed.** That field answers "is this process up and can it reach Postgres", and it answered `true` before the pull as well. The field that moves is `version`, which the running image reports out of its own `package.json` — so the upgrade is confirmed when the number in the response equals the tag you just pulled, and not before: ```json {"ok":true,"database":"connected","version":"0.4.0"} ``` If it still reports the old number, the container was not recreated. `docker compose --profile dashboard up -d dashboard` recreates it only when the image reference it resolves has changed; `docker compose --profile dashboard up -d --force-recreate dashboard` is the hammer. **There is no migration step, and that is by design.** A self-hosted install has no migration runner to hang one off, so the dashboard creates what it needs on first use: `evestack.spans` on the first trace read, the budget tables on the first budget write, and `sql/facts.sql`, `sql/alerts.sql`, `sql/approvals.sql`, `sql/memory-audit.sql` and `sql/query-indexes.sql` applied once per process from disk. Most statements in those files are `CREATE … IF NOT EXISTS`, `CREATE OR REPLACE` or `ADD COLUMN IF NOT EXISTS` — `sql/alerts.sql` already carries two columns added after the first release, with a backfill beside them — so **going forward** a newer image migrates itself on the first request that touches the table. **Going backward, a newer database now REFUSES an older image, and that is the most visible behaviour change on this release.** Two of those files are not idempotent in the sense the note above describes, and both are versioned: - `sql/traces.sql` carries a schema marker for `spans` and `sql/facts.sql` one for `facts`. Each file opens with a guard that raises SQLSTATE **`EV001`** when the database's marker is *ahead* of the version that file understands, and the guard is deliberately the first statement so that nothing is applied before it runs. - `sql/facts.sql` then `DROP TABLE`s all three fact tables whenever the marker is not its own version, and rebuilds them — the fact tables are a cache of a join, so rebuilding is the only strategy that cannot leave a column half-populated. That drop is exactly why the guard exists: without it an older image did not merely fail to upgrade a newer database, it dropped that database's fact tables, rebuilt them in the older shape, and stamped the marker back down to its own version without a word. So rolling the dashboard image back over a database a newer image has written does not half-downgrade it — it stops. What you see: - `/api/health` answers **503** with `{"ok":false,"status":"degraded","reason":"schema-too-new"}`, which also marks the container unhealthy in `docker ps`. The body lists which pages are unavailable, which are degraded and which still work. - `/traces`, `/sessions/[id]`, `/costs`, the overview and `/api/metrics/query` fail; `/sessions` degrades to blank fact columns; `/monitors`, `/approvals`, `/schedules` and `/evals` still work, because they read only the `workflow` tables. That is four working pages, not five — `/charts` is **not** one of them and is deliberately absent from all three lists: `app/charts/page.tsx` calls `notFound()` when `NODE_ENV` is production, so in any image you can pull it is a 404 whatever the schema says. It used to be listed here as "a static demo and stays up", which was wrong in both halves. The health response carries the same three lists, so read it rather than this paragraph if they ever disagree. - `POST /api/ingest/v1/traces` answers 503 to every batch — and the OTLP exporter reports a rejected batch as a success, so on the agent's side this looks like nothing at all. The way out is forward: run the newer image again. Dropping the `evestack` schema also works and costs you every span and every materialized fact in it. Stated from the guards and the health route rather than from a test run: the EV001 path has no automated coverage yet, because exercising it needs two dashboard images and a live Postgres. If you hit it and it behaves differently, that is worth an issue. ### Upgrade eve and the evestack packages ```bash npm install eve@0.30.8 npm install @evestack/budget@latest @evestack/composio@latest @evestack/schedules@latest npm run typecheck npm test npm run verify ``` **Do not `npm install eve@latest` here.** This line said exactly that until 2026-08-09, and `latest` is `0.31.3`, which this stack has not been released against. 0.31.3 stopped returning a continuation token from `POST /eve/v1/session` — its own docs say "session message and control request/response bodies do not accept or return continuation tokens", where 0.30.8's said the response "returns `sessionId` and the `continuationToken`". Dashboard images at `0.3.0` and earlier require that field and answer **502 — "the agent accepted the session but returned no handles"** without it. That is every new chat and every fork, on the dashboard's only route for starting a conversation. Dashboard `0.3.1` and later no longer require it and work against both, so `0.31.3` is safe once you are on `0.3.1` or newer. The pin above is the version the contract suite and the runtime probes are green against; `docs/upstream.mdx` tracks where the pin is going next. Three things to know before you run that: 1. **`eve` below `0.30.0` is a security floor, not a preference.** On 0.29.x, `localDev()` granted a full local-dev principal to anyone who sent a crafted `Host` header. The four `@evestack/*` packages that import eve declare `peerDependencies.eve` as `>=0.30.0 <1.0.0`, so npm will refuse the combination — but if you have `--force` in your muscle memory, know what it would be forcing. [SECURITY.md](https://github.com/SammyTourani/evestack/blob/main/SECURITY.md) has the detail. 2. **`@workflow/world-postgres` is pinned to an exact version, not a range and not a dist-tag.** That is deliberate, and it is the second pin this dependency has had. npm `latest` is on the 4.x line, which eve rejects outright, so the template used to say `"beta"`. A plain `npm install` re-resolves a dist-tag, so that dependency moved under people without any version number in their `package.json` changing — and when upstream raised the World spec version from 5 to 6 inside `5.0.0-beta.*`, every fresh install died at boot with `This Workflow runtime requires a World with matching spec version 5`. The template now declares `"5.0.0-beta.32"` exactly. If you are upgrading a project scaffolded before that, change your own `package.json` to the exact version too; `^` and `~` do not help, because both still admit `5.0.0-beta.34`. If sessions start failing right after an unrelated install, check this first. 3. **`npm run verify` is the real gate**, not `npm test`. Its own header names what it walks: "Postgres, Docker, the agent, the dashboard, an embedding probe" — plus the model key and the schema — against the stack as it is actually running, with the fixing command on every red line. It exits 1 if anything required failed, so you can put it in a script. If a new eve changed the workflow schema, `npm run db:bootstrap` is what applies it. That script is a thin wrapper around `@workflow/world-postgres`'s own setup script — it checks the connection first so a failure names a host and a port instead of a Drizzle stack trace, then hands over unchanged. The migrations are upstream's. ### Pick up template changes This is the step with no automation, so here is the routine that works: ```bash cd /tmp npx create-evestack@latest reference-agent ``` Answer **no** when it offers to bring the stack up — you only want the files. It generates its own credentials, which you will not be copying anywhere. ```bash diff -ru ~/my-agent/scripts /tmp/reference-agent/scripts diff -ru ~/my-agent/deploy /tmp/reference-agent/deploy diff -u ~/my-agent/package.json /tmp/reference-agent/package.json ``` `scripts/` is the highest-value diff and the safest to take wholesale: unless you have edited them, those eight files are evestack's, and `verify.mjs` in particular gains checks as new failure modes are found. `deploy/` is the systemd unit and the launchd plist, and is the one to read rather than copy — both carry absolute paths you filled in for your machine. A project scaffolded before those files existed will show them as additions. Take the `package.json` diff as information rather than as a patch — yours has your project name, and may have dependencies you added. ```bash diff -ru ~/my-agent/agent /tmp/reference-agent/agent diff -ru ~/my-agent/lib /tmp/reference-agent/lib ``` Expect noise here — this is your agent's instructions, tools and channels, and you have presumably changed them. What you are looking for is a *shape* change: a new file, a changed import path, a rewritten auth chain in `agent/channels/eve.ts`. Apply those by hand. ```bash diff -u ~/my-agent/docker-compose.yml /tmp/reference-agent/docker-compose.yml ``` It is generated with the ports *your* machine had free and a project name derived from your directory, so copying it over will point the stack somewhere else. Read it for a new service, a new environment variable, or a new mount, and transplant just that. It holds a generated dashboard password, a trace-ingest token and a Postgres password that are now lying around in `/tmp`. ```bash rm -rf /tmp/reference-agent ``` ### Upgrade the CLI itself `npx evestack@latest …` pins nothing and resolves the newest published version, so there is nothing to maintain. A **global** install does not update itself: ```bash npm i -g evestack@latest ``` ### The order that matters The dashboard image and the agent template are versioned separately and released together, and the scaffolder pins one image tag per template version precisely because that is the combination that was tested. Moving one a long way without the other is untested rather than forbidden. If you are several releases behind, move both, from the same release, and run `npm run verify` afterwards. Nothing here rewrites your data. The one destructive operation in this area is `docker compose down -v`, which deletes the Postgres volume — every session, trace and memory with it — and it is never part of an upgrade. It is only ever needed to make a *new* database password take effect (`EVESTACK_DB_PASSWORD` in `.env`, which Compose interpolates into `POSTGRES_PASSWORD`), because Postgres applies that variable once, when the volume is first created. --- ## Upgrading eve inside this repository Everything below is for work on evestack itself. If you are running a scaffolded project, you want the half above. Vercel merged 252 pull requests into eve in fourteen days. eve shipped 0.30.0, 0.30.1 and 0.30.2 on the same day evestack launched. That cadence is the environment evestack lives in, and it has one practical consequence: Anything evestack builds inside eve's own surface area has a shelf life measured in weeks. Assume every assumption below is temporary and write it down where a machine can check it. ### The policy 1. **Every assumption evestack makes about eve is a contract in `contract/`.** If you find yourself writing "eve does X, so we can do Y", that belongs in the suite before it belongs in the code. 2. **`eve-watch` opens the pull request. A human merges it.** There is no auto-merge and there will not be one. 3. **A red contract is a decision, not a bug.** Sometimes eve is wrong, sometimes we are, sometimes both are right and the contract has gone stale. Those three outcomes need different fixes. 4. **Deterministic checks decide; a model may advise.** See [On putting a model in CI](#on-putting-a-model-in-ci). ### Why a typecheck is not enough evestack once shipped a `strictLocalDev()` wrapper. eve 0.29.x decided "is this request from my own machine" by matching the request's own hostname against an unanchored `/^127\./`, and a request URL is built from the client's `Host` header. So `127.evil.com` — a name anyone can register and point at your agent — received a full local-dev principal with no credentials at all. We measured it: that host answered 200 where a plain foreign host answered 401. eve 0.30.0 fixed it properly upstream. `localDev()` now grants based on the process being an `eve dev` run, and consults nothing in the request. At that moment our wrapper stopped adding protection and started rejecting legitimate local-dev access over a LAN IP, a tunnel, or a container hostname. It had to be deleted. **It typechecked perfectly on both days.** `tsc` had nothing to say when it was load-bearing and nothing to say when it became harmful, because its *types* never changed — only eve's *meaning* did. That is the entire argument for the contract suite. A typecheck cannot catch semantic drift. A behavioural assertion can: ```bash # green against the eve we ship node contract/run.mjs --only=auth # red against the eve that had the bug — naming the exact hostile Host header EVESTACK_CONTRACT_EVE_DIR=node_modules/.pnpm/eve@0.29.5.../node_modules/eve \ node contract/run.mjs --only=auth ``` ### When the suite goes red Read the failure first. Every contract prints the assumption it pins and what evestack does with it, so the report tells you the blast radius before you open a single file. ```bash pnpm install pnpm contract node contract/run.mjs --only= --verbose ``` The suite is free and offline. There is never a reason to debug this from CI logs alone. | What you find | What it means | What to do | |---|---|---| | eve changed behaviour we depend on, deliberately | The contract did its job | Update evestack's code, then update the contract to describe the new behaviour — **same commit** | | eve changed behaviour and it looks like a regression | The contract found an upstream bug | File it at `vercel/eve`, pin the previous version, leave the contract red with a link | | eve is unchanged; our contract was over-specified | The contract is wrong | Loosen it to assert what we actually depend on, not what we happened to observe | A deleted contract is an assumption that silently stopped being checked, and the next person has no way to know it was ever true. Loosen it, or replace it with the narrower thing we genuinely rely on. If evestack really no longer depends on it, delete the dependency in the same commit and say so in the message. The pull request reports the suite against the **current** pin as well as the candidate. If that column is red, `main` was broken before this upgrade and the bump is a distraction — fix `main` first. ### When it stays green and something breaks anyway This is the failure mode worth planning for. A green suite means *every assumption someone wrote down still holds* — not that the release is safe. eve can change something we depend on that no contract names, and the suite will report success with total confidence. When that happens, the fix is not just the patch. **Write the contract that would have caught it**, in the same pull request. That is the only mechanism that makes the suite better over time, and it costs about fifteen minutes while the failure is still fresh. `contract/README.md` describes the shape. ### What is currently pinned | Contract | Assumption | |---|---| | `version/…` | One eve version satisfies every range this repo declares — templates, both published packages' peer ranges, the scaffolder | | `modules/…` | Every `eve/*` subpath evestack imports resolves, and every value it binds is still exported | | `tools/approval-…` | `approval` is the gating field (not the AI SDK's `needsApproval`), and `always()` still returns `"user-approval"` | | `tools/dynamic-…` | `defineDynamic` still produces `kind: "eve:dynamic"`, and `step.started` is still dispatched | | `protocol/…` | The session/cancel/stream routes, the NDJSON content type and the stream resume headers the dashboard drives | | `attributes/…` | Every `$eve.*` run attribute the dashboard reads out of Postgres is still one eve writes | | `auth/…` | `localDev()` grants on process state and never on the request; `httpBasic()` and `routeAuth()` fail closed | | `hooks/…` | Hook handlers return `void` — they cannot block or park a turn | | `sandbox/…` | The `SandboxBackend` shape `@evestack/sandbox-opensandbox` duck-types, which nothing typechecks | The import and attribute lists are derived from evestack's own source at run time, so they widen automatically as the codebase grows. ### The `eve-watch` workflow `.github/workflows/eve-watch.yml` runs daily and on demand. It asks npm for the latest eve; if it is newer than the pinned range it creates a branch, bumps every manifest with a caret pin, installs, typechecks, runs the contract suite, and opens a pull request reporting all of it alongside the eve changelog entries that touch something the suite pins. Two details are deliberate: - **Peer ranges are not bumped.** `peerDependencies.eve` on the published packages is a compatibility promise to users. Widening it automatically would publish support for a version nothing has been tested against. If a new eve falls outside one, the version contract fails and a human decides. - **The baseline runs first.** The suite runs against the current pin before the bump, so a red result after the bump is unambiguous. Run it by hand against a specific version when you want to test a release candidate: ``` Actions → eve-watch → Run workflow → version: 0.31.0 ``` ### On putting a model in CI The honest engineering read, since it comes up every time. **Where a model genuinely helps.** Reading a hundred-line changelog and telling a maintainer which three entries matter to a self-hosted, non-Vercel deployment is a real task with no deterministic solution, and a model is good at it. So is drafting a diff for a mechanical rename once a human has decided the rename is correct. Both are *suggestions to a person who is already reading the PR*. **Where it is dangerous.** Auto-merging a model's fix to an auth path. The `localDev()` story is the argument: the correct response to that change was to *delete* protective code, and deleting security code on a model's say-so, on a schedule, with nobody watching, is how you ship an authentication bypass to everyone who ran `npx create-evestack`. A model asked "does this patch still help?" would very plausibly have said yes — the patch looked defensive and compiled fine. There is also a supply-chain problem. A changelog is third-party text that arrives on a schedule. A model reading it inside a job that can write to the repository is a prompt-injection target with commit access. The mitigation is not a better prompt; it is not giving it commit access. **So the rule is:** Deterministic checks are the mechanism. A model is a convenience layer on top — optional, gated on a secret being present, and able only to comment. The optional step in `eve-watch` follows exactly that. It runs only if `OPENAI_API_KEY` is set, it reads results the deterministic steps already produced, and its single capability is `gh pr comment`. It cannot commit, push, merge, re-run the suite, or change what the PR reports. The model it uses is the repository variable `ADVISOR_MODEL`, defaulting to `gpt-5-mini` — set it under *Settings → Secrets and variables → Actions → Variables* to change it without touching this workflow. The same key drives the nightly evals, so one secret covers both jobs. Its output is labelled as generated and non-authoritative. If it is prompt-injected by a hostile changelog, the worst outcome is a misleading paragraph next to the real contract table — which a human is required to read anyway. If you do not set the secret, everything above still works. That is the test of whether a model is a convenience layer or a dependency, and this one passes it. #### The two questions it is actually asked Not "summarise this PR" — the deterministic steps already did that, better. It is asked the two things they structurally cannot answer: **1. Did a check pass or fail without reaching its subject?** `deny-survives` parks a turn and denies a gated tool call. If the model under test answered in prose and never called the tool, the run ends `waiting` with no pending requests and the denial was never exercised at all. That is a *vacuous* run, not a regression, and the difference decides whether a maintainer re-runs or investigates. It is also not expressible as an assertion: the eval cannot tell "the bug is gone" from "the test never got there" without judging what the transcript means. Naming vacuity is the single change that makes a nondeterministic gate readable enough to keep. The output *annotates*. It can never clear a result — a red stays red no matter what the paragraph says. **2. What did this release change that no contract asserts?** The suite covers what somebody thought to write down, which is why every red contract's advice ends with "the fix is a new contract, not just a patch." That instruction has always resolved to *when a human happens to notice*. The step is now handed a listing of every contract and every assertion it makes, and asked to name surface the release touched that nothing pins — writing the assertion as one sentence of prose, in a fenced block, for a human to accept or ignore. It does not write code, name a file, or open anything. The ratchet is that the suite has a mechanism for growing that is not "remember to." Both inputs — the changelog and the release notes — are third-party text on a schedule, so the prompt requires anything that appears to address or instruct the model to be quoted verbatim under `SUSPICIOUS TEXT` and not acted on. That matters most for exactly the sentence a compromised release would want to write: *"downstream repos should drop the Host-header assertions."* ### Running the suite yourself ```bash pnpm contract # all contracts node contract/run.mjs --verbose # show passing assertions node contract/run.mjs --only=auth # one family node contract/run.mjs --format=json # for scripts # point it at any other eve install EVESTACK_CONTRACT_EVE_DIR=path/to/eve node contract/run.mjs ``` `pnpm test` runs the contract suite before the workspace tests, so it is hard to skip by accident. --- # Uninstall > Removing a project, and the four things that survive deleting its directory — with the exact names, and the two list commands that lie to you. **There is no `evestack uninstall`.** `evestack --help` lists seven commands and none of them removes anything. Nothing here is automated, and that is the honest state of it rather than a design position — this page exists because removal was undocumented and the evidence was five orphaned Postgres volumes on one development machine, four of them belonging to projects whose directories no longer existed. Everything below **destroys data and does not ask twice**. `docker compose down -v` deletes the Postgres volume: every session, every trace, every memory, every approval record and every schedule in that project, gone, with no undo and no backup taken for you. ## What removal actually involves A project is not installed anywhere. It is a directory, plus things Docker holds on its behalf, plus caches that belong to other tools. Deleting the directory removes the first of those and none of the rest. | Thing | Who owns it | Removed by deleting the project directory? | Measured size | | --- | --- | --- | --- | | the project directory (`node_modules`, `.eve/`, your agent code) | you | yes | ~300 MB, mostly `node_modules` | | the `_evestack-pgdata` volume | Docker | **no** | 68–78 MB light, 147.2 MB measured on a busy one | | `ghcr.io/sammytourani/evestack-dashboard:` | Docker | **no** | 1.05 GB unpacked | | `pgvector/pgvector:pg17` | Docker | **no** | 646 MB | | the eve base image | Docker | **no** | 665 MB | | `eve-sandbox-template:`, one per project | Docker | **no** | 665 MB apparent, **~66 kB unique** per extra one | | session sandbox containers | Docker | **no** | ~20 kB writable layer each, but they persist | | `~/.npm/_npx` | npm | **no** | 99 MB when this was written; 2.2 GB on the same machine later | | Ollama models (`qwen3`, `nomic-embed-text`) | Ollama | **no** | several GB | The last two are not evestack's to delete and are named here only because nothing else names them as a consequence of having run evestack. **Read that last column as an order of magnitude, not as your machine.** Every figure in it was measured, but three of them move: the volume grows with your sessions, `~/.npm/_npx` grows with every `npx` invocation of anything, and the dashboard image changes size per tag. The commands in each section below print the real number for *your* machine, which is the one to act on. The sandbox-template row is the one worth reading carefully, because `docker images` reports it misleadingly. Each `eve-sandbox-template` shows a **665 MB** size, but that is almost entirely layers shared with the eve base image; `docker system df -v` breaks the two apart, and the `UNIQUE SIZE` column is what you actually get back: ``` REPOSITORY SIZE SHARED SIZE UNIQUE SIZE eve-sandbox-template 665MB 665.3MB 65.77kB eve-sandbox-template 665MB 665.3MB 65.77kB ghcr.io/vercel/eve 665MB 665.3MB 33.42kB ``` So removing three orphaned sandbox templates recovers about **200 kB**, not 2 GB. Removing the eve base image underneath them is where the 665 MB is — and nothing does that while a sandbox template still references those layers. ## Two list commands that lie, and the rule that follows `docker ps --filter name=X` and `docker volume ls --filter name=X` are **substring** matches, not equality. `--filter name=agent` matches `my-agent`, `agent-2`, and a container called `payments-agent-prod` that belongs to something else entirely. This is not theoretical here: a bulk delete driven by one of those filters destroyed a live container in this project's own history. The data was recoverable; it should not have had to be. **So every removal on this page is two steps, and the first one is a list.** Read the whole output, decide with your eyes, then remove by exact name. Do not pipe a `--filter name=` result into `rm`, `rmi` or `docker rm`, in a script or otherwise. ## 1. Take one project down, from inside it Do this **before** you delete the directory. The compose file is the only thing that knows the project's Compose identity, and once it is gone Docker will not associate the leftovers with it — that is what produced the four orphans. **First: which compose file is this?** Everything below runs `docker compose` with no `-f`, which means *whatever `docker-compose.yml` is in the directory*. In a project from `evestack create` that is evestack's file and this is correct. In a project from **`evestack attach`** it may not be: when the project already had a `docker-compose.yml` of its own, attach refuses to overwrite it and writes **`docker-compose.evestack.yml`** beside it instead. Bare `docker compose` there reads *their* file, resolves *their* project name, and `down -v` deletes *their* volumes. Measured, in a directory holding both files: ``` $ docker compose --profile dashboard config --volumes their-precious-data ``` Not one evestack resource in the answer. So look before anything else: ```bash ls docker-compose*.yml ``` If `docker-compose.evestack.yml` is there, **add `-f docker-compose.evestack.yml` to every `docker compose` command on this page**, and read [attach's own dashboard](#the-attached-dashboard-is-not-a-compose-service) below — an attached project's dashboard is not a Compose service at all. ```bash cd ~/my-agent # The project name and the real volume name, both as Docker stores them. # `config --volumes` alone prints `evestack-pgdata`, which is the key inside the # file and is the same string in every evestack project — not a name you can act on. docker compose --profile dashboard config --format json \ | node -p "const c=JSON.parse(require('fs').readFileSync(0,'utf8')); \ c.name + '\n' + Object.keys(c.volumes||{}).map(v => ' ' + c.name + '_' + v).join('\n')" # and the containers, named exactly docker compose --profile dashboard ps -a ``` Read both. The volume name that command prints is the one — and the only one — that the rest of this page is about. Then, keeping the data: ```bash # containers and network go; the Postgres volume stays docker compose --profile dashboard down --remove-orphans ``` Or destroying it: ```bash # the same, plus the volume — every session, trace, memory, approval and schedule docker compose --profile dashboard down -v --remove-orphans ``` `--profile dashboard` matters, and it is not optional. Without it Compose does not consider the dashboard service part of this run and its container is left behind — **and `--remove-orphans` does not catch it**, because Compose knows that service, it just filtered it out. Measured on Compose v5.1.0, against a two-service project whose second service is profile-gated: ``` $ docker compose down --remove-orphans Network probe_default Removing Network probe_default Resource is still in use $ docker ps -a --format '{{.Names}}' probe-dashboard-1 ``` The container survives, the network cannot be removed because the container is still attached to it, and nothing said the word "orphan". Re-run the same command **with** `--profile dashboard` and both go. ### The attached dashboard is not a Compose service `evestack attach` writes a compose file for **Postgres only**. Its dashboard is started by the `docker run` command attach prints, so no `docker compose down` of any kind touches it — it is not in the file, it carries no Compose project label, and it will still be there after the compose project is gone. It is named after the same Compose project name the volume is, with `-dashboard` on the end, so you can find it exactly: ```bash # list first — read the name, then use the name you read docker ps -a --format 'table {{.Names}}\t{{.Image}}\t{{.Status}}' | grep -- -dashboard docker rm -f ``` Finally the directory: ```bash rm -rf ~/my-agent ``` `.env` and `.env.local` in there hold a generated Postgres password, a dashboard sign-in password and a trace-ingest token, all at mode `0600`. `rm -rf` is enough — none of them is registered anywhere else — but if you keep a backup of the directory, those three secrets are in it. ## 2. If the directory is already gone The volume name is the only surviving reference, and it is derived rather than random: ``` _evestack-pgdata ``` where the Compose project name is the directory's basename, slugged, plus the first six hex characters of the SHA-256 of the directory's **absolute path**. The path is in there deliberately — two projects both called `my-agent` would otherwise be one Compose project sharing one database, which happened twice before the hash was added. It also means **moving a project directory orphans its volume**, which is a second way to arrive at this section. List what is there, read it, and remove by full name: ```bash # every evestack Postgres volume on this machine, with its size docker system df -v | grep evestack-pgdata # or the full picture, unfiltered, if you want to see everything docker volume ls ``` If you remember where the project used to live, you can name its volume exactly: ```bash node -p "const p=require('node:path').resolve('/Users/you/old-project'); \ (p.split(/[\\\\/]/).filter(Boolean).pop().toLowerCase().replace(/[^a-z0-9_-]+/g,'-').replace(/^[^a-z0-9]+/,'')||'evestack') \ + '-' + require('node:crypto').createHash('sha256').update(p).digest('hex').slice(0,6) + '_evestack-pgdata'" ``` For the path in that example it prints exactly this, and you can check the derivation by running it yourself: ``` old-project-c94571_evestack-pgdata ``` **Check it against the list above before you act on it.** The one-liner reconstructs a name; it does not confirm one exists. If the name it prints is not in `docker volume ls`, do not go looking for the nearest match — the two ways to get a near miss are a path that is not byte-for-byte the one the project was scaffolded at (a symlinked `/tmp`, a different capitalisation, a trailing slash: `create` hashes the argument it was given, and `resolve()` here strips a trailing slash it may not have), and a project that was scaffolded somewhere else entirely. Both mean you have the wrong directory, not the wrong volume. Then: ```bash docker volume rm old-project-c94571_evestack-pgdata ``` `docker volume rm` refuses a volume a container still uses, which is a useful second opinion: if it refuses, something is still running and you are about to delete a live project's database. ## 3. The images, once no project needs them Only do this when you have removed every evestack project on the machine — these are shared, so removing them while another project exists just means it pulls or rebuilds them again. ```bash # list first, with sizes, and read it docker images ghcr.io/sammytourani/evestack-dashboard docker images pgvector/pgvector # then remove the exact tags you saw — substitute the tags from YOUR output, # these two are only the shape of the command. # # THE-TAG-YOU-SAW is deliberately not a version number. Do not "fix" it to one: # the publish gate scans the tree for evestack-dashboard:… and requires # every hit to equal packages/dashboard/package.json, so a real version here is # either a duplicate of the pin or a release blocker. It was 0.3.1 once and # would have failed the 0.4.0 release at tag-push time. A placeholder also fails # loudly if pasted verbatim — docker answers "No such image" instead of removing # whichever version you happened to have. docker rmi ghcr.io/sammytourani/evestack-dashboard:THE-TAG-YOU-SAW docker rmi pgvector/pgvector:pg17 ``` **`pgvector/pgvector:pg17` is a public Postgres image, not an evestack one.** Anything else on the machine may be using it — another project's database, something you pulled by hand — and `docker rmi` on it takes it away from them too. `docker rmi` refuses while a container still references it, which catches the common case, but a stopped project you intend to start again is not protected. Check what depends on it before removing it: ```bash docker ps -a --filter ancestor=pgvector/pgvector:pg17 \ --format 'table {{.Names}}\t{{.Status}}' ``` If that lists anything you did not scaffold, leave the image alone. It costs 646 MB and it will be re-pulled on demand; a database you cannot start is worth more than the disk. Naming the repository as an argument, rather than `--filter name=`, is the point: `docker images ` matches that repository and nothing else. ### The per-project sandbox image eve builds an `eve-sandbox-template:` image per project on first run and **nothing removes it** — not `docker compose down -v`, not deleting the project directory, not any evestack command. Three orphaned ones were on the development machine when this was written. The commands are in [Local setup](/docs/local-setup#what-grows-and-the-one-thing-nothing-cleans-up); the short version is that `--filter reference='eve-sandbox-template:*'` matches the *whole* reference as a glob, so it names that repository exactly and leaves only the tag open — unlike `--filter name=`. Do it when no project is running. The tag carries no hint of which project it belongs to, and a live project simply rebuilds its own on the next turn. ### Session sandbox containers The Docker sandbox keeps one long-lived container per durable session, with no idle timeout, so they accumulate for as long as sessions do. List them and read the list before removing anything: ```bash docker ps -a --format 'table {{.Names}}\t{{.Image}}\t{{.Status}}' ``` ## 4. The caches that are not evestack's **`npx`.** Every `npx create-evestack` and `npx evestack` leaves a full install under `~/.npm/_npx` — 99 MB on a machine that had run the scaffolder a few times, 2.2 GB on the same machine after a month of `npx` use of everything else. `npm cache clean --force` does not touch it. It is shared with every other tool you run through `npx`, so `du -sh` first and expect most of it not to be evestack's. ```bash du -sh ~/.npm/_npx # look first rm -rf ~/.npm/_npx # safe: it is a cache, and npx refills it on demand ``` **Ollama.** If you took the $0 local-model path, the models are still on your machine and nothing in evestack references them. They are yours and may be in use by something else, so list before removing: ```bash ollama list ollama rm qwen3 ollama rm nomic-embed-text ``` **The globally installed CLI**, if you installed one rather than using `npx`: ```bash npm ls -g --depth=0 | grep evestack # confirm it is there and named exactly npm rm -g evestack ``` ## 5. Check your work Nothing evestack-shaped should be left. Each of these prints nothing on a clean machine: ```bash docker ps -a --format '{{.Names}}\t{{.Image}}' | grep -E 'evestack|pgvector|eve-sandbox' docker volume ls --format '{{.Name}}' | grep evestack-pgdata docker images --format '{{.Repository}}:{{.Tag}}' | grep -E 'evestack|pgvector|eve-sandbox' ``` **These three are a question, not a to-do list, and the difference matters most here.** They are deliberately wide — wider than evestack — so that nothing hides from them. `pgvector` in particular matches on the *image* column, so every container on the machine running `pgvector/pgvector` prints, whoever owns it. Run on the machine this page was written on, the first line printed a container belonging to no evestack project at all: ``` evestack-audit-pg pgvector/pgvector:pg17 evestack-postgres-1 pgvector/pgvector:pg17 evestack-dashboard-1 evestack-dashboard ``` Only the second and third are a scaffolded project's. The first is a database somebody started by hand and is still using. So do **not** read "whatever these print, delete it". Read each line, decide whether it is one of the projects you removed above, and act only on the ones you recognise — by the exact name printed, one command at a time. A line you cannot account for is a reason to stop and look, not a reason to run `rm`. **What is deliberately not here:** a script. Every candidate one has to select things by pattern, and a pattern is the failure mode this page is written around. The lists above are short enough to read, and reading them is the safeguard. --- # Support > Which versions get fixes, what the version numbers promise, and which platforms are actually tested — including the one that is not. This page is deliberately unambitious. Everything on it is a statement about what CI runs and what the manifests declare, so you can check it yourself rather than take it on trust. Where something is untested, it says so. ## Which versions get fixes The newest published version of the affected package. Nothing else. Fixes land on `main` and go out in that package's next release. There is no long-term-support line, no maintenance branch per version, and nothing is backported. If you are two versions behind and hit a bug, the answer will be "upgrade" — see [Upgrading](/docs/upgrading). Three things ship on separate clocks, which matters when you read "fixed in": | What | How it reaches you | Consequence | | --- | --- | --- | | The npm packages (`evestack`, `create-evestack`, `@evestack/budget`, `@evestack/composio`, `@evestack/schedules`, `@evestack/mcp`, `@evestack/sandbox-opensandbox`) | `npm install` | Ordinary. Your lockfile decides. | | The dashboard | A container tag, `ghcr.io/sammytourani/evestack-dashboard:` | Your project pins a tag, so nothing changes until you repoint and repull | | The agent itself (`templates/default`) | Copied into your project once, at scaffold time | Never updates on its own. Fixes to it are applied by hand | That third row is the one that surprises people. `create-evestack` copies the template into your directory and then has no further relationship with it. There is no `evestack upgrade` command — `packages/evestack-cli/src/` has no such module — so [Upgrading](/docs/upgrading) describes the diff-and-apply routine that stands in for one. ## What the version numbers promise Every published package is `0.x`. Under npm's own rules a caret range on a `0.x` version only admits patch releases — `^0.2.0` will take `0.2.1` and will not take `0.3.0` — and that is the promise being relied on here rather than worked around. - **A minor bump (`0.2.x` → `0.3.0`) may break you.** Pre-1.0, that is what the minor position is for. - **A patch bump is fixes only.** - **No package has reached 1.0**, so nothing here is claiming API stability yet. One compatibility promise is machine-checked rather than declared. `@evestack/budget`, `@evestack/composio`, `@evestack/schedules` and `@evestack/sandbox-opensandbox` each declare `peerDependencies.eve` as `>=0.30.0 <1.0.0`, so npm refuses to install them beside an eve that is too old to have the `localDev()` fix ([SECURITY.md](https://github.com/SammyTourani/evestack/blob/main/SECURITY.md) explains why that version in particular). `contract/contracts/01-version.contract.mjs` fails the suite if one installed eve cannot satisfy every range this repo declares, and the `eve-watch` job is forbidden from widening a peer range on its own — see [Upgrading](/docs/upgrading). `evestack`, `create-evestack` and `@evestack/mcp` declare no eve peer range, because they do not import eve. The scaffolder ships the template; the CLI and the MCP server talk to a dashboard over HTTP. ## Platforms ### Node **Node 24, and the current major alongside it.** All seven published packages declare `"engines": { "node": ">=24" }`, as do the repository root and the template a scaffold gets. CI installs 24 everywhere and 26 as well on the two jobs cheap enough to matrix: `typecheck` and `registry` in `.github/workflows/ci.yml` run `node-version: [24, 26]` with `fail-fast: false`, and the remaining five `actions/setup-node` steps pin 24. So 24 is the floor and 26 is checked; anything between them is inferred, not tested. The floor is enforced rather than merely advertised: `nodeVersionProblem()` in `packages/create-evestack/shared.mjs` refuses an older runtime with a message naming the version, because `engines` is advice that npm prints and installs through anyway. The generated project's scripts use `--env-file-if-exists` (Node 20.12+) and its checks call `URL.parse` (Node 22.1+), so an older Node fails several commands later in a message that never names the real cause. ### Operating systems | Platform | Status | What that is based on | | --- | --- | --- | | Linux `x86_64` | Tested | Every job in `ci.yml` bar the non-blocking `macos` one runs on `ubuntu-latest`, including the runtime tier that boots Postgres, an agent and a Docker sandbox | | Linux `arm64` | Image only | `publish-dashboard.yml` builds `linux/arm64` on an `ubuntu-24.04-arm` runner. The multi-arch image is exercised; the CLI and template are not tested there by CI | | macOS | Partly tested | A non-blocking `macos-latest` job in `ci.yml` runs typecheck, the static contract tier and the POSIX file-mode tests. See below for what it deliberately leaves out | | **Windows** | **Untested** | See below | ### macOS A `macos-latest` job runs on every PR. It is marked `continue-on-error: true`, so it reports without being able to block a merge — the table above says *partly* tested, and that is the honest word for it. **What it runs.** `pnpm -r typecheck`; the static contract tier (`node contract/run.mjs` — 545 assertions, under two seconds on darwin); `create-evestack`'s suite, whose file-mode assertions gate on `process.platform !== "win32"` and check that generated credential files land at `0600`, which is the likeliest place for an APFS-versus-ext4 difference to surface on a security-relevant path; and the template's own suite, which resolves `node_modules/.bin/eve` under a scrubbed PATH and spawns without a shell. **What it deliberately leaves out.** Anything that needs a Docker daemon. Five jobs across this repo's workflows — `runtime` in `ci.yml`, plus `provider-bisect.yml`, `evals.yml`, `eve-watch.yml` and `dashboard-image.yml` — declare a `services:` block, and GitHub-hosted macOS runners have no daemon to back one. Porting them would mean dropping the contract runner's `--require` flag, which downgrades a probe that cannot run from a failure to a skip: green having checked nothing. Playwright is out for the same class of reason — the install step uses `--with-deps`, which is an `apt-get` path. **What is still uncovered.** There are four `darwin` branches in the tree, not three: the `open`-a-browser helpers at `packages/evestack-cli/src/project.mjs:354`, `packages/evestack-cli/src/tour.mjs:393` and `templates/default/scripts/verify.mjs:569`, plus `contract/runtime/repro/eve-turn-wedge.mjs:162`, which shells out to `sysctl -n vm.swapusage`. The macOS job executes none of them — it launches no browser, and the fourth lives in the runtime tier that does not run there. The job narrows the macOS gap; it does not close it. ### Windows **Windows is untested. There is no CI job on any Windows runner, and there never has been.** Twelve `process.platform === "win32"` branches ship in the code, and only one of them is covered by a test — and that one is covered by passing `"win32"` in as an argument, not by running on Windows. Use WSL2, where the Linux paths run. This is stated bluntly because the code *looks* like it supports Windows. The twelve branches, so you can judge the risk yourself: - **Opening a browser** — `packages/evestack-cli/src/project.mjs:357`, `packages/evestack-cli/src/tour.mjs:396`, `templates/default/scripts/verify.mjs:651` each pick `cmd /c start` over `open` / `xdg-open`. Find them with `rg -n "cmd\", \[\"/c\", \"start\"" packages templates` rather than by line number. - **Spawning `npm` and `docker`** — `packages/create-evestack/create.mjs:231`, `:354` and `:1185`, `templates/default/scripts/dev.mjs:81`, `templates/default/scripts/start.mjs:60`, `templates/default/scripts/eval.mjs:148` pass `shell: process.platform === "win32"`, which is how a `.cmd` shim gets found on Windows. Find them with `rg -n 'shell: process.platform' packages templates`. - **Naming the `eve` binary** — `templates/default/scripts/checks.mjs:532` returns `eve` rather than `eve.cmd`. This is the one with a test: `templates/default/test/eve-binary.test.mjs:115` calls `eveBinary(url, "win32")` on Linux, which checks the string and nothing about Windows. - **Terminal glyphs** — `packages/create-evestack/ui.mjs:68` and `templates/default/scripts/ui.mjs:68` fall back to ASCII outside Windows Terminal. That file's own comment is the honest summary: the failure mode is "mojibake in a legacy Windows code page — one nobody here can reproduce." `packages/create-evestack/test/attach-writes.test.mjs` gates its file-permission assertions on `process.platform !== "win32"`, so even if you ran the suite on Windows, the checks on credential file modes would skip rather than fail. If Windows matters to you, a CI job on `windows-latest` is the contribution that would change this row — not a bug report saying it did not work. There is a cheaper contribution than that, though, and `eveBinary` is the worked example of it. Its signature is `eveBinary(scriptUrl, platform = process.platform)`: the default keeps every caller unchanged, and passing `"win32"` explicitly makes the Windows branch reachable from a test on the runner this repo already pays for. That is why one of the twelve is covered and eleven are not — not because the others are harder, but because they read `process.platform` directly instead of taking it as an argument. Threading the same defaulted parameter through the other eleven would convert them into tested branches on Linux, and would leave a `windows-latest` job as a nice-to-have rather than the only way to move this row. ### Everything else - **Docker.** Both compose files run Postgres in it, and the scaffolded sandbox is `eve/sandbox/docker` (`templates/default/agent/sandbox/sandbox.ts`), so every tested path has it. Pointing the agent at a Postgres you host elsewhere is not tested here. - **Postgres** is tested as `pgvector/pgvector:pg17` — the image both compose files run, and the one CI's `services:` block starts. The `vector` extension is not optional if you use memory: `templates/default/lib/memory.ts` runs `CREATE EXTENSION IF NOT EXISTS vector` and stores a `vector(n)` column. - **Browsers.** The dashboard is a Next app and has no browser test matrix. The only browser CI ever launches is Chromium, and that is for the marketing site's Playwright suite, not the dashboard. - **eve** is pinned at `^0.30.8` in `templates/default/package.json`. Below `0.30.0` is a security floor, not a preference. ## What "supported" means here evestack is a small open-source project with a contract suite, not a vendor with an SLA. What you can rely on: 1. **Reported bugs get read.** Open an issue at [github.com/SammyTourani/evestack](https://github.com/SammyTourani/evestack/issues). 2. **Security reports get a reply within 72 hours**, and a fix or mitigation before public disclosure. See [SECURITY.md](https://github.com/SammyTourani/evestack/blob/main/SECURITY.md). 3. **Assumptions about eve are checked by machine, not by memory.** `contract/` is 22 contracts that go red when eve's behaviour drifts under them, and `eve-watch` runs them against every new eve release daily. That is the closest thing here to a compatibility guarantee, and it is stronger than a promise because it is executable. What you should not rely on: a response time on a feature request, a deprecation window, or a version staying installable forever. --- # The @evestack registry > Take one piece of evestack without migrating a project or forking eve. evestack ships as an [`eve` registry](https://eve.dev/docs) in the shadcn registry-item format — the same mechanism `eve add channel/slack` uses for the official catalog. ```bash eve registry add @evestack=https://raw.githubusercontent.com/SammyTourani/evestack/main/registry/r/{name}.json eve add @evestack/memory ``` This is the reason evestack doesn't fork `eve`: anyone already running a vanilla `eve init` project can take exactly the piece they want. | Item | Adds | | --- | --- | | `@evestack/memory` | Semantic long-term memory via pgvector — `remember`/`recall`/`forget` tools | | `@evestack/instrumentation` | OTLP trace export to an evestack dashboard | | `@evestack/docker-sandbox` | Local Docker sandbox instead of hosted Vercel Sandbox | | `@evestack/basic-auth` | HTTP Basic route auth, for agents running off Vercel | | `@evestack/channel-slack` | Slack channel wiring | | `@evestack/channel-telegram` | Telegram channel wiring | | `@evestack/channel-discord` | Discord channel wiring | Seven items, which is what `registry/r/` contains. This table listed four for long enough that three shipped items were reachable only by guessing their names — derive it rather than trust it: ```bash ls registry/r/ # every item, plus registry.json (the index) ``` Verified end to end against a stock `eve init` project: `eve add @evestack/memory` installs all **four** files at their correct targets — `lib/memory.ts` plus the `remember`, `recall` and `forget` tools — the installed content is byte-identical to what ships in `templates/default` (the build script inlines from the real, tested source rather than a maintained copy), and `tsc --noEmit` passes with **no edits to `package.json` or `tsconfig.json`**. That verification was run against eve 0.29.5. `templates/default` now pins `^0.30.8`, and `@evestack/basic-auth` requires `>=0.30`, so the recorded run is older than the code it describes. The CI gate below keeps the *content* honest; nobody has re-run the end-to-end install since. That last part is the part that is easy to get wrong. A stock eve project maps only `"#*": "./agent/*"` and carries no tsconfig `paths`, and TypeScript's `bundler` module resolution ignores package.json `imports` outright — so an item whose files import each other through a `#`-subpath cannot be made to typecheck there without the user editing two files. The memory tools therefore import `lib/memory.ts` by relative path, in `templates/default` too, so the tested code and the shipped item are the same code. Dependencies are pinned to the ranges `templates/default` is tested against, generated from that manifest at build time. `@evestack/memory` carries `pg@^8.20.3`, `@ai-sdk/openai@^4.0.0` and `ai-sdk-ollama@^4.1.0` — only what its own files import, not the template's whole dependency list, and the third is there because the memory tools take either embedding provider. A bare, unversioned name would resolve to whatever npm calls latest on the day you run it, and the pin is what stops an item installing a major version the code was never exercised against. `ai` is deliberately **not** among them even though `lib/memory.ts` imports `embed` from it. Every eve project already depends on `ai` and carries an `overrides` entry pinning it, so an item that declared it directly would fail the install with `EOVERRIDE` rather than resolve. This paragraph used to read `@ai-sdk/openai@^2.0.0` and warn against ending up on v4. The repo moved to v4 deliberately (`@ai-sdk/openai@4.0.30` made `execution-denied` a first-class tool result), so the warning had inverted into a caution against the current, tested state. Read the pin out of the item rather than out of this sentence: ```bash node -e "console.log(require('./registry/r/memory.json').dependencies)" ``` ## Where the registry is hosted The URL is `raw.githubusercontent.com` and not `registry.evestack.dev` because **evestack owns no domain**. `evestack.dev` is unregistered — it returns NXDOMAIN — so the branded URL that used to be documented here died with `getaddrinfo ENOTFOUND` on the first command a stranger ran. Serving the JSON straight off `main` costs nothing and works today. If you ever buy `evestack.dev`, point it at `registry/r/` and change the URL in the five places that carry it: this file, `README.md`, `docs/channels/slack.mdx`, `docs/channels/telegram.mdx`, and the example in the header comment of `registry/build.mjs`. (`llms.txt` links the docs by raw URL for the same reason, but those are docs pages, not registry items.) Nothing in the shipped code reads the registry URL — it only ever lands in a user's own `package.json` under `registries`, written there by `eve registry add`. Two constraints on any replacement host, both read out of eve's registry client (`eve/dist/src/cli/commands/registry-project.js` and eve's bundled shadcn client): - The URL **must** contain the literal `{name}` placeholder. `eve registry add` rejects the mapping outright otherwise, before it ever touches the network. - The `Content-Type` does **not** matter. raw.githubusercontent.com serves these files as `text/plain; charset=utf-8` and eve parses the body as JSON anyway; it reads `content-type` only on the error path, to pull a message out of a non-2xx response. The one real cost of the raw URL is caching: GitHub sends `cache-control: max-age=300`, so a freshly pushed registry change can take up to five minutes to reach users. --- # CLI > The eight `evestack` commands, and the scaffolder under both of its published names. `evestack` is the whole command. Eight things: ```bash evestack create [name] scaffold an agent, a database and a dashboard evestack status is it up? what do I run? evestack tour a guided first run, on a stack that is already up evestack open the dashboard URL and its password, in a browser evestack verify check every part and name the fix for anything broken evestack skills teach your coding agent this project evestack attach [dir] add evestack to an eve project you already have evestack doctor a run stopped moving — read-only forensics ``` `npx create-evestack [name]` is the same scaffolder as `evestack create`, under the name npm's `create-*` convention leads people to. Same code, same prompts, same flags — a bug is fixed once. `-h` / `--help` and `-V` / `--version` work anywhere, and asking never writes, starts or opens anything. Each command's `--help` prints its own options; the top-level one is the command list and nothing else. A bare `evestack` inside a project runs `status`. Outside one it prints the command list and exits `0` — typing the program's name is not an error. A command it does not recognise gets one suggestion if there is a close one (`evestack verfiy` → *Did you mean `evestack verify`?*) and never a guess if there is not. ### Which one am I looking for The three checking commands answer different questions, and it is worth being able to pick: | | Question | Cost | | --- | --- | --- | | `status` | Is it running right now, and where? | three parallel probes plus a config read, read-only | | `verify` | Is it configured correctly, part by part? | talks to Docker, Postgres and the ingest route | | `doctor` | Everything is up and a run still will not move — why? | reads Postgres, writes nothing | ## `evestack create` / `npx create-evestack` ```bash npx evestack create [name] [--yes] [--verbose] npx create-evestack [name] [--yes] [--verbose] ``` | Flag | Effect | | --- | --- | | `[name]` | Project directory. Prompted for if omitted (interactively). | | `--yes` / `-y` | Skip prompts, use defaults. Also triggered automatically when stdin isn't a TTY (CI, a piped script). It additionally declines to start containers, because nobody is there to say no to a 230 MB pull. | | `--verbose` | Show the raw `npm` and `docker` output instead of one progress row each. | Four questions, all asked before any work starts, so the install and the image pull are a wait you can walk away from: 1. **Where** — the directory. 2. **Model** — OpenAI, Anthropic or Ollama, and a key if you have one. 3. **Tools** — Composio, on by default. 4. **Bring it up** — whether to start Postgres, create the schema and pull the dashboard. ### What it generates - The agent project, from `templates/default` - `.env.local` with a **uniquely generated** `EVESTACK_AUTH_PASSWORD` — never a shipped default - `.env` with the generated database password, which is the only file Compose interpolates from - `.gitignore` covering `.env*`, so a generated credential can't be accidentally committed - `docker-compose.yml` for Postgres, with the dashboard behind a profile It finishes by drawing the four moving parts with the ports **this** project chose — they are not always 2000, 5433 and 4000, because a second scaffold on the same machine moves them — and then offers to start the agent, which is the only step left. ### Why the scaffolder has zero dependencies A scaffolder that installs a prompt library before asking its first question is slower than the thing it scaffolds — and every dependency is another supply-chain surface for a tool that writes files and credentials to your disk. It's built on Node's built-in `readline`, `crypto`, and `fs` only. ### Exit codes `0` only when the project was created **and** its dependencies installed. If `npm install` fails, or `node_modules/eve` is missing afterwards, it exits `1` and prints what to run — so a `&&` chain or a CI step stops there rather than walking into an empty `node_modules`. If you accept the offer to start the agent, the exit code becomes the agent's. ### Two real bugs this caught An early version exited `0` — success — having created nothing, whenever stdin wasn't an interactive terminal (a CI run, a piped heredoc). `readline`'s `question()` never resolves after stdin hits EOF, and Node exits cleanly once the event loop empties. Silent success is the worst failure mode for a scaffolder, so `--yes` and non-TTY detection now race every prompt against stdin closing, and fall back to a sane default instead of hanging or vanishing. The second was the same failure mode wearing a different hat. The template copy skipped any source path matching `node_modules`, tested against the **absolute** path — and `npx` stages a package at `~/.npm/_npx//node_modules/create-evestack/template/…`, so every file was skipped and `npx create-evestack` produced an empty directory. It worked perfectly from a monorepo checkout, where no `node_modules` appears in the path, which is exactly why nothing caught it. The filter now matches path *segments* of the path *relative to the template root*, and CI scaffolds from an `npx`-shaped directory on every PR. ## `evestack status` ```bash evestack status [--json] ``` The glance. Four parts — the agent, Postgres and the dashboard probed in parallel, plus the model configuration, which is read from `.env.local` rather than called, because a status command that spends money is one people stop typing — with the command that fixes anything that is down printed under it. Read-only: the Postgres connection is pinned `default_transaction_read_only = on` by the same helper `doctor` uses, so it is safe to point at production. It reports two things about Postgres beyond reachability: whether the workflow schema was ever created (forgetting `db:bootstrap` is a distinct state with a distinct fix, not a green tick), and how many runs and memories are in it. When Postgres *and* the dashboard are both unreachable it checks Docker before printing two compose commands that would fail, and says that instead. | Code | Meaning | | --- | --- | | `0` | Everything this project needs is answering. | | `1` | Something is down, and it printed what to run. | | `2` | Not an evestack project. | ## `evestack tour` ```bash evestack tour [--yes] [--message=TEXT] [--no-open] ``` A guided first run, four steps, on a stack that is already up. It confirms the parts are answering, sends **one real message** to your agent and streams the reply into the terminal, then links you to that same turn in the dashboard and explains what each surface just showed you. That one message is a real model call — a fraction of a cent on `gpt-5-mini`, free and slower on Ollama. Nothing else in the tour calls a model, and it says so before it sends anything. **With no terminal to ask, it refuses and exits 3.** In CI, under a pipe, or with stdin closed, the confirmation has nobody to answer it — so the tour stops rather than treating silence as consent, and names `--yes` as the way to accept the charge up front. It used to send the message. | Flag | Effect | | --- | --- | | `--yes` / `-y` | Don't ask before sending the message. Required when stdin is not a terminal. | | `--message=TEXT` | Send something other than the default question. | | `--no-open` | Never launch a browser. | It needs the agent, Postgres and a model key. The dashboard is step 3 rather than a prerequisite — a turn is a turn whether or not anything is watching — so with the container down it still runs and prints the link. ## `evestack verify` ```bash evestack verify [--open | --no-open] [--json] ``` Runs the project's own `scripts/verify.mjs` and groups its eleven checks in dependency order — **foundation** (`.env.local`, Docker, Postgres, the schema, pgvector), **model** (the provider key, embeddings), then **the stack** (the agent, the dashboard, trace ingest). The eleventh, `dashboard image` — does the running container's version match the tag your compose file pins? — runs only once the dashboard answers, and prints under **other**, the catch-all group that exists so a check can never be silently dropped. For anything broken it names the command that fixes it, and it exits `1` if a required check failed, so CI can run it too. Grouping is the point: a red `dashboard` and a red `postgres` used to read as equally urgent when one of them is why the other failed. Fix the highest red line first. | Flag | Effect | | --- | --- | | `--open` | Open the dashboard afterwards without asking. | | `--no-open` | Never open it. Implied when it is not running in a terminal. | | `--json` | Machine-readable, and opens nothing. | It does not reimplement the checks: it executes the script that shipped with *your* project, so a globally installed CLI cannot report problems against rules your project does not have. A project made by `attach` has no `scripts/` of its own, and there it falls back to the copy inside `create-evestack`. ## `evestack open` ```bash evestack open [--no-open] ``` Prints the dashboard URL, the username and the password, says whether anything is answering there, and opens a browser. It exists because the scaffolder prints the credentials exactly once into a terminal that then scrolls, and the only recovery path was knowing which key of `.env.local` held the password. The port comes from `EVESTACK_DASHBOARD_URL`, which is where the chosen dashboard port is recorded — the scaffolder does not assume 4000. `--no-open` prints and stops. Exit `1` means nothing is answering yet. ## `evestack skills` ```bash evestack skills [--dir=PATH] [--print] [--force] [--json] ``` Writes the evestack skill pack — `SKILL.md` plus four reference files — where your coding agent will find it, so it knows this project without being told. The same pack the landing page's **Set up your agent** button copies. Full detail in [Set up with an agent](/docs/agent-setup). | Flag | Effect | | --- | --- | | `--dir=PATH` | Where to write it. Default: `agent/skills/evestack` inside an eve project, otherwise `.claude/skills/evestack`. | | `--print` | Write the pack to stdout and touch no files. | | `--force` | Overwrite files that already exist. Without it, one existing file stops the run and nothing is written. | | `--json` | Report what was written, as JSON. | The pack is fetched from the site rather than bundled into this package, so it cannot ship stale — which is the trade: it needs a network connection, and says so plainly when it does not have one. | Code | Meaning | | --- | --- | | `0` | Installed, or printed. | | `1` | Could not fetch the pack, or refused to overwrite. | | `2` | Bad arguments. | ## `evestack attach` ```bash npx evestack attach [dir] [--yes] [--dry-run] ``` Adds evestack to an eve project you already have, without overwriting anything, and prints an undo line for everything it writes. | Flag | Effect | | --- | --- | | `[dir]` | The project to attach to. Defaults to the current directory. | | `--yes` / `-y` | Skip the confirmation. Required when stdin is not a terminal — it will not guess at a project it did not create. | | `--dry-run` / `-n` | Print the plan and write nothing. | ## `evestack doctor` ```bash evestack doctor [options] ``` Read-only forensics for a durable job that is dead: it never writes to your database, and when there is something to fix it prints the SQL and lets you decide. Safe to point at production. | Flag | Effect | | --- | --- | | `--schema=NAME` | graphile-worker's schema. Default `graphile_worker`. | | `--workflow=NAME` | eve's workflow schema. Default `workflow`. | | `--url=URL` | Postgres connection string. Default `$WORKFLOW_POSTGRES_URL`, then `$DATABASE_URL`, then `WORKFLOW_POSTGRES_URL` in the project's `.env.local` or `.env`. | | `--agent-url=URL` | The eve agent, for session health. Default `$EVESTACK_AGENT_URL`, then `http://127.0.0.1:2000`. | | `--limit=N` | Max rows listed per section. Default `50`; counts are never capped. | | `--probes=N` | Max sessions probed, `0` to skip. Default `25`. | | `--idle=MINUTES` | How long a session must be quiet before it is worth probing. Default `30`. | | `--timeout=MS` | `statement_timeout` and the HTTP timeout. Default `15000`. | | `--sql` | Print only the remediation SQL, nothing else — so `--sql \| psql` stays your decision. | | `--json` | Print the whole diagnosis as JSON. Same diagnosis as the human report, not a second code path. | | `--verbose` | Add the raw rows behind each finding. | A flag that takes a value needs it as `--flag=VALUE`. `--limit 50` is refused rather than guessed at, because the alternative is a silent run with the default and a stray positional argument. ### Exit codes Doctor's codes are the ones the repro scripts in `contract/runtime/repro` use, so an operator who has run those already knows what they mean. | Code | Meaning | | --- | --- | | `0` | Looked, and found nothing that is costing a run right now. | | `1` | At least one fault — a stranded run, a wedged job, a wedged session. | | `2` | Could not look: not an evestack project, no database, wrong schema, or bad arguments. | `create` and `attach` use the same shape: `0` is a project you can run, `1` is one you cannot. `status`, `verify`, `open`, `tour` and `doctor` all work from anywhere inside the project directory — they walk up looking for the project's env files, because `evestack` is on `PATH` and gets typed from wherever you happen to be. Outside a project they all say so, rather than failing at whatever they tried next. ## Output Every command shares one renderer, so colour, width and glyphs are decided in one place (`packages/create-evestack/ui.mjs`, copied verbatim into each scaffold as `scripts/ui.mjs`). **Colour is off unless stdout is a terminal.** Piping, redirecting to a file, or paging gets plain text — which is what makes a `doctor` report safe to paste into an issue. Three environment variables override the decision: | Variable | Effect | | --- | --- | | `NO_COLOR` | Any value, including empty, turns colour off. ([no-color.org](https://no-color.org)) | | `FORCE_COLOR` | Any value but `0` turns colour on, even through a pipe — for CI logs that render it. | | `EVESTACK_ASCII` | Any value swaps the block and box-drawing glyphs for ASCII. Set automatically on Windows outside Windows Terminal, where a legacy code page mangles them. | `TERM=dumb` is treated as no colour. The brand blue needs a 256-colour terminal and falls back to cyan without one. --- # The dashboard as MCP tools > Point Claude Code at your fleet and ask it questions in English. Read-only until you say otherwise. `@evestack/mcp` speaks [MCP](https://modelcontextprotocol.io) on top of the dashboard's HTTP routes, so an agent can answer *"why did last night's run stop?"* against your own Postgres instead of you opening nine tabs. It is a **thin client**, deliberately. No database connection, no SQL, no price table, no copy of eve's protocol — every tool is a projection of a route `@evestack/dashboard` already serves. The alternative is a second implementation of the cost calculation that disagrees with the UI you are looking at within a week. The consequence is stated rather than hidden: **if the dashboard has no route for something, this package has no tool for it.** ## Setup ```jsonc // .mcp.json, or claude_desktop_config.json { "mcpServers": { "evestack": { "command": "npx", "args": ["-y", "@evestack/mcp"], "env": { "EVESTACK_MCP_DASHBOARD_URL": "http://localhost:4000" } } } } ``` | Variable | Default | What it does | | --- | --- | --- | | `EVESTACK_MCP_DASHBOARD_URL` | `http://localhost:4000` | Where the dashboard is. | | `EVESTACK_MCP_DASHBOARD_AUTH` | — | `user:password` for the dashboard. Every route is behind it. | | `EVESTACK_MCP_ALLOW_CONTROL` | unset | `1` advertises the four mutating tools. See below. | | `EVESTACK_MCP_APPROVER` | — | The name recorded in the audit log for decisions this server makes. | | `EVESTACK_MCP_TIMEOUT_MS` | — | Per-request timeout against the dashboard. | ## The tools Read-only, always available: | Tool | Answers | | --- | --- | | `list_sessions` | What has run recently. | | `get_session` | One session: live waiting state, pending approval requests, usage, and the verbatim budget-stop reason. | | `list_approvals` | Who approved or denied what, and how that identity was established. | | `get_costs` | Caps, per-principal daily spend, stops, lifetime totals. | | `promote_session_to_eval` | Generates eval source from a real session. Writes nothing. | Advertised only with `EVESTACK_MCP_ALLOW_CONTROL=1`: | Tool | Effect | | --- | --- | | `start_session` | Starts a real run. **Spends money.** | | `send_message` | Another turn on a live session. **Spends money.** | | `approve_or_deny` | **Runs the gated tool for real.** Audited. | | `cancel_run` | Cooperative stop between steps; the in-flight model call still bills. | ## Why the default is read-only `approve_or_deny` lets one model approve a tool call another model is parked on. That is the exact gate a human was asked to stand at — the whole reason eve pauses the turn — so enabling it by default would quietly delete the human from human-in-the-loop. The mutating half is therefore **withheld from `tools/list` entirely**, not merely refused on call. A model cannot plan around a capability it has never been told exists, so there is no prompt that talks the server into using one. Turning on `EVESTACK_MCP_ALLOW_CONTROL` means an agent can approve shell commands on your machine and start runs that cost money. Set `EVESTACK_MCP_APPROVER` with it, so the audit log on [Approvals](/docs/dashboard) records which agent decided rather than `unidentified`. ## What it tells you when a route is missing This package and the dashboard version independently, so a tool can outrun the deployment it is pointed at. When that happens a 404 comes back as a tool-execution error naming the missing handler, never as an empty result — because an empty approval log and an absent approval log are opposite answers to "who approved this?", and a model handed the first when the second is true will report that nobody did. If you see one, the fix is upgrading the dashboard image, not this package. ## Provenance Hand-rolled against the MCP spec: no SDK, **no dependencies at all**, `node:` builtins only — the same bet the rest of evestack makes, that a few hundred lines you can read beats a dependency you cannot. Read-only versus mutating is declared three ways, because different clients surface different ones: the first word of every description, `annotations.readOnlyHint`, and whether the tool appears in `tools/list` at all. Source and the full tool reference: [`packages/evestack-mcp`](https://github.com/SammyTourani/evestack/tree/main/packages/evestack-mcp#readme). --- # Troubleshooting > Grouped by symptom — see local-setup for the setup-time issues. ## "eve rejects my workflow world at boot" Two different mistakes produce this, and the second one is the common one now. `@workflow/world-postgres` is on `latest` (the 4.x line) instead of the `5.0.0-beta` line. The `eve` line this repo pins requires `5.0.0-beta` and rejects anything else outright. Or it is on the `beta` dist-tag — or on `^5.0.0-beta.32` / `~5.0.0-beta.32`, which admit the same releases — and npm resolved it forward. The message names spec versions: ``` [env-runner] worker init failed: This Workflow runtime requires a World with matching spec version 5, but the configured World declares spec version 6. Development worker failed before readiness ``` `world-postgres@5.0.0-beta.34` and `.35` depend on `@workflow/world@5.0.0-beta.27`/`.28`, which declare spec **6**; eve 0.30.8's Workflow runtime requires spec **5**. Pin the exact version the template pins — `"@workflow/world-postgres": "5.0.0-beta.32"` — delete `node_modules` and the lockfile entry, and reinstall. ## "The agent will not start any more, with `Invalid input: expected undefined`" The symptom is a boot that never completes, every time, with one or both of: ``` Invalid input: expected undefined, received Date path: completedAt Invalid input: expected undefined, received Uint8Array path: output ``` **One row in `workflow.workflow_runs` is in a shape eve's own read schema forbids**: a non-terminal `status` (`pending` or `running`) on a row that also carries `completed_at`, `output_cbor` or `error_cbor`. `WorkflowRunSchema` is a discriminated union whose pending/running branch declares all three as `undefined`, and `world.start()` re-enqueues active runs by listing and parsing *every* such row before it filters anything. One bad row aborts recovery, so the deployment is dead — and stays dead, because the same row is read on every start. It is also invisible to the obvious query. The dashboard defines an in-flight turn as `completed_at IS NULL`, and the poisoned row *has* a `completed_at` — so a database in this state can report **zero open runs** while refusing to boot. Find it: ```sql SELECT id, status, completed_at, output_cbor IS NOT NULL AS has_output, error_cbor IS NOT NULL AS has_error FROM workflow.workflow_runs WHERE (status IN ('pending', 'running') AND (completed_at IS NOT NULL OR output_cbor IS NOT NULL OR error_cbor IS NOT NULL)) OR (status IN ('completed', 'failed', 'cancelled') AND completed_at IS NULL); ``` The second half of that `WHERE` catches the mirror-image shape, which is equally fatal and far less obvious: a **terminal** row whose `completed_at` is NULL. Every terminal branch of the union requires `completedAt`. A NULL column reaches the schema as `undefined` rather than null, because world-postgres maps one to the other in `compact()` — and where `null` would have been quietly coerced to 1970-01-01 and booted, `undefined` builds an Invalid Date and throws. Repair it by making the status agree with the payload the row already carries — the run really did finish, only the status says otherwise: ```sql UPDATE workflow.workflow_runs SET status = (CASE WHEN error_cbor IS NOT NULL THEN 'failed' ELSE 'completed' END)::workflow.status WHERE status IN ('pending', 'running') AND (completed_at IS NOT NULL OR output_cbor IS NOT NULL OR error_cbor IS NOT NULL); ``` The `::workflow.status` cast is required, not decorative: `status` is an enum, and without it Postgres refuses the statement with `column "status" is of type workflow.status but expression is of type text`. The version first published here omitted it and did not run. For the mirror-image shape, give the row the completion time it should already have had. `updated_at` is the closest honest value — it is when the engine last wrote the row, which is when it finished: ```sql UPDATE workflow.workflow_runs SET completed_at = updated_at WHERE status IN ('completed', 'failed', 'cancelled') AND completed_at IS NULL; ``` Back the table up first, and read the rows before you write. This is your durable session state, and the statement above is a judgement — that a run holding an output or an error finished — not something the engine can confirm after the fact. **How a row gets that way.** Anything that writes `status` back over a row the engine owns, without checking the engine has not moved on in between. Our own runtime probe did exactly that and has been fixed. Upstream there is a narrower race with the same outcome: `@workflow/world-postgres` guards `run_completed`, `run_failed` and `run_cancelled` with `notInArray(status, TERMINAL_WORKFLOW_RUN_STATUSES)` and does **not** guard `run_started`, so two workers racing on one run can leave a completed row marked `running`. Rare, and worth knowing about if you see this without having written to the table yourself. ## "Under systemd/launchd it says eve is not installed, but it is" ``` eve is not installed in this project. Run npm install, then npm run build before npm run start. ``` The message is wrong and the install is fine. `npm run start` prepends `node_modules/.bin` to PATH; a service manager does not — systemd's default is `/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin`, launchd's is shorter — so `scripts/start.mjs` spawning the bare name `eve` found nothing and reported the only cause it knew about. Fixed in the template: `eveBinary()` in `scripts/checks.mjs` resolves `node_modules/.bin/eve` by absolute path, and the message now prints the path it looked for. If you see this on an older scaffold, either update `scripts/start.mjs` and `scripts/checks.mjs` from [templates/default](https://github.com/SammyTourani/evestack/tree/main/templates/default/scripts), or add the directory to the unit's `Environment=PATH=…`. ## "The agent runs as a service but every bash command fails" ``` DockerUnavailableError: The Docker sandbox backend requires Docker, but the `docker` CLI was not found. ``` The agent boots, serves HTTP and calls the model perfectly well — only the sandbox is dead, which is why this is easy to miss until someone asks the agent to run something. eve spawns the `docker` CLI **by name** (`process.env.EVE_DOCKER_PATH ?? "docker"`, no shell), so it needs it on the supervised process's PATH. Docker Desktop installs it to `/usr/local/bin` and Homebrew to `/opt/homebrew/bin`; launchd's default PATH (`/usr/bin:/bin:/usr/sbin:/sbin`) contains neither. Set PATH in the unit or plist — the shipped `deploy/dev.evestack.agent.plist` already does — or name the binary outright with `EVE_DOCKER_PATH=$(command -v docker)`. See [Operations](/docs/operations#the-two-traps-that-only-appear-under-a-supervisor). ## "npm run db:prune says there is nothing to prune" Almost always one run stuck in `running`. A session is pruned whole or not at all, and any run in it that is `pending` or `running` protects the whole family — deliberately, because that is also exactly what an open session looks like. ```sql SELECT COALESCE(attributes ->> '$rootRunId', id) AS session_id, id, status, name, updated_at FROM workflow.workflow_runs WHERE status IN ('pending', 'running') ORDER BY updated_at; ``` A genuinely open session is fine and should stay. A row that finished months ago and still says `running` is the defect above under "Invalid input: expected undefined" — same repair, same warning about backing the table up first. ## "I changed EVESTACK_MODEL and it's still calling the old provider" It always will. `EVESTACK_MODEL` names a model; `EVESTACK_PROVIDER` picks who to ask. Set only the first and the new model name goes to the provider you were already on, which fails in whatever way that provider fails on a name it doesn't recognise. The local case is the loudest — ``` Cannot compile agent compaction because the primary compaction trigger model "openai/qwen3" does not have known AI Gateway context window metadata. ``` — because eve looks the context window up in the AI Gateway catalog, and `openai/qwen3` is not in it. Anthropic model names sent to OpenAI produce a 404 from OpenAI instead. Set both variables together; the table in [Local setup](/docs/local-setup#picking-a-provider) lists which key each provider reads. A provider value that isn't `openai`, `anthropic` or `ollama` now throws at boot naming the three valid ones. Older templates fell back to OpenAI silently, so `EVESTACK_PROVIDER=claude` looked configured and behaved as though it wasn't set at all. ## "`eve dev` rebuilds whenever a file changes in my home directory" The giveaway is a doubled path in the log line, naming files you never put in the project: ``` [eve:dev] change detected (5 events: unlink /Users/me/Users/me/.npmrc, add /Users/me/.gemrc, add /Users/me/.npmrc, add /Users/me/.yarnrc, add /Users/me/.yarnrc.yml), rebuilding authored artifacts... ``` **The project has no source-root marker.** eve resolves its dev source root by walking up from the app directory until it finds `.git`, `pnpm-workspace.yaml`, or a `package.json` with a `workspaces` key — three markers, not two, and the third is easy to miss because an npm or yarn workspace root does not have to contain either of the other files. A project with none of them keeps walking, and on a machine where the dotfiles live in git, the walk stops at `$HOME` and calls it the source root. Everything follows from that: the lockfile watcher watches 20 paths, 15 of them outside the project and none of them existing, and chokidar reacts to a watch target that does not exist by watching its parent directory instead. That parent is your home directory. It is not only noise. Files matching eve's workspace-metadata list are **copied** into `.eve/dev-runtime/snapshots//source/`, and `.npmrc` is on that list — so a registry credential in `~/.npmrc` gets duplicated into the project directory, byte for byte, with nothing logged to say so. The scaffolded `.gitignore` covers `.eve/`, so it will not be committed, and nothing uploads it. It is still a credential somewhere nobody asked for it to be. **Fix, and it is one command:** ```bash git init ``` Both resolvers use the same marker list and stop at the first marker they find, so an empty repository at the project root ends the walk. Measured before and after: the source root returns to the project, watch paths drop from 20 to 5, and nothing outside the project is watched or copied. `create-evestack` does this for you — `create` runs `git init` at the end, and `attach` adds the same empty repository to a project that has no marker of its own. That is also what finally makes the `.gitignore` either of them writes mean anything. Both of them can only act at the moment they run, and **eve itself never mentions it again** — `npm run dev` does not look, and eve logs nothing when its source root leaves the project. So the fence can go silently, in every one of these: | How the marker disappears | What warns you | | --- | --- | | `git` was not installed, or `git init` failed | `create` and `attach` both say so, once, at the time | | You deleted `.git` — including by following the `rm -rf .git` line in `attach`'s own undo list | `npm run verify` | | The project reached this machine as a tarball, a zip, `git archive`, or an `rsync --exclude=.git` | `npm run verify` | | It was scaffolded before the scaffolder shipped the fence | `npm run verify` | | It is a copy inside an image built with `.git` in `.dockerignore` | `npm run verify` — though this only bites if the image's working directory sits under a marked ancestor | **`npm run verify` is the check.** It walks the same markers eve walks and reports a `fence` line with one of three outcomes: a pass when the marker is this project, a warning naming your home directory when there is no marker anywhere above, and a warning naming the directory and the file when the walk lands somewhere that holds an `.npmrc`. A marker above you that has nothing to copy is a pass, not a warning, on purpose — a line that is yellow in every workspace on every run is a line people stop reading. This is newer than most installs, so if your project's `scripts/verify.mjs` has no `fence` line it predates the check, and the one-command version is: ```bash node -e "console.log(require('node:fs').existsSync('.git') ? 'fenced' : 'no marker in this directory — see below before running git init')" ``` Read the next paragraph before acting on that output: it checks this directory only, and there is one shape of project where `git init` is the wrong response. `attach` is idempotent about this: re-running `npx evestack attach` in a project whose `.git` has gone will put the empty repository back. **The one case where `git init` is the wrong answer** is a project that really is a package inside a workspace — a `pnpm-workspace.yaml` or a `workspaces` package.json above you. eve reaches your workspace siblings by walking out of the project and into that root, so a marker here would put them outside the source root and the build would stop finding them. `attach` detects this and deliberately does not fence; it reads the workspace root's `.npmrc` and `package.json` instead and warns only if one of them holds a literal credential. If it does, move the secret to your own `~/.npmrc` or replace it with an environment reference (`_authToken=${NPM_TOKEN}`) — the copies then carry a variable name instead of a token. **In an npm or yarn workspace, check that root's `.npmrc` by hand.** `npm run verify`'s fence walk looks for `.git` and `pnpm-workspace.yaml` only — it does not know the third marker, a `package.json` carrying a `workspaces` key. So in a package whose only marker above it is an npm/yarn workspace root, verify reports *"no `.git` here or above, so eve's dev watcher will walk to your home directory"* and suggests `git init`. Both halves are wrong for that project: eve stops at the workspace root, and `git init` is the one thing you should not do there. Measured on a tree of exactly that shape — a `packages/my-agent` under a root whose `package.json` declares `workspaces` and whose `.npmrc` holds a token — verify's walk returns no marker while eve's returns the workspace root, and the `.npmrc` sitting in it is the file that gets copied. Until the walk learns the third marker, in an npm or yarn workspace do this instead of trusting the `fence` line: ```bash # the nearest package.json above you that declares workspaces, and whether it holds a credential node -e "const{existsSync,readFileSync}=require('node:fs'),{join,resolve,dirname}=require('node:path');\ let d=resolve('.');for(;;){try{if(JSON.parse(readFileSync(join(d,'package.json'),'utf8')).workspaces){\ console.log('workspace root:',d);console.log('.npmrc there?',existsSync(join(d,'.npmrc')));break}}catch{}\ const p=dirname(d);if(p===d){console.log('no workspace root above this directory');break}d=p}" ``` The underlying bug is eve's, not evestack's, and it is still present in the newest published eve. A full write-up with a minimal reproduction is drafted and ready to file at `.github/upstream/eve-dev-watcher-source-root.md`; it has not been posted. ## "The agent forgets everything between restarts" `WORKFLOW_POSTGRES_URL` isn't set, or Postgres isn't reachable at that URL. `eve` falls back silently to an on-disk world (`.eve/.workflow-data`) rather than failing loudly — silent in the sense that the agent keeps working, just without the durability you expected from Postgres. ## "Recall returns nothing even though I definitely saved that fact" If you're running a customized memory setup rather than the shipped `@evestack/memory` registry item, check the vector index type. IVFFlat built on a table that started empty can return zero rows for a query that should obviously match — see [Memory](/docs/memory) for the full explanation and the fix (HNSW). ## "A stop/cancel button in my own UI doesn't seem to work" It almost certainly did — cancellation is cooperative, not immediate. Read [Architecture § cooperative cancellation](/docs/architecture#cooperative-cancellation) for the measured ~90-second tail and the actual event ordering. ## "The dashboard shows sessions starting in the future" This was a real bug we found and fixed: `eve`'s `workflow` tables store UTC timestamps in columns with no timezone offset, so `pg` parsed them in the local timezone of whatever machine ran the dashboard. On a non-UTC machine, every run rendered shifted by that offset. Fixed in `packages/dashboard/lib/db.ts` with a type parser scoped to that one column type — if you're seeing this, you're likely on an old build; update. ## "The agent stopped answering and I get MODEL_CALL_FAILED" Read the provider's own message inside `details.message` — `eve` buries it as escaped JSON, and the dashboard's chat view digs it out for you. The most common cause on a new account is the daily request cap rather than anything in your setup: an OpenAI account with no payment method allows **50 requests per day**, and a day of building against it exhausts that faster than you would guess. The message names the limit and how long to wait. ## "I denied a tool approval and the whole session died" You are on `@ai-sdk/openai` v2. Under v2 the `output.type = "execution-denied"` that eve records for a denied call was outside the SDK's tool-output contract, so on the next turn the converter fell through and serialized it as `output: undefined`, OpenAI answered 400 (`Missing required parameter: 'input[N].output'`), and the session failed permanently. Approving never triggered it; only denial did, on both eve versions we tested (0.30.2 and 0.30.6). **Upgrade to `@ai-sdk/openai` v4, which is what the template pins.** In `@ai-sdk/provider` 4.0.5 `execution-denied` is a first-class member of `LanguageModelV4ToolResultOutput`, and `@ai-sdk/openai` 4.0.30 has an explicit case in both converters that sends `output.reason ?? "Tool call execution denied."`. The template's `surviveDeniedToolResults` middleware deliberately does **not** touch a denial any more — intercepting it on v4 would replace eve's prose reason with a JSON blob and defeat the SDK's own double-send guards. What the middleware still catches is an output type new to both eve and its list: v4's `function_call_output` switch has no default branch, so an unrecognized type still serializes to `output: undefined` and still 400s exactly as v2 did. ## "A turn shows as completed but the agent produced nothing" `eve`'s event stream and its workflow store disagree here, and the store is the misleading one. A turn killed mid-flight — a rate limit, a dropped connection — emits `turn.failed` on the stream, while the workflow row still reads `status = 'completed'`, because the workflow caught the error and finished cleanly. Nothing failed as far as it is concerned. Any dashboard trusting `status` alone therefore paints a healthy badge on a turn that did nothing. `$eve.model` is only written once a model call reports usage, so its absence on a finished turn is the surviving evidence. evestack uses exactly that signal and labels those turns **"no model call — turn produced nothing"**. ## "Composio tools aren't showing up" Check `COMPOSIO_API_KEY` is set in the **agent's** `.env.local` (not just the dashboard's — they're separate processes with separate environment files). With no key, `composioTools()` resolves to zero tools and logs one line; the agent still runs normally. ## Still stuck Open an issue with the checklist in the bug report template — a repro against a fresh `npx create-evestack test` is worth far more than a description of what you were doing in a much larger project. --- # Upstream eve > evestack is a distribution, not a fork — what eve documents, what evestack documents, and why the best eve docs are already on your disk. evestack packages [Vercel's `eve`](https://github.com/vercel/eve). It does not fork it, wrap its API, or re-document it. Every page here covers a seam that self-hosting creates — the Postgres world, auth off-platform, the dashboard, the Docker sandbox — or a subject eve has no page for at all. For the framework itself, these docs link out. That is a deliberate policy, not a gap in coverage. ## Why evestack does not mirror eve's docs **eve ships about two releases a day.** Measured from npm on 2026-08-05 (`npm view eve time --json`): **37 releases in the 20 days ending 2026-08-04**, or 1.85 per day. Seven of them landed on 2026-08-04 alone — `0.30.0` through `0.30.6`, from 14:26 to 23:01 UTC. **It is pre-1.0, so breaking changes ride minor versions.** Seven minors shipped between `0.24.0` (2026-07-14) and `0.30.0` (2026-08-04). Two of them, from eve's own `CHANGELOG.md`: | Version | Date | What broke | |---|---|---| | `0.29.0` | 2026-07-30 | `eve trace` renamed to `eve traces`; the singular form removed. `eve channels add` and the `/channels` and `/connect` TUI commands removed. | | `0.30.0` | 2026-08-04 | `localDev()` grants on process state instead of the request host; the exported `isLoopbackRequest` helper removed; the default channel auth fallback changed to `[vercelOidc(), localDev(), placeholderAuth()]`. | Those two landed five days apart. A mirrored copy of eve's docs would have been stale within a week and actively wrong about auth defaults and CLI commands within a month — which is worse than no copy, because a reader has no way to tell a stale page from a current one. eve is Apache-2.0 and its docs ship inside the npm package under that same license, so evestack could legally mirror them. It chooses not to. The constraint here is accuracy, not permission. ## Read the docs that match your pin eve's `package.json` lists `docs` in its `files` array, so **every install carries the full documentation set for exactly the version you resolved**. In a scaffolded evestack project: ```bash ls node_modules/eve/docs node -p "require('eve/package.json').version" ``` Against the version this repo pins today — `eve@0.30.8` — that is **82 Markdown pages in 12 subdirectories, 872 KB**. These figures move with the pin; the two commands above are how you re-measure them rather than trusting this paragraph. Some practical uses: ```bash # every page mentioning the auth primitive you're debugging — 43 hits in 12 files grep -rn "localDev" node_modules/eve/docs # the guide for the version you actually run, not the version eve.dev documents less node_modules/eve/docs/guides/auth-and-route-protection.md # the changelog ships too — 1,008 lines of it at 0.30.8, and it grows every release less node_modules/eve/CHANGELOG.md ``` This beats the website for any pinned project, and eve says so itself. From [eve.dev/llms.txt](https://eve.dev/llms.txt): "prefer `node_modules/eve/docs/`: those docs match the installed eve version". [eve.dev](https://eve.dev/docs/getting-started) always documents the latest release. On a day like 2026-08-04 that is up to seven versions ahead of what you installed this morning. When a website page and your local copy disagree, your local copy is the one describing the code you are running. The site does serve Markdown directly — append `.md` to any docs URL, and [eve.dev/sitemap.md](https://eve.dev/sitemap.md) lists every page — which is the right source when you are deliberately reading ahead of your pin, for example while reviewing an [upgrade](/docs/upgrading). ## Who documents what ### eve owns the framework Everything about authoring and running an agent is upstream. These links point at eve.dev because they describe eve's surface, not evestack's; read the matching file under `node_modules/eve/docs/` when the version matters. | Topic | Upstream page | |---|---| | Agent configuration — model, reasoning effort, compaction | [agent-config](https://eve.dev/docs/agent-config) | | Tools | [tools](https://eve.dev/docs/tools) | | Approvals and human-in-the-loop gating | [human-in-the-loop](https://eve.dev/docs/human-in-the-loop) | | Subagents | [subagents](https://eve.dev/docs/subagents) | | Skills | [skills](https://eve.dev/docs/skills) | | Channels — HTTP, Slack, Discord, GitHub, Linear | [channels/overview](https://eve.dev/docs/channels/overview) | | Connections — MCP and OpenAPI servers | [connections](https://eve.dev/docs/connections) | | Extensions | [extensions](https://eve.dev/docs/extensions) | | Sandbox — shell, filesystem, lifecycle, network policy | [sandbox](https://eve.dev/docs/sandbox) | | Schedules | [schedules](https://eve.dev/docs/schedules) | | Evals | [evals/overview](https://eve.dev/docs/evals/overview) | | CLI reference | [reference/cli](https://eve.dev/docs/reference/cli) | | TypeScript API | [reference/typescript-api](https://eve.dev/docs/reference/typescript-api) | | Execution model and durability | [concepts/execution-model-and-durability](https://eve.dev/docs/concepts/execution-model-and-durability) | | Security model | [concepts/security-model](https://eve.dev/docs/concepts/security-model) | | Auth and route protection | [guides/auth-and-route-protection](https://eve.dev/docs/guides/auth-and-route-protection) | | Instrumentation and OpenTelemetry export | [guides/instrumentation](https://eve.dev/docs/guides/instrumentation) | | Deploying to Vercel, and framework-level self-hosting | [guides/deployment/self-hosting](https://eve.dev/docs/guides/deployment/self-hosting) | ### evestack owns the seam | Topic | Page | |---|---| | The dashboard — sessions, turns, tokens, computed cost, driving the agent | [/docs/dashboard](/docs/dashboard) | | Running the whole stack on your own hardware | [/docs/self-hosting](/docs/self-hosting) | | Postgres as the workflow world, and the `@beta` pinning trap | [/docs/local-setup](/docs/local-setup) | | Long-term memory on pgvector | [/docs/memory](/docs/memory) | | The `@evestack` registry, and taking one piece into an existing eve project | [/docs/registry](/docs/registry) | | Keeping up with a framework that ships daily | [/docs/upgrading](/docs/upgrading) | | Failures specific to running off-platform | [/docs/troubleshooting](/docs/troubleshooting) | The dividing rule: **a page belongs here when self-hosting changes the answer.** eve's [self-hosting guide](https://eve.dev/docs/guides/deployment/self-hosting) tells you the framework supports running off Vercel; [/docs/self-hosting](/docs/self-hosting) is the runbook for the machine you actually have. eve's [instrumentation guide](https://eve.dev/docs/guides/instrumentation) tells you how to export spans; [Architecture](/docs/architecture) explains why evestack's dashboard reads Postgres instead. ## Memory: eve ships no first-party store, on purpose `https://eve.dev/docs/memory` returns 404, and no file under `node_modules/eve/docs/` has "memory" in its name except `patterns/multi-tenant-memory.md`. That page is a composition pattern — route auth, dynamic instructions, and ordinary tools — and it states plainly that the storage implementation sits outside eve, naming PostgreSQL, a durable KV store, or a vector database as equally valid. An earlier version of this page, and the comparison table on [Introduction](/docs/index), reached **"not included"** from that — by quoting the page's second paragraph and cutting its first. Restored in full, the page opens: > You can add long-term memory from the integration gallery using the Memory filter, or build > tenant-aware memory from your own application store by composing three existing eve > primitives. So memory *is* available upstream, via the gallery. Reading only the sentence that suited us is exactly the failure mode these docs exist to avoid, and it is corrected here. The accurate claim is narrower and still worth making. eve ships **no first-party memory store** — "The storage implementation is deliberately outside eve" — and the gallery's packaged options are **third-party hosted services**: [Mem0](https://eve.dev/integrations/mem0) ("persistent memory for AI agents and assistants") and [Upstash AgentKit](https://eve.dev/integrations/upstash-agentkit) ("long-term memory, Redis Search, and durable chat history"). Both are somebody else's SaaS holding your agent's recollections. That is the part a self-hosted operator has to decide, and there is nothing upstream to link for it. [/docs/memory](/docs/memory) fills it: `remember` and `recall` on the Postgres already running your sessions, indexed with HNSW rather than IVFFlat, with the measured reason that choice is a correctness fix and not a preference. Same data ownership as the rest of the stack — no third party, no extra container. ## Attribution evestack is built on [vercel/eve](https://github.com/vercel/eve) and is licensed Apache-2.0, same as eve. eve is a trademark of Vercel; evestack is an independent project and is not affiliated with or endorsed by Vercel. It is also not the first self-hosted eve distribution. [vercel-labs/steve](https://github.com/vercel-labs/steve) — "Self-hosted eve poc" — was published on 2026-06-24 by a Vercel employee, weeks before this project started. Any claim that self-hosting is a thing Vercel withheld is refuted by their own repo. The contract suite, and what to do when it goes red.