Skip to content
▚ evestack docs

Troubleshooting

Grouped by symptom — see local-setup for the setup-time issues.

"eve rejects my workflow world at boot"

Two different mistakes produce this, and the second one is the common one now.

@workflow/world-postgres is on latest (the 4.x line) instead of the 5.0.0-beta line. The eve line this repo pins requires 5.0.0-beta and rejects anything else outright.

Or it is on the beta dist-tag — or on ^5.0.0-beta.32 / ~5.0.0-beta.32, which admit the same releases — and npm resolved it forward. The message names spec versions:

[env-runner] worker init failed: This Workflow runtime requires a World with matching spec
version 5, but the configured World declares spec version 6.
Development worker failed before readiness

world-postgres@5.0.0-beta.34 and .35 depend on @workflow/world@5.0.0-beta.27/.28, which declare spec 6; eve 0.30.8's Workflow runtime requires spec 5. Pin the exact version the template pins — "@workflow/world-postgres": "5.0.0-beta.32" — delete node_modules and the lockfile entry, and reinstall.

"The agent will not start any more, with Invalid input: expected undefined"

The symptom is a boot that never completes, every time, with one or both of:

Invalid input: expected undefined, received Date          path: completedAt
Invalid input: expected undefined, received Uint8Array    path: output

One row in workflow.workflow_runs is in a shape eve's own read schema forbids: a non-terminal status (pending or running) on a row that also carries completed_at, output_cbor or error_cbor. WorkflowRunSchema is a discriminated union whose pending/running branch declares all three as undefined, and world.start() re-enqueues active runs by listing and parsing every such row before it filters anything. One bad row aborts recovery, so the deployment is dead — and stays dead, because the same row is read on every start.

It is also invisible to the obvious query. The dashboard defines an in-flight turn as completed_at IS NULL, and the poisoned row has a completed_at — so a database in this state can report zero open runs while refusing to boot.

Find it:

SELECT id, status, completed_at, output_cbor IS NOT NULL AS has_output,
       error_cbor IS NOT NULL AS has_error
  FROM workflow.workflow_runs
 WHERE (status IN ('pending', 'running')
          AND (completed_at IS NOT NULL OR output_cbor IS NOT NULL OR error_cbor IS NOT NULL))
    OR (status IN ('completed', 'failed', 'cancelled') AND completed_at IS NULL);

The second half of that WHERE catches the mirror-image shape, which is equally fatal and far less obvious: a terminal row whose completed_at is NULL. Every terminal branch of the union requires completedAt. A NULL column reaches the schema as undefined rather than null, because world-postgres maps one to the other in compact() — and where null would have been quietly coerced to 1970-01-01 and booted, undefined builds an Invalid Date and throws.

Repair it by making the status agree with the payload the row already carries — the run really did finish, only the status says otherwise:

UPDATE workflow.workflow_runs
   SET status = (CASE WHEN error_cbor IS NOT NULL THEN 'failed' ELSE 'completed' END)::workflow.status
 WHERE status IN ('pending', 'running')
   AND (completed_at IS NOT NULL OR output_cbor IS NOT NULL OR error_cbor IS NOT NULL);

The ::workflow.status cast is required, not decorative: status is an enum, and without it Postgres refuses the statement with column "status" is of type workflow.status but expression is of type text. The version first published here omitted it and did not run.

For the mirror-image shape, give the row the completion time it should already have had. updated_at is the closest honest value — it is when the engine last wrote the row, which is when it finished:

UPDATE workflow.workflow_runs
   SET completed_at = updated_at
 WHERE status IN ('completed', 'failed', 'cancelled')
   AND completed_at IS NULL;

Back the table up first, and read the rows before you write. This is your durable session state, and the statement above is a judgement — that a run holding an output or an error finished — not something the engine can confirm after the fact.

How a row gets that way. Anything that writes status back over a row the engine owns, without checking the engine has not moved on in between. Our own runtime probe did exactly that and has been fixed. Upstream there is a narrower race with the same outcome: @workflow/world-postgres guards run_completed, run_failed and run_cancelled with notInArray(status, TERMINAL_WORKFLOW_RUN_STATUSES) and does not guard run_started, so two workers racing on one run can leave a completed row marked running. Rare, and worth knowing about if you see this without having written to the table yourself.

"Under systemd/launchd it says eve is not installed, but it is"

  eve is not installed in this project.

  Run npm install, then npm run build before npm run start.

The message is wrong and the install is fine. npm run start prepends node_modules/.bin to PATH; a service manager does not — systemd's default is /usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin, launchd's is shorter — so scripts/start.mjs spawning the bare name eve found nothing and reported the only cause it knew about.

Fixed in the template: eveBinary() in scripts/checks.mjs resolves node_modules/.bin/eve by absolute path, and the message now prints the path it looked for. If you see this on an older scaffold, either update scripts/start.mjs and scripts/checks.mjs from templates/default, or add the directory to the unit's Environment=PATH=….

"The agent runs as a service but every bash command fails"

DockerUnavailableError: The Docker sandbox backend requires Docker,
but the `docker` CLI was not found.

The agent boots, serves HTTP and calls the model perfectly well — only the sandbox is dead, which is why this is easy to miss until someone asks the agent to run something.

eve spawns the docker CLI by name (process.env.EVE_DOCKER_PATH ?? "docker", no shell), so it needs it on the supervised process's PATH. Docker Desktop installs it to /usr/local/bin and Homebrew to /opt/homebrew/bin; launchd's default PATH (/usr/bin:/bin:/usr/sbin:/sbin) contains neither. Set PATH in the unit or plist — the shipped deploy/dev.evestack.agent.plist already does — or name the binary outright with EVE_DOCKER_PATH=$(command -v docker). See Operations.

"npm run db:prune says there is nothing to prune"

Almost always one run stuck in running. A session is pruned whole or not at all, and any run in it that is pending or running protects the whole family — deliberately, because that is also exactly what an open session looks like.

SELECT COALESCE(attributes ->> '$rootRunId', id) AS session_id,
       id, status, name, updated_at
  FROM workflow.workflow_runs
 WHERE status IN ('pending', 'running')
 ORDER BY updated_at;

A genuinely open session is fine and should stay. A row that finished months ago and still says running is the defect above under "Invalid input: expected undefined" — same repair, same warning about backing the table up first.

"I changed EVESTACK_MODEL and it's still calling the old provider"

It always will. EVESTACK_MODEL names a model; EVESTACK_PROVIDER picks who to ask. Set only the first and the new model name goes to the provider you were already on, which fails in whatever way that provider fails on a name it doesn't recognise. The local case is the loudest —

Cannot compile agent compaction because the primary compaction trigger model
"openai/qwen3" does not have known AI Gateway context window metadata.

— because eve looks the context window up in the AI Gateway catalog, and openai/qwen3 is not in it. Anthropic model names sent to OpenAI produce a 404 from OpenAI instead. Set both variables together; the table in Local setup lists which key each provider reads.

A provider value that isn't openai, anthropic or ollama now throws at boot naming the three valid ones. Older templates fell back to OpenAI silently, so EVESTACK_PROVIDER=claude looked configured and behaved as though it wasn't set at all.

"eve dev rebuilds whenever a file changes in my home directory"

The giveaway is a doubled path in the log line, naming files you never put in the project:

[eve:dev] change detected (5 events: unlink /Users/me/Users/me/.npmrc, add /Users/me/.gemrc,
add /Users/me/.npmrc, add /Users/me/.yarnrc, add /Users/me/.yarnrc.yml), rebuilding authored artifacts...

The project has no source-root marker. eve resolves its dev source root by walking up from the app directory until it finds .git, pnpm-workspace.yaml, or a package.json with a workspaces key — three markers, not two, and the third is easy to miss because an npm or yarn workspace root does not have to contain either of the other files. A project with none of them keeps walking, and on a machine where the dotfiles live in git, the walk stops at $HOME and calls it the source root. Everything follows from that: the lockfile watcher watches 20 paths, 15 of them outside the project and none of them existing, and chokidar reacts to a watch target that does not exist by watching its parent directory instead. That parent is your home directory.

It is not only noise. Files matching eve's workspace-metadata list are copied into .eve/dev-runtime/snapshots/<id>/source/, and .npmrc is on that list — so a registry credential in ~/.npmrc gets duplicated into the project directory, byte for byte, with nothing logged to say so. The scaffolded .gitignore covers .eve/, so it will not be committed, and nothing uploads it. It is still a credential somewhere nobody asked for it to be.

Fix, and it is one command:

git init

Both resolvers use the same marker list and stop at the first marker they find, so an empty repository at the project root ends the walk. Measured before and after: the source root returns to the project, watch paths drop from 20 to 5, and nothing outside the project is watched or copied.

create-evestack does this for you — create runs git init at the end, and attach adds the same empty repository to a project that has no marker of its own. That is also what finally makes the .gitignore either of them writes mean anything.

Both of them can only act at the moment they run, and eve itself never mentions it again — npm run dev does not look, and eve logs nothing when its source root leaves the project. So the fence can go silently, in every one of these:

How the marker disappearsWhat warns you
git was not installed, or git init failedcreate and attach both say so, once, at the time
You deleted .git — including by following the rm -rf .git line in attach's own undo listnpm run verify
The project reached this machine as a tarball, a zip, git archive, or an rsync --exclude=.gitnpm run verify
It was scaffolded before the scaffolder shipped the fencenpm run verify
It is a copy inside an image built with .git in .dockerignorenpm run verify — though this only bites if the image's working directory sits under a marked ancestor

npm run verify is the check. It walks the same markers eve walks and reports a fence line with one of three outcomes: a pass when the marker is this project, a warning naming your home directory when there is no marker anywhere above, and a warning naming the directory and the file when the walk lands somewhere that holds an .npmrc. A marker above you that has nothing to copy is a pass, not a warning, on purpose — a line that is yellow in every workspace on every run is a line people stop reading.

This is newer than most installs, so if your project's scripts/verify.mjs has no fence line it predates the check, and the one-command version is:

node -e "console.log(require('node:fs').existsSync('.git') ? 'fenced' : 'no marker in this directory — see below before running git init')"

Read the next paragraph before acting on that output: it checks this directory only, and there is one shape of project where git init is the wrong response.

attach is idempotent about this: re-running npx evestack attach in a project whose .git has gone will put the empty repository back.

The one case where git init is the wrong answer is a project that really is a package inside a workspace — a pnpm-workspace.yaml or a workspaces package.json above you. eve reaches your workspace siblings by walking out of the project and into that root, so a marker here would put them outside the source root and the build would stop finding them. attach detects this and deliberately does not fence; it reads the workspace root's .npmrc and package.json instead and warns only if one of them holds a literal credential. If it does, move the secret to your own ~/.npmrc or replace it with an environment reference (_authToken=${NPM_TOKEN}) — the copies then carry a variable name instead of a token.

In an npm or yarn workspace, check that root's .npmrc by hand. npm run verify's fence walk looks for .git and pnpm-workspace.yaml only — it does not know the third marker, a package.json carrying a workspaces key. So in a package whose only marker above it is an npm/yarn workspace root, verify reports "no .git here or above, so eve's dev watcher will walk to your home directory" and suggests git init. Both halves are wrong for that project: eve stops at the workspace root, and git init is the one thing you should not do there.

Measured on a tree of exactly that shape — a packages/my-agent under a root whose package.json declares workspaces and whose .npmrc holds a token — verify's walk returns no marker while eve's returns the workspace root, and the .npmrc sitting in it is the file that gets copied. Until the walk learns the third marker, in an npm or yarn workspace do this instead of trusting the fence line:

# the nearest package.json above you that declares workspaces, and whether it holds a credential
node -e "const{existsSync,readFileSync}=require('node:fs'),{join,resolve,dirname}=require('node:path');\
let d=resolve('.');for(;;){try{if(JSON.parse(readFileSync(join(d,'package.json'),'utf8')).workspaces){\
console.log('workspace root:',d);console.log('.npmrc there?',existsSync(join(d,'.npmrc')));break}}catch{}\
const p=dirname(d);if(p===d){console.log('no workspace root above this directory');break}d=p}"

The underlying bug is eve's, not evestack's, and it is still present in the newest published eve. A full write-up with a minimal reproduction is drafted and ready to file at .github/upstream/eve-dev-watcher-source-root.md; it has not been posted.

"The agent forgets everything between restarts"

WORKFLOW_POSTGRES_URL isn't set, or Postgres isn't reachable at that URL. eve falls back silently to an on-disk world (.eve/.workflow-data) rather than failing loudly — silent in the sense that the agent keeps working, just without the durability you expected from Postgres.

"Recall returns nothing even though I definitely saved that fact"

If you're running a customized memory setup rather than the shipped @evestack/memory registry item, check the vector index type. IVFFlat built on a table that started empty can return zero rows for a query that should obviously match — see Memory for the full explanation and the fix (HNSW).

"A stop/cancel button in my own UI doesn't seem to work"

It almost certainly did — cancellation is cooperative, not immediate. Read Architecture § cooperative cancellation for the measured ~90-second tail and the actual event ordering.

"The dashboard shows sessions starting in the future"

This was a real bug we found and fixed: eve's workflow tables store UTC timestamps in columns with no timezone offset, so pg parsed them in the local timezone of whatever machine ran the dashboard. On a non-UTC machine, every run rendered shifted by that offset. Fixed in packages/dashboard/lib/db.ts with a type parser scoped to that one column type — if you're seeing this, you're likely on an old build; update.

"The agent stopped answering and I get MODEL_CALL_FAILED"

Read the provider's own message inside details.message — eve buries it as escaped JSON, and the dashboard's chat view digs it out for you. The most common cause on a new account is the daily request cap rather than anything in your setup: an OpenAI account with no payment method allows 50 requests per day, and a day of building against it exhausts that faster than you would guess. The message names the limit and how long to wait.

"I denied a tool approval and the whole session died"

You are on @ai-sdk/openai v2. Under v2 the output.type = "execution-denied" that eve records for a denied call was outside the SDK's tool-output contract, so on the next turn the converter fell through and serialized it as output: undefined, OpenAI answered 400 (Missing required parameter: 'input[N].output'), and the session failed permanently. Approving never triggered it; only denial did, on both eve versions we tested (0.30.2 and 0.30.6).

Upgrade to @ai-sdk/openai v4, which is what the template pins. In @ai-sdk/provider 4.0.5 execution-denied is a first-class member of LanguageModelV4ToolResultOutput, and @ai-sdk/openai 4.0.30 has an explicit case in both converters that sends output.reason ?? "Tool call execution denied.". The template's surviveDeniedToolResults middleware deliberately does not touch a denial any more — intercepting it on v4 would replace eve's prose reason with a JSON blob and defeat the SDK's own double-send guards. What the middleware still catches is an output type new to both eve and its list: v4's function_call_output switch has no default branch, so an unrecognized type still serializes to output: undefined and still 400s exactly as v2 did.

"A turn shows as completed but the agent produced nothing"

eve's event stream and its workflow store disagree here, and the store is the misleading one. A turn killed mid-flight — a rate limit, a dropped connection — emits turn.failed on the stream, while the workflow row still reads status = 'completed', because the workflow caught the error and finished cleanly. Nothing failed as far as it is concerned.

Any dashboard trusting status alone therefore paints a healthy badge on a turn that did nothing. $eve.model is only written once a model call reports usage, so its absence on a finished turn is the surviving evidence. evestack uses exactly that signal and labels those turns "no model call — turn produced nothing".

"Composio tools aren't showing up"

Check COMPOSIO_API_KEY is set in the agent's .env.local (not just the dashboard's — they're separate processes with separate environment files). With no key, composioTools() resolves to zero tools and logs one line; the agent still runs normally.

Still stuck

Open an issue with the checklist in the bug report template — a repro against a fresh npx create-evestack test is worth far more than a description of what you were doing in a much larger project.