> ## Documentation Index
> Fetch the complete documentation index at: https://docs.avocadostudio.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Docker deployment

> Deploy the orchestrator as a self-contained Docker image. Avocado Studio is Apache 2.0 licensed and self-hostable — Docker is the supported path for running it on your own infrastructure.

<Note>
  Avocado Studio is **Apache 2.0 licensed and self-hostable**: you run the whole stack on your own infrastructure, on your own Anthropic, OpenAI or Google API keys, with no per-seat licence. The orchestrator (the brain that runs sessions, calls the LLMs, and serves draft state) is the only stateful service in the stack, and Docker is the supported path for running it. If you need help self-hosting, get in touch — we would rather hear about the sharp edge than have you work around it.
</Note>

This page is for the self-hoster running the **standalone** orchestrator: the Fastify server from `apps/orchestrator`, as a Docker container on any host that supports long-lived containers with a persistent volume. The site and the editor deploy separately, to Vercel, Netlify, or any host that runs Next.js or serves static files. If your orchestrator runs inside your Next.js app (library mode), you do not need this image; see [state and backups](/operations/state-and-backups) for what that process has to keep.

## Where to host the orchestrator

The orchestrator is a small Node.js 22 Fastify service. It is stateful but not heavy.

### What the host has to give you

Any host works if it offers three things: **one** long-lived container, a persistent volume you can mount at `/app/.data`, and request timeouts you can raise. The list below is how that maps onto the usual candidates — it is a shape guide, not a compatibility matrix, and we have not run a deployment on every row.

| Host | What to set |
| - | - |
| **Self-managed VPS (Hetzner, Linode, etc.)** | The simplest option, and the one closest to how the image is meant to run. `docker compose up -d` plus a reverse proxy (Caddy, nginx, Traefik) for HTTPS. |
| **Fly.io** | The always-on machine model fits the orchestrator's long-lived sessions. Mount a Fly Volume at `/app/.data`. |
| **Railway** | Provision a volume — the default ephemeral filesystem loses every session on redeploy. |
| **DigitalOcean App Platform** | A Web Service with an attached volume. |
| **Render** | A Web Service with a Render disk mounted at `/app/.data`. Free tier is too small for real use; the smallest paid tier is fine. |
| **AWS ECS / Fargate, GCP Cloud Run, Azure Container Apps** | A single task/instance with a persistent volume (EFS, Filestore, Azure Files). Cloud Run needs `min instances = 1` so the container doesn't cold-start mid-session. |
| **Kubernetes** | A single-replica `StatefulSet` with a `PersistentVolumeClaim` at `/app/.data`. **Do not** run multiple replicas — see the constraints below. |

<Note>
  **What we run ourselves.** Our own orchestrator is on Render, but it is deployed from source with Render's native Node buildpack (`rootDir: apps/orchestrator`, `tsx src/index.ts`) rather than from this image — the [Vercel deployment guide](/operations/vercel-deployment#orchestrator-from-source) has that recipe. The image packages the same service; it is not yet the thing that has the most production hours on it. If you hit a rough edge self-hosting it, tell us — that feedback is more useful to us than a workaround is to you.
</Note>

### Resource requirements

* **Memory**: \~300–500 MB at idle, \~600–800 MB under typical load. A 1 GB instance is comfortable; the smallest "\$7/month-ish" tier on most hosts is enough for early use.
* **CPU**: Mostly I/O-bound — the orchestrator spends most of its time waiting on LLM API calls, not computing. 0.5 vCPU is fine for a handful of concurrent sessions; 1 vCPU is comfortable for a small team.
* **Disk**: A persistent volume for `/app/.data` (session state, telemetry, generated images). 1–5 GB is plenty for early use; generated-image storage is what grows fastest if you use AI image generation heavily.
* **Network**: Outbound HTTPS to your chosen LLM providers (`api.anthropic.com`, `api.openai.com`, `generativelanguage.googleapis.com`); inbound HTTPS from your editor and site origins. No inbound from end-users — only your editor and Next.js site need to reach it.

### Critical hosting constraints

A few things matter regardless of which host you pick:

<Warning>
  **Single instance only.** The orchestrator persists all session state — draft pages, undo history, chat threads, version log, site configs — to a local SQLite database via `better-sqlite3`. **Do not run multiple replicas behind a load balancer**: each replica would open its own SQLite file (or fight over a shared one), so requests round-robined between processes would see divergent state and sessions would appear to randomly lose work. Horizontal scaling is on the roadmap but not implemented today. Configure your host for exactly one instance and scale vertically (more memory / CPU on a single machine) if you need more capacity.
</Warning>

* **Persistent volume is required.** The orchestrator stores its SQLite database at `/app/.data/orchestrator.db` (plus `-wal` / `-shm` sidecar files and rolling backups). If you mount that path on an ephemeral filesystem (Cloud Run without a volume, Heroku-style ephemeral dynos, default container hosts without disk attachment), every redeploy or container reschedule will wipe all sessions and undo history. Always attach a persistent volume — even 1 GB is enough.
* **SSE-friendly reverse proxy.** The chat endpoint streams server-sent events for live editor updates. Some reverse proxies and CDNs buffer responses by default, which makes the editor look frozen until the full response lands. If you put a reverse proxy in front of the orchestrator, disable response buffering on `/chat/*` — and on `/sites-agent/*` if you turn that surface on (in nginx: `proxy_buffering off`; in Caddy: `flush_interval -1` on `reverse_proxy`; in Cloudflare: bypass cache for these paths).
* **Long timeouts, if you enable the agent surface.** Chat turns finish in seconds, so a default 30s or 60s request timeout is survivable for editing alone. Sites-agent runs (full URL migration, repo integration) take several minutes and will be killed mid-run — so if you set `AGENT_SURFACE=on` (see [The agent surface](#the-agent-surface-off-by-default) below), raise request timeouts to at least 10 minutes on the orchestrator's service.
* **CORS for the editor's origin.** Set `ORCHESTRATOR_CORS_ORIGINS` to include both your site and editor origins (HTTPS, no trailing slash). See [CORS configuration](#cors-configuration) below.
* **HTTPS reachable from your editor and site, and from nobody else.** Both the editor (in the browser) and your site (server-side draft fetches) call the orchestrator, so it needs a URL both can reach, such as `https://orchestrator.example.com`. The standalone server enforces the access password only on the agent surface: `/chat`, `/ops`, `/history/*` and `/telemetry/*` answer anyone who can reach them. Restrict who can reach it at the network level. See [security and access](/reference/security).

## Building the image

The orchestrator source lives at `apps/orchestrator` in the repo. Build the image from the repository root:

```bash theme={null}
docker build -f apps/orchestrator/Dockerfile -t avocado-orchestrator:latest .
```

The build is multi-stage and uses the pnpm workspace to install only the orchestrator's dependencies (`packages/shared`, `packages/migration-sdk`, `packages/orchestrator-core`, `packages/richtext`). It needs BuildKit, which is the default in Docker 23 and later — the Dockerfile's `# syntax` directive requires it, and so does the context ignore list, which lives at `apps/orchestrator/Dockerfile.dockerignore` because the build context is the repo root.

The container runs the TypeScript source with `tsx`; there is no separate compile step. The runtime stage is `node:22-slim` rather than Alpine on purpose: `better-sqlite3` ships prebuilt binaries against glibc, and on musl the install falls back to compiling from source with a toolchain the slim Node images don't carry into production. The build stage installs `python3`, `make` and `g++` as a fallback for targets with no prebuild; they never reach the final image.

## Running standalone

```bash theme={null}
docker run -d \
  --name avocado-orchestrator \
  -p 4200:4200 \
  --env-file .env \
  -v avocado-data:/app/.data \
  avocado-orchestrator:latest
```

### Required environment variables

At minimum you need **one** AI provider key:

| Variable | Description |
| - | - |
| `ANTHROPIC_API_KEY` | Claude API key (recommended — most battle-tested) |
| `OPENAI_API_KEY` | OpenAI API key |
| `GOOGLE_GENAI_API_KEY` | Google Gemini API key |

For an editor hosted on a public URL, also set these on the orchestrator. The
editor fetches the first two from `GET /editor/credentials` after sign-in, and
that route answers 403 in production until a credential is configured:

| Variable | Description |
| - | - |
| `ACCESS_PASSWORD_HASH` | SHA-256 hex digest of the editor password |
| `DRAFT_MODE_SECRET` | The site's draft secret |
| `PUBLISH_TOKEN` | The site's publish secret |

### CORS configuration

<Note>
  With `ORCHESTRATOR_CORS_ORIGINS` unset, the orchestrator falls back to four development origins: `http://localhost:3000`, `http://127.0.0.1:3000`, `http://localhost:4100` and `http://127.0.0.1:4100`. `localhost` and `127.0.0.1` are distinct origins to a browser, which is why both spellings are listed. For production, set the allowed origins explicitly — the fallback is replaced, not extended.
</Note>

```bash theme={null}
ORCHESTRATOR_CORS_ORIGINS=https://your-site.example.com,https://editor.example.com
```

### State persistence

The orchestrator writes session state, telemetry, and generated images to `/app/.data` inside the container. Mount a volume there to persist data across restarts.

The image pre-configures these paths:

* `ORCHESTRATOR_DB_FILE=/app/.data/orchestrator.db` — the live state, plus its `-wal` / `-shm` sidecars and rolling `.db.backup-<ts>` snapshots
* `ORCHESTRATOR_STATE_FILE=/app/.data/orchestrator-state.json` — legacy, and only meaningful if you are carrying over a volume written before the SQLite store. Nothing creates this file today
* `CHAT_TELEMETRY_FILE=/app/.data/chat-telemetry.ndjson`
* `ORCHESTRATOR_GENERATED_IMAGE_DIR=/app/.data/generated-images`

The server snapshots the database into the same volume every 24 hours, keeping
14, and the first snapshot is taken about a minute after start. Copy snapshots
off the volume if the drafts matter to you. See
[state and backups](/operations/state-and-backups).

On restart, drafts, history and the version log come back from the volume.
Held approval plans and editor sign-in tokens do not: people are asked for the
password again.

## The agent surface, off by default

The container sets `NODE_ENV=production`, and in production the orchestrator **does not mount** `/agent/*` or `/sites-agent/*` — the routes behind site onboarding, URL migration and repo integration. Calls to them return 404 until you opt in.

That is deliberate rather than an oversight. Those routes run open-ended multi-turn agent loops with file and shell tools; the `useCliAgent` variant spawns the Claude CLI with `--permission-mode bypassPermissions` and passes the request body through as the prompt. That is arbitrary code execution as the orchestrator's process user, by design — it is what makes onboarding work — so it cannot be made safe by narrowing what it may run. It can only be kept off and put behind a credential.

To turn it on you need both of these:

```bash theme={null}
AGENT_SURFACE=on
# and one credential, so callers have something to present:
ACCESS_PASSWORD_HASH=<sha256 hex>      # editor exchanges the password at /auth/verify
# or
ORCHESTRATOR_ACCESS_TOKEN=<static token>
```

`AGENT_SURFACE=on` with neither credential set **refuses to mount and says why in the boot log** rather than mounting an open surface. `AGENT_CLI=1` is a separate opt-in, off everywhere by default, for the variant that spawns the CLI on the host.

If you only need chat editing and publishing, leave all of this alone — `/chat`, `/ops` and `/publish` are unaffected. The same goes for `SITE_OPS_AGENTS`, which turns on pre-alpha features and is off by default.

## Using docker-compose

A `docker-compose.yml` at the repo root runs the orchestrator with sensible defaults:

```bash theme={null}
cp .env.example .env    # required — see below
docker compose up -d
docker compose logs -f orchestrator
docker compose down
```

The compose file uses a named volume (`orchestrator-data`), sets `ORCHESTRATOR_CORS_ORIGINS` to `http://localhost:3000,http://localhost:4100` for local development, and loads env vars from `.env` at the repo root. That `env_file` entry is **not optional**: with no `.env` present, `docker compose up` fails before it starts anything. Copy `.env.example` and put at least one provider key in it first.

## Health check

The container includes a health check that polls `http://127.0.0.1:4200/health` every 30 seconds. Check status with:

```bash theme={null}
docker inspect --format='{{.State.Health.Status}}' avocado-orchestrator
```

## Environment reference

See `.env.example` at the repo root for the complete list of environment variables. Common Docker overrides:

| Variable | Purpose |
| - | - |
| `PORT` | HTTP port (default: 4200) |
| `HOST` | Bind address (default: `::`, all interfaces). Container platforms sometimes set `HOST` for you |
| `NODE_ENV` | `production` by default in the image |
| `ORCHESTRATOR_CORS_ORIGINS` | Comma-separated list of allowed origins |
| `ORCHESTRATOR_DB_FILE` | The SQLite database — the live state. Set to `/app/.data/orchestrator.db` in the image; the literal `:memory:` forces an ephemeral store |
| `ORCHESTRATOR_DB_BACKUP_INTERVAL_HOURS` | How often a `VACUUM INTO` snapshot is taken next to the database (default: 24) |
| `ORCHESTRATOR_DB_BACKUP_LIMIT` | How many rolling `.db.backup-<ts>` snapshots to keep (default: 14) |
| `ORCHESTRATOR_STATE_FILE` | Legacy session-state JSON, read **once** at first boot and migrated into SQLite, then renamed `<path>.migrated-<ts>` and never written again |
| `ORCHESTRATOR_JSON_MIGRATION_TTL_DAYS` | How long that archived JSON is kept before it is swept (default: 14) |
| `CHAT_TELEMETRY_FILE` | Path to telemetry NDJSON |
| `ORCHESTRATOR_GENERATED_IMAGE_DIR` | Directory for generated images |
| `IMAGE_GEN_PROVIDER` | Backend for AI image generation — `gemini` (default) or `openai`. Falls back if the chosen provider has no key |
| `PUBLISH_TOKEN` | Three jobs. When set, `POST /publish` requires it in the `x-publish-token` header; the `site-contract` target sends it as that header to your site; and `GET /editor/credentials` hands it to a signed-in editor. Optional here, **not** optional on the site: a site running under `NODE_ENV=production` with no matching publish secret answers 401 and every publish fails |
| `DRAFT_MODE_SECRET` | The site's draft secret, handed to a signed-in editor by `GET /editor/credentials` |
| `ACCESS_PASSWORD_HASH` | SHA-256 hex digest of a password. `/auth/verify` exchanges the password for a bearer token the editor sends as `x-access-token`. In the standalone server only the agent surface and `GET /editor/credentials` enforce it |
| `ORCHESTRATOR_ACCESS_TOKEN` | A static bearer token, as an alternative to the password hash. Either one satisfies the agent surface's credential requirement |
| `AGENT_SURFACE` | `on` mounts `/agent/*` and `/sites-agent/*`, which are otherwise absent in production. See [The agent surface](#the-agent-surface-off-by-default) |
| `AGENT_CLI` | `1` additionally allows the variant that spawns the Claude CLI on the host. Off everywhere by default |

## Running locally without Docker

For local development you can run the orchestrator directly via pnpm from the repository root — that's the faster dev loop. Docker is the supported path for production self-hosting; the source-based workflow is for anyone iterating on the orchestrator itself.

```bash theme={null}
pnpm install
pnpm dev:start        # starts site + editor + orchestrator via tsx
# or run just the orchestrator:
pnpm --filter @ai-site-editor/orchestrator dev
```

The Dockerfile is an **additional** distribution option, not a replacement for the source-based dev workflow.

## Troubleshooting

### Container exits immediately

Check logs: `docker logs avocado-orchestrator`. Look for an invalid `.env` file, or a `better-sqlite3` load error when the image was built for a different platform.

### Editor shows the password prompt again after a deploy

Expected. Sign-in tokens are held in memory, so a restart voids them. `ORCHESTRATOR_ACCESS_TOKEN` is the credential that survives restarts, for scripts.

### CORS errors from editor or site

Set `ORCHESTRATOR_CORS_ORIGINS` to include both the site and editor origins (no trailing slashes). Remember that setting it *replaces* the localhost defaults, so a value that lists only your production site will lock out a locally-running editor.

### `/sites-agent/*` returns 404

Expected in a container: the agent surface is unmounted under `NODE_ENV=production` unless you opt in. See [The agent surface](#the-agent-surface-off-by-default). The boot log says which decision was made and why — look for `[agent-surface]`.

### State not persisting

Ensure the volume is mounted at `/app/.data`. The image runs as root, so permissions are rarely the problem — check that the mount actually landed (`docker inspect -f '{{json .Mounts}}' avocado-orchestrator`) and that `ORCHESTRATOR_DB_FILE` still points inside it if you overrode it.

### Health check failing

Wait for the 10-second start period. If it still fails, check the logs for `Orchestrator listening on 4200`. The server binds `::` (all interfaces) by default; a `HOST` value set by the platform can override that. Also check that no firewall is blocking the port.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.