> ## Documentation Index
> Fetch the complete documentation index at: https://docs.avocadostudio.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Asset Manager & AI Images

> Image sources, AI generation with OpenAI and Gemini (nano-banana), multi-turn image chat, and how the editor and chat pipeline resolve images.

<div className="relative w-full rounded-xl overflow-hidden border border-gray-200 dark:border-gray-800 my-6" style={{ aspectRatio: "1736 / 1080" }}>
  <video className="absolute inset-0 h-full w-full" width="1736" height="1080" src="https://mintcdn.com/avocadostudioai/_ZLLYIZBMfUDPbdL/images/asset-picker-ai.mp4?fit=max&auto=format&n=_ZLLYIZBMfUDPbdL&q=85&s=45f9cae5a614a5b3a33f07606a592361" poster="/images/asset-picker-ai-thumb.jpg" preload="none" controls playsInline data-path="images/asset-picker-ai.mp4" />

  <button
    type="button"
    aria-label="Play the Asset Manager demo"
    className="group absolute inset-0 flex h-full w-full cursor-pointer items-center justify-center border-0 bg-transparent p-0"
    onClick={(e) => {
  const overlay = e.currentTarget;
  const video = overlay.parentElement.querySelector("video");
  overlay.style.display = "none";
  Promise.resolve(video.play()).catch(() => { overlay.style.display = ""; });
}}
  >
    <span className="flex h-16 w-16 items-center justify-center rounded-full bg-black/55 backdrop-blur-sm transition-transform duration-150 group-hover:scale-110">
      <svg viewBox="0 0 24 24" fill="white" className="ml-1 h-7 w-7" aria-hidden="true">
        <path d="M8 5v14l11-7z" />
      </svg>
    </span>
  </button>
</div>

## Overview

**For the editor:** the **Asset Manager** opens whenever you change an image —
**Change** or **Choose image** on an image field in the Properties panel, the
**Change image** button on a picture in the preview (with the element picker
on), or an image row in the page's **SEO** settings. Its tabs offer AI
generation (**Generate**), your Google **Drive** folder, **Unsplash** photos,
your CMS's media library, and **Upload** from your computer or a pasted URL.
Which tabs appear depends on how your site is set up. In chat you can also just
ask: *"use a warmer photo for the hero"*.

The rest of this page is for developers: where each source comes from, which
keys switch it on, and the routes behind it.

Avocado Studio has two entry points for images: the **Asset Manager modal** in
the editor, and the **`image.generate` tool** that the chat planner can call
during natural-language edits. Both paths share the same providers and the
same on-disk storage.

```mermaid theme={null}
flowchart TD
    editor["Editor image field click"]
    chat["Chat message:<br/>'generate a hero image…'"]

    editor --> modal["Asset Manager modal"]
    chat --> tool["image.generate tool"]

    modal --> generate["Generate tab"]
    modal --> drive["Drive tab"]
    modal --> unsplash["Unsplash tab"]
    modal --> cms["CMS libraries"]
    modal --> upload["Upload tab"]

    generate --> openai["OpenAI<br/>gpt-image-1"]
    generate --> gemini["Google Gemini<br/>gemini-3.1-flash-lite-image"]
    tool --> openai
    tool --> gemini

    openai --> storage[("Generated-images store")]
    gemini --> storage
    upload --> storage
```

The modal is defined in `apps/editor/src/components/ImagePickerModal.tsx`.
The multi-turn chat view in `apps/editor/src/components/ImageGenerateChat.tsx`
is an assistant-style chat that talks to `POST /image/generate/chat` when
Gemini is configured. The `image.generate` tool manifest lives in
`packages/orchestrator-core/src/tools/builtins/image-generate.ts`.

## Image sources (modal tabs)

Which tabs appear depends on orchestrator env vars and the CMS media config
the site passes to the editor. Tab selection is computed in
`ImagePickerModal.tsx` (`availableTabs`).

| Tab | Enabled when | Backing source |
| - | - | - |
| **Generate** | Always shown. Generates when `OPENAI_API_KEY` or `GOOGLE_GENAI_API_KEY` is set; otherwise it says *Image generation is off* and names the keys | AI image generation (OpenAI or Gemini) |
| **Drive** | `GOOGLE_DRIVE_FOLDER_ID`, `GOOGLE_SERVICE_ACCOUNT_KEY_JSON` or `GOOGLE_API_KEY` set — **standalone orchestrator only**; library mode reports it off | Google Drive folder listing; a site's own folder ID in Site settings narrows it |
| **Unsplash** | `UNSPLASH_ACCESS_KEY` set | Unsplash search API |
| **CMS** | The site's adapter implements `getMedia`, or the site has a `cmsMedia` config | CMS asset library |
| **Upload** | Always available | Local file → `POST /image/upload` |

<Tip>
  **Fastest way to configure the env-var sources** (Generate / Drive / Unsplash) is the first-time setup script. Running `pnpm dev:setup` walks you through each prompt, opens the provider's signup/console page in your browser, and validates the key before writing it to `.env`. The CMS tab has no env vars: it comes from the site's adapter, or from the per-site config in the editor's Site Config drawer.
</Tip>

The feature flags the editor reads (`imageGenerate`, `imageGenerateChat`,
`googleDrive`, `unsplash`, `cmsMedia`) are published by the orchestrator at
`GET /status/planner`. All but `cmsMedia` are set purely from env vars (and
library mode always reports `googleDrive: false`); `cmsMedia` is derived from
whether the mount's adapter implements `getMedia`. There is no per-session
toggle today.

### CMS asset libraries

One tab, filled two ways. Both answer `POST /media/cms`; neither reaches a
CMS from the browser.

**The adapter path (library mode).** Implement `getMedia` on your `CmsAdapter`
and the tab appears — no editor configuration, no credentials leaving the
process that already holds them:

```ts theme={null}
import { cmsMediaSource, type CmsAdapter } from "@avocadostudio-ai/site-sdk/server"

export function sanityAdapter(): CmsAdapter {
  return {
    id: "sanity",
    getPages: (options) => getSanityPages(options?.perspective),
    getMedia: cmsMediaSource({
      provider: "sanity",
      projectId: process.env.NEXT_PUBLIC_SANITY_PROJECT_ID!,
      dataset: process.env.NEXT_PUBLIC_SANITY_DATASET
    })
  }
}
```

`cmsMediaSource` is a convenience for Contentful, Sanity and Strapi. It is not
the seam — `getMedia` is. A site on a CMS none of those three describes writes
the method itself:

```ts theme={null}
getMedia: async ({ query, page, limit }) => {
  const res = await myCms.assets.search({ q: query, page, perPage: limit })
  return {
    items: res.assets.map((a) => ({ id: a.id, name: a.filename, imageUrl: a.url, thumbUrl: a.thumb, alt: a.altText })),
    totalPages: res.pageCount,
    label: "My CMS"
  }
}
```

`GET /status/planner` reports `features.cmsMedia`, and `/whoami` reports
`capabilities.readsMedia`; both are derived from the presence of the method,
never declared. An adapter cannot claim a library it did not implement.

**The per-site path (standalone).** The standalone multi-site orchestrator
wires no adapter, so the editor sends the connection details it holds for the
active site. `CmsMediaConfig` (`apps/editor/src/lib/editor-types.ts`) is a
discriminated union over the same three providers, stored on the editor's
site entry:

| Provider | Fields |
| - | - |
| Contentful | `{ provider: "contentful", spaceId, deliveryToken, environment? }` |
| Sanity | `{ provider: "sanity", projectId, dataset?, token? }` |
| Strapi | `{ provider: "strapi", url, token? }` |

Whichever is configured, the request goes to the orchestrator and the vendor
call is made there. That is why the route is a POST: the body carries a token,
and a token in a query string is a token in every access log between the
orchestrator and the browser.

<Note>
  The adapter wins when both are available. A site that implemented `getMedia`
  meant it, and its own reader can see things a generic one cannot.
</Note>

**Click path to configure a site's `cmsMedia`:**

1. Open the editor at `http://localhost:4100`.
2. Click the **site name** in the top bar to open its menu, then choose
   **Site settings**, or open the **Sites** page and press the gear button on a
   site card.
3. In the **Site Config** drawer, open the **Deploy** tab and scroll to
   **CMS Media**.
4. Pick a **Provider** (Contentful / Sanity / Strapi / none) and fill in
   the provider-specific fields:
   * **Contentful** — Space ID and Delivery token (environment defaults
     to `master`).
   * **Sanity** — Project ID and optional dataset (defaults to
     `production`). No token needed for public datasets.
   * **Strapi** — Base URL (e.g. `https://cms.example.com`) and optional
     API token. Public upload endpoints work without a token.
5. Close the drawer. The CMS tab appears in the asset picker on next open.

<Warning>
  This config lives in the editor's own site list, not in the orchestrator's
  registry — it is per-browser, and it is where the tokens you type are kept.
  A library-mode site should use the adapter path instead, where the
  credentials never leave the server.
</Warning>

### Documents, and adding one

A CMS library holds more than pictures. `getMedia` takes a `kind` — `"image"`
(the default, and what every caller meant before documents existed) or
`"file"` — and the link field's document picker asks for the latter. For
Sanity that is a second asset type, `sanity.fileAsset`, which `cmsMediaSource`
queries when `kind` is `"file"`; a PDF uploaded in the Studio is one of those,
and a query for `sanity.imageAsset` finds none of them.

The write half is `CmsAdapter.uploadMedia`, and `cmsMediaUploader` implements
it for Sanity from the same connection details:

```ts theme={null}
import { cmsMediaSource, cmsMediaUploader, type CmsAdapter } from "@avocadostudio-ai/site-sdk/server"

const media = {
  provider: "sanity",
  projectId: process.env.NEXT_PUBLIC_SANITY_PROJECT_ID!,
  dataset: process.env.NEXT_PUBLIC_SANITY_DATASET,
  token: process.env.SANITY_API_TOKEN
} as const

const uploadMedia = cmsMediaUploader(media)

export function sanityAdapter(): CmsAdapter {
  return {
    id: "sanity",
    getPages: (options) => getSanityPages(options?.perspective),
    getMedia: cmsMediaSource(media),
    ...(uploadMedia ? { uploadMedia } : {})
  }
}
```

The file goes to Sanity's asset API and comes back with a
`cdn.sanity.io/files/…` URL — durable, CDN-served, and unaffected by any
redeploy of the site that links to it. That is the whole reason to prefer this
over writing into the site's own `public/` directory: a file written to the
running host's disk does not survive an ephemeral deployment, and a link to it
is a 404 waiting for the next deploy.

<Note>
  `cmsMediaUploader` returns **`null`** when the provider has no uploader
  (Contentful's is a three-step asynchronous create/process/publish; Strapi's is
  not written yet) or when no token is configured — hence the spread. The editor
  derives `capabilities.writesMedia` from whether the method *exists*, so an
  absent one hides the upload control, and a present one that always fails
  would be a button that is always there and never works.

  A **read-only token is the trap**: Sanity's query API answers one, so the
  picker fills with images and only the upload fails. The uploader says so in
  the refusal it hands back.
</Note>

### Unsplash: licensing & attribution

<Note>
  **Why Unsplash is in the picker.** The Unsplash tab ships primarily as a
  **demo / placeholder convenience** — it lets stakeholders, evaluators, and
  new sites populate a site with real-looking imagery during a live session
  without leaving the editor or wiring up a DAM. Production sites should
  usually graduate to branded photography, generated imagery, or a CMS-backed
  asset library. If you *do* ship Unsplash photos to end users, the rules
  below apply — attribution and API compliance are the adopter's
  responsibility, not the orchestrator's.
</Note>

Photos returned by the Unsplash tab (and by the `unsplash.search` tool) are
served under the [Unsplash License](https://unsplash.com/license). The
license is permissive — photos are free for commercial and non-commercial
use and no permission from the photographer is required — but the Unsplash
API Terms impose a few concrete obligations that adopters are responsible
for meeting:

* **Credit the photographer and Unsplash** wherever a selected photo is
  rendered. Note that the `unsplash.search` **tool** does not carry the
  photographer: its `author` is the literal `"Unsplash"` and its `sourceUrl` is
  the CDN image URL. To build real attribution, use the HTTP route
  `GET /unsplash/search`, which returns the photographer's name — for example:
  `Photo by <a href="{sourceUrl}">Photographer Name</a> on <a href="https://unsplash.com">Unsplash</a>`
* **Do not resell, redistribute, or host** unmodified Unsplash photos as a
  stock-photo service, wallpaper pack, or competing search product.
* **Do not imply endorsement** by photographers or by Unsplash of your
  product, brand, or customers.
* **Track downloads** when building your own Unsplash-powered integration.
  The API requires a `GET /photos/:id/download` trigger per selected photo;
  the `unsplash.search` tool in this repo does not do this automatically —
  adopters integrating Unsplash into a custom flow outside of the built-in
  tab should implement it to stay within Unsplash's API guidelines.

Generated images (OpenAI, Gemini) and assets returned from CMS tabs
(Contentful, Sanity, Strapi) have their own licensing and usage rules;
consult the relevant provider terms before shipping content publicly.

## AI providers

### OpenAI (default)

* Models: `OPENAI_IMAGE_MODEL` (default `gpt-image-2`) for `quality: "final"`,
  `OPENAI_IMAGE_MODEL_DRAFT` (default `gpt-image-1-mini`) for
  `quality: "draft"`.
* Sizes map from aspect ratio: `landscape → 1536x1024`, `square → 1024x1024`,
  `portrait → 1024x1536`.
* Native transparency via the `background` parameter (`transparent`,
  `opaque`, `auto`).
* Output formats: `png`, `webp`, `jpeg`.

### Google Gemini — "nano-banana"

Google's Gemini image models are publicly nicknamed **"nano-banana"**. Avocado
calls them via the `@google/genai` SDK, an optional peer dependency in library
mode: install it yourself, or the Gemini paths cannot load.

* Model: `GOOGLE_GENAI_IMAGE_MODEL`, defaulting to `gemini-3.1-flash-lite-image`.
* Aspect ratios are mapped to Gemini's `imageConfig.aspectRatio` strings:
  `landscape → 3:2`, `square → 1:1`, `portrait → 2:3`. `16:9` and `9:16` can
  be triggered from prompt hints (e.g. "wide 16:9").
* Quality tiers map to Gemini's `imageSize`: `draft → 1K`, `final → 2K`.
* **No native transparency.** When `background: "transparent"` is requested,
  the orchestrator appends a prompt hint telling the model to render on a
  fully transparent background — results vary.
* Supports multi-turn chat and reference images (see next section).

### Provider selection

The `image.generate` tool (the one the AI planner calls) and
`POST /image/generate` pick their backend from `IMAGE_GEN_PROVIDER` — `openai`
(default) or `gemini` — and fall back to the other provider when the chosen one
has no key. The editor's Generate tab routes to `POST /image/generate/chat`
when `imageGenerateChat` is enabled (i.e. `GOOGLE_GENAI_API_KEY` is set), and
otherwise makes single-shot calls to `POST /image/generate`, which use the
draft-quality model.

## Multi-turn image chat (Gemini only)

When `GOOGLE_GENAI_API_KEY` is set, the Generate tab becomes a full assistant
chat built on `@assistant-ui/react`, streaming against
`POST /image/generate/chat`. The route lives in
`apps/orchestrator/src/routes/media.ts` and in library mode's
`createOrchestrator`; both call the same
`orchestrator-core/src/http/image-generate-actions.ts`, so the chat sessions,
aspect-ratio rules and SSE frames are one implementation, not two.

### Request body

```jsonc theme={null}
{
  "prompt": "...",                 // required
  "chatId": "…",                   // optional — resume an existing session
  "aspectRatio": "3:2",            // optional — "landscape" | "square" | "portrait" | explicit ratio
  "stream": true,                  // optional — SSE vs JSON response
  "referenceImageUrl": "https://…",// optional — primary reference
  "referenceImageUrls": ["…"]      // optional — up to 14 additional references, ≤5 MB each
}
```

### SSE event stream (`stream: true`)

```
event: chatId   → { chatId, aspectRatio }
event: status   → { stage: "Generating image…" }
event: text     → { text }        // streamed text tokens from Gemini
event: image    → { url, alt }    // generated image ready
event: error    → { error }
event: done     → {}
```

If `stream` is omitted, the endpoint returns a single JSON payload with
`{ chatId, url, alt, text, aspectRatio }`.

### UI modes

`ImageGenerateChat.tsx` switches between three modes depending on whether an
image already exists in the field:

* **`choose`** — current image exists; user picks "Edit this image" or
  "Generate a new one".
* **`edit`** — re-generate using the current image as reference context, so
  Gemini can honor existing composition, palette, or subjects.
* **`new`** — fresh generation from the prompt alone.

The modal also measures the current image's natural dimensions and snaps to
the closest Gemini-supported aspect ratio (1:1, 3:2, 2:3, 16:9, 9:16) so the
new image matches the slot it will fill.

### Session management

* Sessions are kept in an in-memory map on the orchestrator, capped at **200
  concurrent** sessions.
* Idle sessions are evicted after **30 minutes** (LRU).
* **Changing aspect ratio mid-session creates a new session.** Gemini's
  `imageConfig` is immutable once a chat is created, so the orchestrator
  detects the change, deletes the old session, and starts a new one.
* Sessions do not survive an orchestrator restart.

### Reference images

Up to **14 reference images** can be attached to the first message in a
session. Each is fetched server-side, validated against a **5 MB** size cap,
and forwarded to Gemini as base64 `inlineData` parts alongside the prompt.
Failed fetches are logged and skipped — the request does not fail if some
references are unreachable.

## `image.generate` tool (chat pipeline path)

When the user asks the chat to produce or replace an image, the planner can
call the `image.generate` tool. The full manifest is in
`packages/orchestrator-core/src/tools/builtins/image-generate.ts`.

```jsonc theme={null}
{
  "name": "image.generate",
  "capability": "read",
  "timeoutMs": 90000,
  "retryPolicy": { "maxAttempts": 1 },
  "idempotent": false,
  "inputSchema": {
    "prompt":      "string (required)",
    "aspectRatio": "landscape | square | portrait",
    "quality":     "draft | final",
    "style":       "string (optional style guidance)",
    "background":  "transparent | opaque | auto",
    "outputFormat":"png | webp | jpeg",
    "blockType":   "Hero | Card | …",
    "blockId":     "string",
    "pageSlug":    "string"
  },
  "outputSchema": {
    "imageUrl": "string",
    "alt":      "string",
    "width":    "number",
    "height":   "number"
  }
}
```

### Prompt enrichment

When `blockType`, `blockId`, or `pageSlug` are provided, the tool enriches
the raw prompt with the block's composition hint, the page title, and the
block's existing `heading` / `subheading` / `title` props. It also appends
the default constraints: *no text overlays, no logos, no watermarks*. Known
composition hints include `Hero`, `Banner`, `CTA`, `Card`, `CardGrid`,
`FeatureGrid`, `Gallery`, `Carousel`, and `TwoColumn`.

### Progress streaming

The handler emits five progress stages via `context.onImageProgress`, and
those events surface on the chat SSE stream as `image_progress` events (see
[How It Works › Step 5](/how-it-works#step-5-the-live-preview-updates)):

```
 0%  Understanding prompt…
15%  Composing scene…
40%  Rendering image…
75%  Finalizing details…
95%  Almost there…
100% Done
```

### Deferred image resolution

With `CHAT_DEFER_IMAGE_RESOLUTION=1` (the default), the chat pipeline applies
text and structural ops **immediately** and resolves image tool calls in the
background. The editor receives the text updates at once and the image URLs
patch in via follow-up SSE events once generation completes.

## Other endpoints

Registered in `apps/orchestrator/src/routes/media.ts` and answered identically by
library mode, from the shared `orchestrator-core/src/http/*-actions.ts` modules — except
the Drive routes, which are registered in `apps/orchestrator/src/routes/gdrive.ts` and are
**not** served by `createOrchestrator()`.

| Endpoint | Purpose |
| - | - |
| `POST /image/generate` | Single-shot generation proxy. Uses `IMAGE_GEN_PROVIDER` (OpenAI fallback). Returns `{ url, alt }`. |
| `POST /image/generate/chat` | Multi-turn Gemini chat (documented above). Requires `GOOGLE_GENAI_API_KEY`. |
| `POST /image/upload` | Multipart form upload. Writes to `ORCHESTRATOR_GENERATED_IMAGE_DIR` and returns `{ url, bytes, mimeType }`. |
| `GET /unsplash/search?q=&page=&limit=` | Unsplash proxy used by the modal's Unsplash tab and by the `unsplash.search` tool. |
| `GET /gdrive/images` / `GET /gdrive/images/:fileId` | Google Drive browse/download with server-side resize, WebP conversion, and EXIF strip. **Standalone orchestrator only.** |

## Environment variables

| Variable | Default | Purpose |
| - | - | - |
| `IMAGE_GEN_PROVIDER` | `openai` | `openai` or `gemini` — which backend `image.generate` uses. |
| `OPENAI_API_KEY` | — | Enables OpenAI generation and the `imageGenerate` feature flag. |
| `OPENAI_IMAGE_MODEL` | `gpt-image-2` | Final-quality OpenAI image model. |
| `OPENAI_IMAGE_MODEL_DRAFT` | `gpt-image-1-mini` | Draft-quality OpenAI image model. |
| `GOOGLE_GENAI_API_KEY` | — | Enables Gemini generation and the `imageGenerateChat` feature flag. |
| `GOOGLE_GENAI_IMAGE_MODEL` | `gemini-3.1-flash-lite-image` | Gemini image model. |
| `UNSPLASH_ACCESS_KEY` | — | Enables the Unsplash tab and the `unsplash.search` tool. |
| `GOOGLE_DRIVE_FOLDER_ID` | — | Enables the Drive tab (alternatively `GOOGLE_SERVICE_ACCOUNT_KEY_JSON` or `GOOGLE_API_KEY`). |
| `ORCHESTRATOR_GENERATED_IMAGE_DIR` | `.data/generated-images` | Filesystem target for uploaded and generated images. |
| `CHAT_DEFER_IMAGE_RESOLUTION` | `1` | Apply text/structural ops first and resolve images in the background. |

## Limits & constraints

* **Reference images**: ≤14 per message, ≤5 MB each. Must be reachable over
  HTTP(S) from the orchestrator, or uploaded first via `POST /image/upload`.
* **Gemini sessions**: in-memory only. Capped at 200 concurrent with 30-minute
  LRU eviction. An orchestrator restart wipes all sessions.
* **Gemini transparency**: no native support — `background: "transparent"`
  falls back to a prompt hint and the result is best-effort.
* **`image.generate` timeout**: 90 seconds, no automatic retry.
* **Image storage**: generated images are written to
  `ORCHESTRATOR_GENERATED_IMAGE_DIR`. For production, mount a persistent
  volume or upload to an external bucket as part of your
  [publish target](/how-it-works#the-publishtarget-interface).

## Related pages

* [How It Works › Step 5](/how-it-works#step-5-the-live-preview-updates) —
  where `image_progress` events fit in the chat SSE stream.
* [Tools MVP](/integration/tools-mvp) — the broader tool contract that
  `image.generate` and `unsplash.search` implement.
* [Custom Blocks › Image fields](/integration/custom-blocks) — how to declare
  image fields on your own blocks so the Asset Manager opens for them.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.