> ## Documentation Index
> Fetch the complete documentation index at: https://docs.avocadostudio.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# AI providers and model routing

> Run the planner on Anthropic, OpenAI, or Google Gemini with your own keys, and choose which model tier plans each edit.

Avocado Studio is **provider-agnostic**. The planner runs the same prompts and emits the same operations whether you wire it to Anthropic Claude, OpenAI GPT, or Google Gemini. You bring the API keys; no per-seat pricing, no vendor lock-in. This page is for the developer or self-hoster deciding which keys to set and which models plan the edits.

A single chat turn plans on one provider. Image generation is configured separately, so you can plan on Claude and generate images with Gemini or OpenAI.

## Supported providers

| Provider | Status | Notes |
| - | - | - |
| **Anthropic Claude** | Recommended | The planner's prompts and its eval are tuned against Claude Sonnet 5.5. Required for [extended thinking](#extended-thinking-anthropic-only). When an Anthropic key is set, the editor's model picker offers only Claude models. |
| **OpenAI** | Supported | Also powers `gpt-image-2` (final) / `gpt-image-1-mini` (draft) image generation. |
| **Google Gemini** | Supported | Often the cheapest option. Default backend for AI image generation. Required for conversational image editing. |

At least **one** provider key is required for AI editing. With none, chat still answers from a built-in planner that handles simple, literal edits, and visual editing works in full.

Each request names its provider. A request that names none — from a script or an MCP client, say — plans on the first configured key in the order OpenAI, Anthropic, Gemini. Set `CHAT_PLANNER_FORCE_SONNET=1` to send those requests to Claude Sonnet instead whenever an Anthropic key is set.

## Bring your own keys

Set whichever you have in `.env`:

```bash theme={null}
ANTHROPIC_API_KEY=sk-ant-...
OPENAI_API_KEY=sk-...
GOOGLE_GENAI_API_KEY=...
```

The editor reads `/status/planner` on boot to find out which providers have keys. Providers without keys are never offered.

Keys never leave your orchestrator. The editor and your site talk to the orchestrator over HTTP; the orchestrator is the only process that holds API credentials.

## Model tiers

Every provider has four tiers. Each tier is a separate env var, so you can change the model behind a tier without changing code:

| Tier | When it is used |
| - | - |
| `FAST` | The intent router, and plans for edits the router judges simple. Cheap models win here. |
| `BALANCED` | The default for every chat turn. The workhorse tier. |
| `REASONING` | Only when someone picks it in the model picker. Nothing switches to it automatically. |
| `CODEX` | The largest model on offer (Opus 5.5 on Anthropic), when someone picks it in the model picker. |

Agent mode does not use these tiers. It has its own two variables, `AGENT_ANTHROPIC_MODEL` (default `claude-sonnet-5-5`) and `AGENT_OPENAI_MODEL` (default `gpt-5.6-terra`).

### Defaults

```bash theme={null}
# Anthropic (defaults baked into the orchestrator)
ANTHROPIC_MODEL_FAST=claude-haiku-4-5-20251001
ANTHROPIC_MODEL_BALANCED=claude-sonnet-5-5
ANTHROPIC_MODEL_REASONING=claude-sonnet-5-5
ANTHROPIC_MODEL_CODEX=claude-opus-5-5

# OpenAI
OPENAI_MODEL_FAST=gpt-5.4-nano
OPENAI_MODEL_BALANCED=gpt-5.6-terra
OPENAI_MODEL_REASONING=gpt-5.6-terra
OPENAI_MODEL_CODEX=gpt-5.3-codex

# Gemini
GOOGLE_GENAI_MODEL_FAST=gemini-3.5-flash-lite
GOOGLE_GENAI_MODEL_BALANCED=gemini-3.6-flash
GOOGLE_GENAI_MODEL_REASONING=gemini-2.5-pro
GOOGLE_GENAI_MODEL_CODEX=gemini-2.5-pro
```

Override any of them in `.env`. The orchestrator reads these variables once, when it starts, so **restart it** after changing one.

<Note>
  The model picker shows each tier by its default model's name. It cannot see an override yet, so after you set `ANTHROPIC_MODEL_BALANCED` it still reads "Sonnet 5.5".
</Note>

## Tiered routing in action

When a user types in the editor:

```mermaid theme={null}
flowchart TD
  msg["User message<br/>'Add a testimonial section'"]
  router["Intent router<br/>(FAST tier)"]
  planner["Full planner<br/>(the tier you picked, BALANCED by default)"]
  fast["Planner on the FAST tier"]
  ops["Operations"]
  msg --> router
  router -->|simple edit| fast
  router -->|anything else| planner
  planner --> ops
  fast --> ops
```

1. **Intent router** runs on the `FAST` tier. It decides whether this is a chat-only message ("what does Hero do?"), a real edit, or ambiguous, and how complex the edit is.
2. **Full planner** runs on the tier you picked, `BALANCED` unless you changed it. The router and planner run in parallel (`CHAT_PARALLEL_PLANNER=1`, the default). When the router judges an edit simple, the planner drops to the `FAST` model for that turn.
3. **Extended thinking** switches on for complex prompts on Anthropic (`CHAT_AUTO_REASONING=1`, the default). Signals: multi-step asks, conditional language ("if there's already a CTA, ..."), structural verbs ("restructure," "rewrite tone of"), long prompts. The model stays the same; only thinking is added.

So a "fix typo in hero" request typically costs one cheap Haiku call, while "restructure the homepage to be more conversion-focused" runs on Sonnet with extended thinking. No request is moved up to a larger tier automatically.

## Extended thinking (Anthropic only)

When `CHAT_AUTO_REASONING=1` is on (default), the planner enables Anthropic's [extended thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking) for complex prompts. SSE events stream the thinking tokens back to the editor:

* `thinking_start` — model begins reasoning
* `thinking_token` — incremental text deltas
* `thinking_end` — reasoning complete

The editor renders these as a collapsible "Thinking…" block above the change log.

```bash theme={null}
CHAT_AUTO_REASONING=1                # default on
CHAT_AUTO_REASONING_BUDGET=2048      # thinking budget, mapped to an effort level (min 1024)
```

OpenAI and Gemini models do their own reasoning, but the planner does not stream it. Thinking events come from Anthropic only.

## Image generation routing

Image generation is decoupled from the planner. Two env vars control it:

```bash theme={null}
VARIATION_DEFAULT_IMAGE_SOURCE=unsplash   # or "ai" / "gemini" / "openai"
IMAGE_GEN_PROVIDER=gemini                 # or "openai" — which backend handles AI gen
                                          # (variations default to gemini; the image.generate
                                          #  tool and POST /image/generate default to openai)
```

| Setting | Result |
| - | - |
| `VARIATION_DEFAULT_IMAGE_SOURCE=unsplash` | Stock photos by default; AI generation only on explicit mention |
| `VARIATION_DEFAULT_IMAGE_SOURCE=ai` + `IMAGE_GEN_PROVIDER=gemini` | AI by default, Gemini backend |
| Explicit keyword in chat (`"generate via openai"`, `"unsplash"`) | Overrides any default |

If the configured provider has no API key, the orchestrator falls back to the other backend rather than failing.

## Switching providers from the editor

Every chat message carries an optional `provider` (`openai` | `anthropic` | `gemini`) and `modelKey` (`fast` | `balanced` | `reasoning` | `codex`). The editor sends both on every message. You choose them in **Settings** → **Model**, and the choice is remembered in your browser.

The picker offers Claude models whenever the orchestrator has an Anthropic key. OpenAI and Gemini are still wired end to end, but they are not held to the same quality bar, so the picker lists them only on a deployment that has no Anthropic key.

<Note>
  There is no server-side way to override a provider the request names — the per-message
  `provider` from the client wins. To constrain a deployment, supply only the API key for
  the provider you want. A request that names a provider without a key falls back to one
  that has one.
</Note>

## See also

* [How it works](/how-it-works) — the planner → ops → preview pipeline end-to-end
* [Chat troubleshooting](/observability/chat-troubleshooting) — debugging planner output
* [Token usage tracking](/observability/token-usage-tracking) — measuring per-provider spend
* [MCP server](/integration/mcp-server) — drive the same planner from Claude Desktop / Cursor


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.