> ## Documentation Index
> Fetch the complete documentation index at: https://docs.avocadostudio.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Token usage tracking

> How the orchestrator tracks LLM token usage and estimated cost across debug panel, logs, and telemetry.

This page is for the self-hoster watching what the planner costs. The
orchestrator records token usage and an estimated cost for every AI-planned chat
turn and every variation request. You can read it in four places: the editor's
debug panel, the browser's network tab, the server log, and the telemetry
endpoints.

## Editor debug panel

1. Open **Settings** in the editor's top bar. On a public demo editor, add
   `?dev=1` to the URL first.
2. Turn on **Developer mode**, then the **Debug mode** toggle it reveals. Both
   persist in the browser's local storage.
3. Send a chat message.
4. Each assistant response shows a **Debug** section with `traceId`, `outcome`,
   `intent`, the model used, `opCount`, `ops`, `tokens` (in, out and total) and
   `cost`.

Cache counts (`cacheReadInputTokens`, `cacheCreationInputTokens`) are in the
debug payload but not in the panel. Use the network tab or the telemetry API for
those.

## Browser network tab

Every `/chat` POST response includes a `debug` object:

```jsonc theme={null}
{
  "status": "applied",
  "summary": "Updated hero heading.",
  "debug": {
    "traceId": "abc-123",
    "promptHash": "d408ff2e63278b72",
    "promptExcerpt": "change hero heading",
    "outcome": "applied",
    "intent": "edit_plan",
    "opCount": 1,
    "opTypes": ["update_props"],
    // Token usage fields:
    "inputTokens": 1842,
    "outputTokens": 312,
    "totalTokens": 2154,
    "cacheReadInputTokens": 640,
    "cacheCreationInputTokens": 1200,
    "estimatedUsd": 0.00773
  }
}
```

For the **variation endpoint** (`POST /chat/variations`), usage is a top-level field:

```jsonc theme={null}
{
  "status": "ok",
  "summary": "Generated 3 variations for Hero.",
  "variations": [ ... ],
  "usage": {
    "inputTokens": 920,
    "outputTokens": 1450,
    "totalTokens": 2370,
    "cacheReadInputTokens": 320,
    "cacheCreationInputTokens": 610,
    "estimatedUsd": 0.0168
  }
}
```

## Server logs

The orchestrator logs every telemetry event with its token data. Look for
`"event":"chat_telemetry"` lines in the orchestrator's log, which is the site's
server log in library mode:

```jsonc theme={null}
{
  "event": "chat_telemetry",
  "phase": "result",
  "outcome": "applied",
  "modelUsed": "claude-sonnet-5-5",
  "inputTokens": 1842,
  "outputTokens": 312,
  "totalTokens": 2154,
  "cacheReadInputTokens": 640,
  "cacheCreationInputTokens": 1200,
  "estimatedUsd": 0.00773
}
```

Key phases that include token data:

* `plan_generated` — after the AI returns a plan
* `result` — final outcome (`applied`, `needs_clarification`, `plan_ready_for_approval`, etc.)

## Telemetry API

The standalone server exposes stored telemetry on two endpoints. Library mode
does not serve them; read the log or the NDJSON file instead. See
[chat telemetry events](/observability/chat-telemetry-events).

### List events

```
GET http://localhost:4200/telemetry/chat?session=<id>&limit=50
```

Returns recent telemetry rows with summary stats. Each row includes `inputTokens`, `outputTokens`, `totalTokens`, `cacheReadInputTokens`, `cacheCreationInputTokens`, and `estimatedUsd` when available.

### Failure review

```
GET http://localhost:4200/telemetry/chat/review?session=<id>&limit=300
```

Returns failure analysis: rates, top failed prompts, and recommendations.

## Pricing table

Cost estimates use a built-in pricing table. An exact model id wins; otherwise the
first key, in table order, that the model id starts with is used. For example,
`claude-sonnet-5` matches `claude-sonnet`:

| Model key | Input (\$/1M tokens) | Output (\$/1M tokens) |
| - | -: | -: |
| `gpt-5.6-sol` | 5.00 | 30.00 |
| `gpt-5.6-terra` | 2.50 | 15.00 |
| `gpt-5.6-luna` | 1.00 | 6.00 |
| `gpt-5.3-codex` | 1.75 | 14.00 |
| `gpt-5.4-nano` | 0.20 | 1.25 |
| `gpt-4o-mini` | 0.15 | 0.60 |
| `gpt-4o` | 2.50 | 10.00 |
| `gpt-5` | 1.25 | 10.00 |
| `gemini-3.6-flash` | 1.50 | 7.50 |
| `gemini-3.5-flash-lite` | 0.30 | 2.50 |
| `gemini-2.5-pro` | 1.25 | 10.00 |
| `gemini-2.5-flash` | 0.30 | 2.50 |
| `claude-sonnet-5-5` | 2.00 | 10.00 |
| `claude-opus-5-5` | 4.00 | 20.00 |
| `claude-haiku` | 1.00 | 5.00 |
| `claude-sonnet` | 3.00 | 15.00 |
| `claude-opus` | 5.00 | 25.00 |

So `claude-sonnet-5-5` resolves to its own row, `claude-sonnet-5` to
`claude-sonnet`, and the default fast model `claude-haiku-4-5-20251001` to
`claude-haiku`. If no key matches, `estimatedUsd` is `null`.

To update pricing, edit `USD_PER_MTOK` in `packages/orchestrator-core/src/telemetry/usage.ts`.

## Limitations

* **OpenAI streaming** returns zero token counts, because the stream carries no usage data
* **Anthropic streaming** captures usage from the final message after the stream completes
* **Gemini** reports input and output tokens from `usageMetadata`, with no cache fields
* `cacheReadInputTokens` maps to OpenAI `cached_tokens` and Anthropic `cache_read_input_tokens`
* `cacheCreationInputTokens` is currently provided by Anthropic (`cache_creation_input_tokens`)
* **Cache pricing is not applied.** The estimate multiplies input and output tokens by the list rates, so it ignores cache-read discounts and cache-write premiums
* **Image generation and audio transcription** are not tracked
* Cost is an estimate based on list pricing; actual billing may differ

## What to read next

<CardGroup cols={2}>
  <Card title="Chat telemetry events" icon="list-check" href="/observability/chat-telemetry-events">
    The phase events each turn emits, from `received` to `result`.
  </Card>

  <Card title="AI providers" icon="microchip" href="/ai-providers">
    Model tiers, and which provider answers what.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.