Skip to main content
This page is for the self-hoster watching what the planner costs. The orchestrator records token usage and an estimated cost for every AI-planned chat turn and every variation request. You can read it in four places: the editor’s debug panel, the browser’s network tab, the server log, and the telemetry endpoints.

Editor debug panel

  1. Open Settings in the editor’s top bar. On a public demo editor, add ?dev=1 to the URL first.
  2. Turn on Developer mode, then the Debug mode toggle it reveals. Both persist in the browser’s local storage.
  3. Send a chat message.
  4. Each assistant response shows a Debug section with traceId, outcome, intent, the model used, opCount, ops, tokens (in, out and total) and cost.
Cache counts (cacheReadInputTokens, cacheCreationInputTokens) are in the debug payload but not in the panel. Use the network tab or the telemetry API for those.

Browser network tab

Every /chat POST response includes a debug object:
For the variation endpoint (POST /chat/variations), usage is a top-level field:

Server logs

The orchestrator logs every telemetry event with its token data. Look for "event":"chat_telemetry" lines in the orchestrator’s log, which is the site’s server log in library mode:
Key phases that include token data:
  • plan_generated — after the AI returns a plan
  • result — final outcome (applied, needs_clarification, plan_ready_for_approval, etc.)

Telemetry API

The standalone server exposes stored telemetry on two endpoints. Library mode does not serve them; read the log or the NDJSON file instead. See chat telemetry events.

List events

Returns recent telemetry rows with summary stats. Each row includes inputTokens, outputTokens, totalTokens, cacheReadInputTokens, cacheCreationInputTokens, and estimatedUsd when available.

Failure review

Returns failure analysis: rates, top failed prompts, and recommendations.

Pricing table

Cost estimates use a built-in pricing table. An exact model id wins; otherwise the first key, in table order, that the model id starts with is used. For example, claude-sonnet-5 matches claude-sonnet: So claude-sonnet-5-5 resolves to its own row, claude-sonnet-5 to claude-sonnet, and the default fast model claude-haiku-4-5-20251001 to claude-haiku. If no key matches, estimatedUsd is null. To update pricing, edit USD_PER_MTOK in packages/orchestrator-core/src/telemetry/usage.ts.

Limitations

  • OpenAI streaming returns zero token counts, because the stream carries no usage data
  • Anthropic streaming captures usage from the final message after the stream completes
  • Gemini reports input and output tokens from usageMetadata, with no cache fields
  • cacheReadInputTokens maps to OpenAI cached_tokens and Anthropic cache_read_input_tokens
  • cacheCreationInputTokens is currently provided by Anthropic (cache_creation_input_tokens)
  • Cache pricing is not applied. The estimate multiplies input and output tokens by the list rates, so it ignores cache-read discounts and cache-write premiums
  • Image generation and audio transcription are not tracked
  • Cost is an estimate based on list pricing; actual billing may differ

Chat telemetry events

The phase events each turn emits, from received to result.

AI providers

Model tiers, and which provider answers what.