> ## Documentation Index
> Fetch the complete documentation index at: https://docs.avocadostudio.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Chat telemetry events

> Reference for all telemetry events emitted by the orchestrator chat pipeline.

This page is for the self-hoster and the developer reading the orchestrator's
chat telemetry: every event the chat pipeline records, its fields, and the
endpoints that serve it.

## Scope

This covers events pushed via `ctx.chatTelemetry.push(...)` in:

* `packages/orchestrator-core/src/chat/chat-pipeline.ts`

And serialized/logged by:

* `packages/orchestrator-core/src/telemetry/chat-telemetry.ts`

It does **not** cover generic operational logs such as `chat_pipeline_start` or image rewrite logs.

## Event transport

Each telemetry entry is:

1. Buffered in memory.
2. Optionally persisted as NDJSON.
3. Emitted to server logs as a structured log with `event: "chat_telemetry"`.

Persistence is on by default, except under `NODE_ENV=test`. Set
`CHAT_TELEMETRY_PERSIST=0` to keep telemetry in memory only.

* File: `CHAT_TELEMETRY_FILE`, else `.data/chat-telemetry.ndjson` in the data
  directory. The Docker image sets `/app/.data/chat-telemetry.ndjson`.
* Buffer: the newest `CHAT_TELEMETRY_LIMIT` entries, default `500`, which are
  also what is reloaded from the file at startup.

Rows carry a prompt excerpt, the first 180 characters of the message, and a
16-character SHA-256 prefix of the full prompt.

## Event schema

Fields from `ChatTelemetryEntry`:

* Required: `id`, `at`, `phase`, `session`, `requestedSlug`, `effectiveSlug`, `plannerSource`, `modelKey`, `modelUsed`, `promptHash`, `promptExcerpt`, `promptLength`
* Optional classification: `outcome`, `reason`, `reasonCategory`
* Optional plan shape: `intent`, `opCount`, `opTypes`
* Optional usage/cost: `inputTokens`, `outputTokens`, `totalTokens`, `cacheReadInputTokens`, `cacheCreationInputTokens`, `estimatedUsd`
* Optional timing: `totalDurationMs`, `planningDurationMs`, `firstPlanningTokenMs`, `applyDurationMs`, `imageResolutionDurationMs`, `planningAttempts`, and `timelineStage` on `milestone` rows (`request_received`, `first_token`, `first_structured_progress`, `plan_ready`, `first_op_applied`, `done`)
* Optional apply detail: `skippedOpCount`, `imagesRequested`, `imagesResolved`
* Optional planner context: `plannerTier` (`forced_deterministic`, `deterministic`, `llm_intent_router`, `full_llm`, `demo`), `contractMode` (`minimal`, `targeted`, `full`), `contractBytes`, `contractBlockCount`, `contextPackBytes`, `strictJsonEnabled`, `schemaRetryUsed`, `plannerRefusal`, `plannerIncomplete`, `compactContextEnabled`, `minimalContextEnabled`
* Optional change-log drift: `changelogMissingCount`, `changelogExtraCount`, `changelogFieldMislabelCount`
* Optional tool calls: `toolName`, `toolOk`, `toolLatencyMs`, `toolAttempts`, `toolErrorCode`, `correlationId`
* Optional suggestion pills: `suggestionIds`, `suggestionSources`, `clickedSuggestionId`

`reasonCategory` values are from guardrail classification:

* `schema_violation`
* `ambiguity`
* `not_found`
* `no_effective_change`
* `planner_refusal`
* `incomplete_output`
* `malformed_output`
* `internal_error`
* `canceled`
* `operation_failed`
* `unsupported_by_site` — the operation is valid, but the site declared it cannot honour it

## Phases

`phase` is one of:

1. `received`
2. `milestone`
3. `forced_plan`
4. `deterministic_plan_generated`
5. `plan_attempt_failed`
6. `plan_generated`
7. `plan_apply_failed`
8. `repair_attempt`
9. `repair_generated`
10. `tool_call`
11. `result`

## Outcomes

### Emitted telemetry outcomes (`chatTelemetry.push`)

1. `guardrail_failure`
2. `needs_clarification`
3. `plan_ready_for_approval`
4. `no_effective_change`
5. `applied`
6. `apply_failed`
7. `apply_pending_plan_error`
8. `forced_duplicate_page`
9. `forced_create_page`
10. `planner_exception`
11. `deterministic_plan_ready`
12. `compound_deterministic_plan_ready`
13. `attempt_${attempt}_failed` (dynamic, e.g. `attempt_1_failed`)
14. `planning_exhausted`
15. `planning_missing`
16. `planning_refusal`
17. `planning_incomplete`
18. `repair_started`
19. `repair_plan_generated`
20. `repair_failed`
21. `content_answer`
22. `variation_request_redirect`
23. `blocked_structural_capability`
24. `llm_router_plan_ready`
25. `llm_router_needs_clarification`
26. `tool_ok`
27. `tool_error`
28. `api_error`
29. `empty_edit_plan`

### Debug-only outcomes (response payload, not telemetry rows)

1. `validation_error`
2. `pending_plan_missing`
3. `pending_plan_mismatch`
4. `info`
5. `blocked_demo_mode`

## Phase to outcome mapping

Common mappings in current implementation:

* `received`: no `outcome`
* `forced_plan`: `forced_duplicate_page`, `forced_create_page`
* `deterministic_plan_generated`: `deterministic_plan_ready`
* `plan_attempt_failed`: `attempt_${attempt}_failed`
* `milestone`: no `outcome`; `timelineStage` says which point the turn reached
* `plan_generated`: usually no `outcome` (plan metadata + optional usage)
* `tool_call`: `tool_ok` or `tool_error`
* `plan_apply_failed`: `apply_failed`
* `repair_attempt`: `repair_started`
* `repair_generated`: `repair_plan_generated`
* `result`: terminal or branch outcomes such as `applied`, `needs_clarification`, `planning_exhausted`, `repair_failed`, etc.

## Telemetry APIs

The list and review endpoints are served by the **standalone server**. Library
mode serves only the feedback endpoint below. On the standalone server these
routes are not behind the access gate, and rows include prompt excerpts, so do
not expose the server publicly. A `PUBLIC_DEMO` answers 403 on them.

List entries (`limit` default 100, max 1000):

* `GET /telemetry/chat?limit=<n>&outcome=<outcome>&phase=<phase>&session=<session>`

Review aggregate (`limit` default 300, max 2000):

* `GET /telemetry/chat/review?limit=<n>&session=<session>`

`/telemetry/chat/review` currently treats these as failure outcomes:

* `guardrail_failure`
* `apply_failed`
* `repair_failed`
* `planner_exception`
* `planning_exhausted`
* `planning_missing`
* `planning_refusal`
* `planning_incomplete`
* `empty_edit_plan` — an `edit_plan` that carried no operations. It answers 200 and reads
  as success to the user, which is exactly why it counts as a failure here.

Thumbs-up and thumbs-down feedback, on both transports:

* `POST /telemetry/chat/feedback` with `{ traceId, session, rating: "up" | "down", note? }`
* `GET /telemetry/chat/feedback?session=<session>&rating=<rating>&traceId=<id>&limit=<n>`

## Cache metrics

Cache-related fields:

* `cacheReadInputTokens`
* `cacheCreationInputTokens`

Provider mapping:

* OpenAI: `cacheReadInputTokens` maps to `cached_tokens` (from usage details)
* Gemini: no cache fields are recorded
* Anthropic: `cacheReadInputTokens` maps to `cache_read_input_tokens`; `cacheCreationInputTokens` maps to `cache_creation_input_tokens`

## Source references

* `packages/orchestrator-core/src/telemetry/chat-telemetry.ts`
* `packages/orchestrator-core/src/chat/chat-pipeline.ts`
* `apps/orchestrator/src/index.ts` (list and review endpoints)
* `packages/orchestrator-core/src/errors.ts` (reason categories)

## What to read next

<CardGroup cols={2}>
  <Card title="Token usage tracking" icon="coins" href="/observability/token-usage-tracking">
    What each turn actually cost, per provider.
  </Card>

  <Card title="Chat troubleshooting" icon="bug" href="/observability/chat-troubleshooting">
    The playbook for a turn that produced the wrong operation, or none.
  </Card>
</CardGroup>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.