Skip to main content
This page is for the self-hoster and the developer reading the orchestrator’s chat telemetry: every event the chat pipeline records, its fields, and the endpoints that serve it.

Scope

This covers events pushed via ctx.chatTelemetry.push(...) in:
  • packages/orchestrator-core/src/chat/chat-pipeline.ts
And serialized/logged by:
  • packages/orchestrator-core/src/telemetry/chat-telemetry.ts
It does not cover generic operational logs such as chat_pipeline_start or image rewrite logs.

Event transport

Each telemetry entry is:
  1. Buffered in memory.
  2. Optionally persisted as NDJSON.
  3. Emitted to server logs as a structured log with event: "chat_telemetry".
Persistence is on by default, except under NODE_ENV=test. Set CHAT_TELEMETRY_PERSIST=0 to keep telemetry in memory only.
  • File: CHAT_TELEMETRY_FILE, else .data/chat-telemetry.ndjson in the data directory. The Docker image sets /app/.data/chat-telemetry.ndjson.
  • Buffer: the newest CHAT_TELEMETRY_LIMIT entries, default 500, which are also what is reloaded from the file at startup.
Rows carry a prompt excerpt, the first 180 characters of the message, and a 16-character SHA-256 prefix of the full prompt.

Event schema

Fields from ChatTelemetryEntry:
  • Required: id, at, phase, session, requestedSlug, effectiveSlug, plannerSource, modelKey, modelUsed, promptHash, promptExcerpt, promptLength
  • Optional classification: outcome, reason, reasonCategory
  • Optional plan shape: intent, opCount, opTypes
  • Optional usage/cost: inputTokens, outputTokens, totalTokens, cacheReadInputTokens, cacheCreationInputTokens, estimatedUsd
  • Optional timing: totalDurationMs, planningDurationMs, firstPlanningTokenMs, applyDurationMs, imageResolutionDurationMs, planningAttempts, and timelineStage on milestone rows (request_received, first_token, first_structured_progress, plan_ready, first_op_applied, done)
  • Optional apply detail: skippedOpCount, imagesRequested, imagesResolved
  • Optional planner context: plannerTier (forced_deterministic, deterministic, llm_intent_router, full_llm, demo), contractMode (minimal, targeted, full), contractBytes, contractBlockCount, contextPackBytes, strictJsonEnabled, schemaRetryUsed, plannerRefusal, plannerIncomplete, compactContextEnabled, minimalContextEnabled
  • Optional change-log drift: changelogMissingCount, changelogExtraCount, changelogFieldMislabelCount
  • Optional tool calls: toolName, toolOk, toolLatencyMs, toolAttempts, toolErrorCode, correlationId
  • Optional suggestion pills: suggestionIds, suggestionSources, clickedSuggestionId
reasonCategory values are from guardrail classification:
  • schema_violation
  • ambiguity
  • not_found
  • no_effective_change
  • planner_refusal
  • incomplete_output
  • malformed_output
  • internal_error
  • canceled
  • operation_failed
  • unsupported_by_site — the operation is valid, but the site declared it cannot honour it

Phases

phase is one of:
  1. received
  2. milestone
  3. forced_plan
  4. deterministic_plan_generated
  5. plan_attempt_failed
  6. plan_generated
  7. plan_apply_failed
  8. repair_attempt
  9. repair_generated
  10. tool_call
  11. result

Outcomes

Emitted telemetry outcomes (chatTelemetry.push)

  1. guardrail_failure
  2. needs_clarification
  3. plan_ready_for_approval
  4. no_effective_change
  5. applied
  6. apply_failed
  7. apply_pending_plan_error
  8. forced_duplicate_page
  9. forced_create_page
  10. planner_exception
  11. deterministic_plan_ready
  12. compound_deterministic_plan_ready
  13. attempt_${attempt}_failed (dynamic, e.g. attempt_1_failed)
  14. planning_exhausted
  15. planning_missing
  16. planning_refusal
  17. planning_incomplete
  18. repair_started
  19. repair_plan_generated
  20. repair_failed
  21. content_answer
  22. variation_request_redirect
  23. blocked_structural_capability
  24. llm_router_plan_ready
  25. llm_router_needs_clarification
  26. tool_ok
  27. tool_error
  28. api_error
  29. empty_edit_plan

Debug-only outcomes (response payload, not telemetry rows)

  1. validation_error
  2. pending_plan_missing
  3. pending_plan_mismatch
  4. info
  5. blocked_demo_mode

Phase to outcome mapping

Common mappings in current implementation:
  • received: no outcome
  • forced_plan: forced_duplicate_page, forced_create_page
  • deterministic_plan_generated: deterministic_plan_ready
  • plan_attempt_failed: attempt_${attempt}_failed
  • milestone: no outcome; timelineStage says which point the turn reached
  • plan_generated: usually no outcome (plan metadata + optional usage)
  • tool_call: tool_ok or tool_error
  • plan_apply_failed: apply_failed
  • repair_attempt: repair_started
  • repair_generated: repair_plan_generated
  • result: terminal or branch outcomes such as applied, needs_clarification, planning_exhausted, repair_failed, etc.

Telemetry APIs

The list and review endpoints are served by the standalone server. Library mode serves only the feedback endpoint below. On the standalone server these routes are not behind the access gate, and rows include prompt excerpts, so do not expose the server publicly. A PUBLIC_DEMO answers 403 on them. List entries (limit default 100, max 1000):
  • GET /telemetry/chat?limit=<n>&outcome=<outcome>&phase=<phase>&session=<session>
Review aggregate (limit default 300, max 2000):
  • GET /telemetry/chat/review?limit=<n>&session=<session>
/telemetry/chat/review currently treats these as failure outcomes:
  • guardrail_failure
  • apply_failed
  • repair_failed
  • planner_exception
  • planning_exhausted
  • planning_missing
  • planning_refusal
  • planning_incomplete
  • empty_edit_plan — an edit_plan that carried no operations. It answers 200 and reads as success to the user, which is exactly why it counts as a failure here.
Thumbs-up and thumbs-down feedback, on both transports:
  • POST /telemetry/chat/feedback with { traceId, session, rating: "up" | "down", note? }
  • GET /telemetry/chat/feedback?session=<session>&rating=<rating>&traceId=<id>&limit=<n>

Cache metrics

Cache-related fields:
  • cacheReadInputTokens
  • cacheCreationInputTokens
Provider mapping:
  • OpenAI: cacheReadInputTokens maps to cached_tokens (from usage details)
  • Gemini: no cache fields are recorded
  • Anthropic: cacheReadInputTokens maps to cache_read_input_tokens; cacheCreationInputTokens maps to cache_creation_input_tokens

Source references

  • packages/orchestrator-core/src/telemetry/chat-telemetry.ts
  • packages/orchestrator-core/src/chat/chat-pipeline.ts
  • apps/orchestrator/src/index.ts (list and review endpoints)
  • packages/orchestrator-core/src/errors.ts (reason categories)

Token usage tracking

What each turn actually cost, per provider.

Chat troubleshooting

The playbook for a turn that produced the wrong operation, or none.