--- title: Observability description: Send traces, metrics, and logs from DocsGPT to any OpenTelemetry-compatible backend (Axiom, Honeycomb, Grafana, Datadog, Jaeger, etc.). --- import { Callout } from 'nextra/components' # Observability DocsGPT bundles the OpenTelemetry SDK and auto-instrumentation packages in `docsgpt/requirements.txt` — they install with the rest of the backend deps. Telemetry is **off by default**; opt in by prefixing the launch command with `opentelemetry-instrument` and setting OTLP env vars. Auto-instrumentation covers Flask, Starlette, Celery, SQLAlchemy, psycopg, Redis, requests, and Python logging. Agent runs, LLM calls, tool calls and retrieval are recorded by DocsGPT itself and exported as OpenTelemetry GenAI spans — see [Execution traces](#execution-traces). ## Enabling Set these env vars in your `.env` (or compose `environment:` block): ```bash OTEL_SDK_DISABLED=false OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf OTEL_EXPORTER_OTLP_ENDPOINT=https://your-collector.example.com OTEL_EXPORTER_OTLP_HEADERS=Authorization=Bearer%20 OTEL_TRACES_EXPORTER=otlp OTEL_METRICS_EXPORTER=otlp OTEL_LOGS_EXPORTER=otlp OTEL_PYTHON_LOG_CORRELATION=true OTEL_RESOURCE_ATTRIBUTES=service.name=docsgpt-backend,deployment.environment=prod ``` Then prefix the process command with `opentelemetry-instrument`. The simplest way is a compose override (no image rebuild): ```yaml # deployment/docker-compose.override.yaml services: backend: command: > opentelemetry-instrument gunicorn -w 1 -k uvicorn_worker.UvicornWorker --bind 0.0.0.0:7091 --config docsgpt/gunicorn_conf.py docsgpt.asgi:asgi_app environment: - OTEL_SERVICE_NAME=docsgpt-backend worker: command: opentelemetry-instrument celery -A docsgpt.app.celery worker -l INFO -B environment: - OTEL_SERVICE_NAME=docsgpt-celery-worker ``` For local dev, prepend `dotenv run --` so the `OTEL_*` vars from `.env` reach `opentelemetry-instrument` before it boots the SDK: ```bash dotenv run -- opentelemetry-instrument flask --app docsgpt/app.py run --port=7091 dotenv run -- opentelemetry-instrument celery -A docsgpt.app.celery worker -l INFO --pool=solo ``` Logs are exported in-process when `OTEL_LOGS_EXPORTER=otlp` is set — `docsgpt/core/logging_config.py` detects the flag and preserves the OTEL log handler. Without it, `logging` writes only to stdout. ## Execution traces Every request records an **execution trace**: a timed tree of the steps behind it. Traces are recorded for chat turns (`/stream`, `/api/answer`, `/v1/chat/completions`, including each round of a tool-approval pause), scheduled and webhook runs, workflows, the research agent, `/api/search`, the MCP `search_docs` tool, and graph builds. | Step | Recorded when | | --- | --- | | `invoke_agent` | An agent (or a workflow node's agent) runs | | `chat` | An LLM call made during the request, including retries, fallbacks, query rephrasing, prescreening, history compression and guardrail judges | | `execute_tool` | A tool is executed, paused for approval, denied or skipped | | `retrieval` | A retriever or the multi-source dispatcher searches | | `embeddings` | The query is embedded | | `search` | One source is searched | | `rerank` | Prescreening filters retrieved chunks | | `guardrail` | A guardrail calls a remote check or fires | | `step` | A workflow node or research phase runs | Traces are stored in the `request_traces` table and shown in the app: open **Settings → Logs** (or an agent's **Logs** tab), expand an entry and choose **View trace** to see a waterfall of every step with its timing, tokens, cost and details. Stored traces keep short previews — tool arguments and results, retrieved chunk titles and snippets, rephrased queries, answer excerpts — truncated and with secret-named fields redacted. Full prompts are never stored. When a guardrail fires during a request, every preview is dropped from its trace. ```bash TRACES_ENABLED=true # record traces at all TRACES_CAPTURE_CONTENT=true # keep previews in stored traces TRACES_PREVIEW_CHARS=2000 # characters kept per preview TRACES_MAX_SPANS=500 # steps kept per trace; the rest are counted TRACES_RETENTION_DAYS=30 # a daily task deletes older traces TRACES_OTEL_EXPORT=true # also export traces as OTel GenAI spans ``` ### GenAI spans and metrics When DocsGPT runs under `opentelemetry-instrument`, each finished trace is also exported as spans that follow the [OpenTelemetry GenAI semantic conventions](https://opentelemetry.io/docs/specs/semconv/gen-ai/): `invoke_agent {agent}`, `chat {model}`, `execute_tool {tool}`, `embeddings {model}` and `retrieval`, with attributes such as `gen_ai.provider.name`, `gen_ai.request.model`, `gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens`, `gen_ai.usage.cache_read.input_tokens`, `gen_ai.conversation.id`, `gen_ai.agent.id` and `gen_ai.tool.name`. DocsGPT-specific details use the `docsgpt.*` prefix (`docsgpt.request_id`, `docsgpt.token_source`, `docsgpt.cache_hit`, `docsgpt.ttft_ms`, ...). The trace's root span is a child of the request's HTTP server span, and the stored trace keeps the OTel trace id so you can move between the two. Two metrics are recorded for every model call: `gen_ai.client.token.usage` and `gen_ai.client.operation.duration`. Backends that understand the GenAI conventions — Langfuse (`/api/public/otel`), Arize Phoenix, Datadog LLM Observability, Grafana — render these as LLM traces with token and cost views. Prompt and tool content is **not** exported by default, because the OTLP backend may be a third party. Opt in with the standard variable: ```bash OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=SPAN_ONLY ``` This adds the same redacted previews the app stores (for example `gen_ai.tool.call.arguments` and `gen_ai.tool.call.result`). GenAI spans are exported when the request finishes, with their original timestamps. Consequences: a long research run appears only when it ends; HTTP and database spans made during a step sit beside the step rather than under it; and log records carry the request's span ids, not the step's. The GenAI conventions are still in development upstream, so attribute names may change in later releases. ## Backend examples ### Axiom ```bash OTEL_EXPORTER_OTLP_ENDPOINT=https://api.axiom.co OTEL_EXPORTER_OTLP_HEADERS=Authorization=Bearer%20xaat-XXXX,X-Axiom-Dataset=docsgpt OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf ``` `%20` is the URL-encoded space between `Bearer` and the token. Create the dataset in the Axiom UI before sending. ### Self-hosted OTLP collector / Jaeger / Tempo ```bash OTEL_EXPORTER_OTLP_ENDPOINT=http://otel-collector:4317 OTEL_EXPORTER_OTLP_PROTOCOL=grpc ``` ### Honeycomb / Grafana Cloud / Datadog Each vendor publishes a single-line `OTEL_EXPORTER_OTLP_ENDPOINT` plus `OTEL_EXPORTER_OTLP_HEADERS` recipe — drop them in alongside the service-name override. ## Caveats - The Dockerfile uses `gunicorn -w 1`. If you raise worker count, move SDK init into a `post_worker_init` hook to avoid one-thread-per-process exporter contention. - `asgi.py` wraps Flask in Starlette's `WSGIMiddleware`. Both instrumentors are installed, so each request produces a Starlette span enclosing a Flask span. Drop `opentelemetry-instrumentation-flask` from `requirements.txt` if the duplication is noisy. - OTEL packages add ~50 MB to the image. They install on every build — the runtime cost is zero unless you set `opentelemetry-instrument` on the command and set the OTLP env vars. - The OTEL exporter ecosystem currently caps `protobuf` at `<7`, so the backend runs on protobuf 6.x. This will catch up in a future OTEL release.