Files
DocsGPT/docsgpt/alembic
arc53-machine b17f40a993 feat(usage): persist LLM call latency
The wrappers in docsgpt/usage.py already measured how long each call took --
for the llm_gen_finished / llm_stream_finished log lines -- and then threw the
number away. Nothing in the schema recorded it, so an operator could see what
an instance spent but never how slow it was.

Adds token_usage.duration_ms and token_usage.ttft_ms, and threads the
measurements the wrappers already take into the insert. The stream wrapper now
also stamps the moment of the first yielded chunk.

Both columns are nullable, and deliberately so: ttft_ms is NULL for a
non-streaming call and for a stream that failed before yielding anything, and
every row written before this migration is NULL too. A 0 there would drag a
p50 toward an instant first token that never happened.
2026-09-22 10:28:20 +01:00
..