mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-04 18:13:03 +00:00
The wrappers in docsgpt/usage.py already measured how long each call took -- for the llm_gen_finished / llm_stream_finished log lines -- and then threw the number away. Nothing in the schema recorded it, so an operator could see what an instance spent but never how slow it was. Adds token_usage.duration_ms and token_usage.ttft_ms, and threads the measurements the wrappers already take into the insert. The stream wrapper now also stamps the moment of the first yielded chunk. Both columns are nullable, and deliberately so: ttft_ms is NULL for a non-streaming call and for a stream that failed before yielding anything, and every row written before this migration is NULL too. A 0 there would drag a p50 toward an instant first token that never happened.