mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-04 10:13:06 +00:00
Three gaps in one view. Cost: token_usage.cost has been written on every call since quotas landed, and tokens_by_model already returned it, but bucketed_totals and top_token_users selected tokens only. An admin could set a USD quota in /admin/quotas and had no way to see the spend it was capping. Buckets and top users now carry cost, the endpoint reports a window total and a per-model split, and the chart takes a Tokens/Cost toggle with a currency axis. Group-by: the endpoint has supported group_by=model|agent|source from the start and the UI only ever sent bucket=day. The selector is now wired, and a grouped series pads missing buckets so a model that was idle on Tuesday plots a zero instead of shifting its whole row one bar left. Latency: surfaced from the columns the previous commit added, as p50/p95 with a median time-to-first-token underneath. Unmeasured rows are excluded rather than counted as zero, and the card says so when nothing was measured. Also surfaces the prompt-cache hit rate, computed only over rows whose provider reported a cache breakdown -- NULL means "not reported", and folding those in as 0% would understate it. Per-user drill-down: GET /api/admin/users/<id>/usage returns a daily series plus splits by model and by flow, reachable from the top-users table and from the user detail dialog. The detail dialog previously showed one tokens_30d number, which answers neither "what is this person costing" nor "what is driving it".