Files
DocsGPT/tests/storage
arc53-machine a56f9df122 feat(admin): cost, group-by and latency in usage
Three gaps in one view.

Cost: token_usage.cost has been written on every call since quotas landed, and
tokens_by_model already returned it, but bucketed_totals and top_token_users
selected tokens only. An admin could set a USD quota in /admin/quotas and had
no way to see the spend it was capping. Buckets and top users now carry cost,
the endpoint reports a window total and a per-model split, and the chart takes
a Tokens/Cost toggle with a currency axis.

Group-by: the endpoint has supported group_by=model|agent|source from the
start and the UI only ever sent bucket=day. The selector is now wired, and a
grouped series pads missing buckets so a model that was idle on Tuesday plots
a zero instead of shifting its whole row one bar left.

Latency: surfaced from the columns the previous commit added, as p50/p95 with
a median time-to-first-token underneath. Unmeasured rows are excluded rather
than counted as zero, and the card says so when nothing was measured.

Also surfaces the prompt-cache hit rate, computed only over rows whose
provider reported a cache breakdown -- NULL means "not reported", and folding
those in as 0% would understate it.

Per-user drill-down: GET /api/admin/users/<id>/usage returns a daily series
plus splits by model and by flow, reachable from the top-users table and from
the user detail dialog. The detail dialog previously showed one tokens_30d
number, which answers neither "what is this person costing" nor "what is
driving it".
2026-09-22 10:33:56 +01:00
..
2026-03-30 16:13:08 +01:00