Correctness
- /api/remote never recorded source.created, so URL, GitHub and connector
sources had a source.deleted with no matching creation. All three creation
paths now go through one _audit_source_created helper.
- The prompt-cache rate divided cached tokens by a whole bucket's prompt
tokens. A bucket is a day and mixes calls whose provider reports a cache
breakdown with calls whose provider does not, so filtering buckets in the
client could not separate them and the rate was understated by however much
traffic ran on a non-reporting provider. The denominator is now computed in
SQL over the reporting rows.
- The outcome pill matched values nothing writes. Guardrails emit triggered /
not_evaluated and the device feed emits dispatched; the map had blocked /
denied / allowed, so a guardrail that fired rendered neutral grey -- the one
signal the merged feed exists to surface. Fixtures were seeding the
fictional values, so the tests passed on it too.
- Stream duration_ms timed the consumer. stream_token_usage is a generator,
so start-to-exhaustion includes the agent loop's tool handling and the SSE
client's pace; a slow browser recorded ~30s for a sub-second call. It now
accumulates only the time spent inside next().
Safety
- Activity filters failed open: an unknown facet or unparseable timestamp was
dropped, and no filter means every row, so a typo widened an audit view and
on the export streamed the full history. Both are now a 400.
- The search term was interpolated into an ILIKE pattern, so "100%" matched
everything and "q1_report" matched more than it should. Escaped.
- 0034 set actor_id NOT NULL with no default. A previous-release process
inserting mid-rollout would raise, and in admin/routes.py that insert shares
the request transaction, so a role grant beside it would roll back too.
Noise and dead code
- The per-user panel is a security panel: data-plane events file under the
actor, so an active account's routine deletes pushed a denied login out of
the 20-row window. It now excludes them; the Activity tab shows everything.
- device_audit_log had no created_at-leading index, so the merged feed
sequentially scanned that branch every page (migration 0036).
- conversation.deleted_all no longer records when nothing was deleted, and
agent.updated no longer records an empty field list.
- Dropped by_model from /admin/usage (no consumer; an extra aggregate per page
load), the duplicate filter surface on AuthEventsRepository that nothing
called, and the unreachable FLOW_LABELS.schedule entry.
- Type hints on record_event's conn and the remaining unannotated helpers.
Three gaps in one view.
Cost: token_usage.cost has been written on every call since quotas landed, and
tokens_by_model already returned it, but bucketed_totals and top_token_users
selected tokens only. An admin could set a USD quota in /admin/quotas and had
no way to see the spend it was capping. Buckets and top users now carry cost,
the endpoint reports a window total and a per-model split, and the chart takes
a Tokens/Cost toggle with a currency axis.
Group-by: the endpoint has supported group_by=model|agent|source from the
start and the UI only ever sent bucket=day. The selector is now wired, and a
grouped series pads missing buckets so a model that was idle on Tuesday plots
a zero instead of shifting its whole row one bar left.
Latency: surfaced from the columns the previous commit added, as p50/p95 with
a median time-to-first-token underneath. Unmeasured rows are excluded rather
than counted as zero, and the card says so when nothing was measured.
Also surfaces the prompt-cache hit rate, computed only over rows whose
provider reported a cache breakdown -- NULL means "not reported", and folding
those in as 0% would understate it.
Per-user drill-down: GET /api/admin/users/<id>/usage returns a daily series
plus splits by model and by flow, reachable from the top-users table and from
the user detail dialog. The detail dialog previously showed one tokens_30d
number, which answers neither "what is this person costing" nor "what is
driving it".
A resource restriction was checked on ids in the request, not on what the
addressed row pulls in or belongs to. Closed:
- Workflow writes for tokens restricted on sources, tools or prompts (a graph
names those inside its nodes), and attaching a workflow to an agent unless
the token is restricted on workflows too.
- Chat for tools-restricted tokens (chat executes tools; rejected at token
creation as well), and agent-less chat for tokens restricted on prompts or
workflows.
- conversation_id on chat: it must belong to the agent being run, or to no
agent for agent-less chat. Otherwise the server continued, appended to, or
resumed pending tool calls of another agent's conversation.
- Schedules for tokens restricted on anything but agents; schedule-id routes
for every restricted token.
- Conversations and analytics for every restricted token, not only
agent-restricted ones.
Also: create, first publish and adopt return the agent API key masked to a
token without agents:keys; token ids must be canonical UUIDs (urn:uuid: gave
a 500); an expired token is reported as expired; token creation takes a
per-user advisory lock so the cap cannot be raced; admin revoke-sessions
writes a pat_revoked event per token. The UI drops a row whose revoke returns
404 and does not offer a tools restriction next to chat:run.
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.
Kept for one release:
- A top-level application package whose meta-path finder resolves
application.x.y to the already-imported docsgpt.x.y object, so old imports
and entry points (celery -A application.app.celery,
uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
docsgpt.* task on start-up, so messages queued by the previous release still
run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
the previous release wrote are left unread instead of firing twice.
The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.