Covers what is recorded, where it is shown, the TRACES_* settings, the
gen_ai.* spans and metrics, content capture, and the limits of exporting
spans when the request ends.
TRACES_* settings for the per-request trace timeline and its GenAI OTel
export. opentelemetry-api was only transitive; the tracing package
imports it directly.
The DEFAULT 'unknown' added in the last pass stopped a previous-release
process from raising NotNullViolation mid-rollout, but it threw the
attribution away to do it. The highest-volume auth_events writers are OIDC
login and silent renewal, where the actor is simply the user, so a rollout
window would have flattened exactly the rows that were trivially recoverable.
Lifts both derivation rules into SQL functions -- auth_events_derive_actor
and auth_events_derive_target -- so the backfill and a BEFORE INSERT trigger
share one definition rather than two copies of the same CASE drifting apart.
The trigger guards on NEW.actor_id IS NULL, which identifies a legacy insert
exactly because the repository always supplies one. That distinction matters
for target_id: a current writer sets it NULL deliberately for events with no
user target, and deriving there would name the actor as their own target.
Kept rather than scheduled for removal: it is a no-op on the path the
repository takes, and it keeps the column's contract true for any writer that
bypasses it.
Correctness
- /api/remote never recorded source.created, so URL, GitHub and connector
sources had a source.deleted with no matching creation. All three creation
paths now go through one _audit_source_created helper.
- The prompt-cache rate divided cached tokens by a whole bucket's prompt
tokens. A bucket is a day and mixes calls whose provider reports a cache
breakdown with calls whose provider does not, so filtering buckets in the
client could not separate them and the rate was understated by however much
traffic ran on a non-reporting provider. The denominator is now computed in
SQL over the reporting rows.
- The outcome pill matched values nothing writes. Guardrails emit triggered /
not_evaluated and the device feed emits dispatched; the map had blocked /
denied / allowed, so a guardrail that fired rendered neutral grey -- the one
signal the merged feed exists to surface. Fixtures were seeding the
fictional values, so the tests passed on it too.
- Stream duration_ms timed the consumer. stream_token_usage is a generator,
so start-to-exhaustion includes the agent loop's tool handling and the SSE
client's pace; a slow browser recorded ~30s for a sub-second call. It now
accumulates only the time spent inside next().
Safety
- Activity filters failed open: an unknown facet or unparseable timestamp was
dropped, and no filter means every row, so a typo widened an audit view and
on the export streamed the full history. Both are now a 400.
- The search term was interpolated into an ILIKE pattern, so "100%" matched
everything and "q1_report" matched more than it should. Escaped.
- 0034 set actor_id NOT NULL with no default. A previous-release process
inserting mid-rollout would raise, and in admin/routes.py that insert shares
the request transaction, so a role grant beside it would roll back too.
Noise and dead code
- The per-user panel is a security panel: data-plane events file under the
actor, so an active account's routine deletes pushed a denied login out of
the 20-row window. It now excludes them; the Activity tab shows everything.
- device_audit_log had no created_at-leading index, so the merged feed
sequentially scanned that branch every page (migration 0036).
- conversation.deleted_all no longer records when nothing was deleted, and
agent.updated no longer records an empty field list.
- Dropped by_model from /admin/usage (no consumer; an extra aggregate per page
load), the duplicate filter surface on AuthEventsRepository that nothing
called, and the unreachable FLOW_LABELS.schedule entry.
- Type hints on record_event's conn and the remaining unannotated helpers.
Covers what the previous commits changed: the actor_id / target_id split and
the system actors, the data-plane events, the merged Activity feed and its
export, and the new usage and per-user spend endpoints.
Also corrects "there is no UI for these events yet" in the OIDC login
auditing section, which is no longer true.
- check_usage treats a request through a keyless (draft) agent as agent
traffic, matching how its usage rows are bucketed and the headless rule
- dashboard edits carry the stored enabled flag instead of re-enabling the
policy; disabled policies are labelled in the Quotas tab and the editor
- quota 429s send x-should-retry: false so OpenAI SDK clients do not retry
a refusal that cannot succeed before the reset
- cached-input and cache-write rates for Anthropic, OpenRouter and Groq
gpt-oss-120b; refresh OpenRouter deepseek-v3.2 list prices
- UsageQuota reuses usagePercent; docs note that a user override needs an
existing user
- A tool continuation refused for usage now releases the resume claim it
took; before, retries got a 409 until the stale claim was reverted.
- Agent traffic is any row with an agent key or an agent id, so keyless
agents and workflow nodes count toward the agent bucket, not direct.
- The user quota modal discards responses for a previously opened user.
- The usage meter shows every limited bucket, not only 'all'.
- Restore the class separator on the analytics stat card that a formatter
run removed, and align the OpenRouter DeepSeek description with its rates.
How the instance default, team allowances and user overrides resolve
(including users in several teams), the quota window, who is charged for
agent traffic, how cost budgets price models and what happens to unpriced
ones, and the admin and user API.
Rename the unused *_cost_per_token capability fields to USD per 1M tokens,
add prompt-cache read/write rates, and ship list prices for the hosted
catalogs. The old per-token keys still load, scaled, with a warning.
docsgpt/pricing.py turns a call's token bins into a USD cost. Models with
no declared rate cost $0 unless QUOTA_UNPRICED_RATE_PER_MILLION is set.
POST /api/user/tokens/<id>/regenerate swaps the secret of an existing token
in place: name, scopes and restrictions stay, the old secret stops matching
at once, and the expiry is reset. The lifetime defaults to the one the token
was last issued with (clamped to today's policy) or to expires_in_days when
given. An expired token can be renewed this way; a revoked one cannot. It is
session only like the rest of token management, and writes a pat_regenerated
audit event. regenerated_at records the rotation.
A resource restriction was checked on ids in the request, not on what the
addressed row pulls in or belongs to. Closed:
- Workflow writes for tokens restricted on sources, tools or prompts (a graph
names those inside its nodes), and attaching a workflow to an agent unless
the token is restricted on workflows too.
- Chat for tools-restricted tokens (chat executes tools; rejected at token
creation as well), and agent-less chat for tokens restricted on prompts or
workflows.
- conversation_id on chat: it must belong to the agent being run, or to no
agent for agent-less chat. Otherwise the server continued, appended to, or
resumed pending tool calls of another agent's conversation.
- Schedules for tokens restricted on anything but agents; schedule-id routes
for every restricted token.
- Conversations and analytics for every restricted token, not only
agent-restricted ones.
Also: create, first publish and adopt return the agent API key masked to a
token without agents:keys; token ids must be canonical UUIDs (urn:uuid: gave
a 500); an expired token is reported as expired; token creation takes a
per-user advisory lock so the cap cannot be raced; admin revoke-sessions
writes a pat_revoked event per token. The UI drops a row whose revoke returns
404 and does not offer a tools restriction next to chat:run.
The ASGI message events route now accepts the same scopes as its Flask
sibling (conversations:read or chat:run) through a shared constant. A token
request that fails routing gets Flask's 404/405 instead of a 403. Creating a
token retires an expired token that still held the name. allowed_ids uses
is_pat instead of a bare literal. PAT_ENABLED is documented as the master
switch it is: turning it off stops every existing token from authenticating.
The docs explain that sources are matched by name (oldest wins) and point CI
flows at sources upload --replace.
Message tail and the ASGI reconnect stream cannot tie a message to an
allowlist, so any token with a resource filter is refused there. The tools
listing now applies the allowlist to default and builtin rows as well. Token
creation answers 400 instead of 500 for a JSON body that is not an object.
A personal_access_tokens table (migration 0032) holds scoped user-level API
credentials. Only the SHA-256 of the secret is stored, like device session
tokens. Lookups exclude revoked and expired tokens and the tokens of
deactivated users. PAT_* settings cover the feature switch, default and
maximum lifetime, the operator opt-in for non-expiring tokens and the
per-user cap.
Graph retrieval tied plain vector search at best and never beat it. Measured
across five corpora, the bottleneck was seeding, not the graph: the walk
started from nodes whose embeddings were computed from bare entity names, and
a whole question shares almost nothing with a name like "Quill".
Extraction now embeds each node from "name (type): description" and each
relationship as the fact it asserts ("Alder streams_to Quill: ..."), stored on
a new nullable graph_edges.fact_embedding column that ensure_vector_schema adds
in place. Entity names are canonicalised (case, punctuation, word breaks and a
cautious plural) so "VECTOR_STORE" and "vector stores" land on one node. Extraction calls run
concurrently (GRAPHRAG_EXTRACTION_WORKERS, default 8) while embedding and graph
writes stay serial on the task thread, so ordering and idempotency are
unchanged; that measured 8.4x faster with identical output.
Retrieval gains per-source options, stored under retrieval.graph and read live
at query time:
- seed_strategy: start from matching entities (default) or matching
relationships, which can reach an entity the question never names;
- passage_nodes (on): walk the source's passages alongside entities, with
PageRank damping 0.5 instead of 0.85;
- blend_vector (on): fuse the graph ranking with the source's vector ranking
by reciprocal rank.
The defaults are the measured-best configuration. Through GraphRAGRetriever,
the new seeding moved recall@4 from 0.41 to 0.68 on a multi-hop corpus and
from 0.50 to 1.00 on the docs corpus, and regressed none of the corpora
measured. Existing graphs keep name-only embeddings until rebuilt.
The docs site failed to build: a bare "<= 1" in the prose of the
generated page is parsed by MDX as the start of a JSX tag ("Unexpected
character '=' before name"). Constraints are rendered as code spans now,
where MDX leaves them alone, and a test rejects any bare <, { or } outside
a code span so a future description cannot reintroduce the failure.
Verified with a local next build of the docs site.
Review follow-up. The per-group secret validators normalised a hand-picked
list of API keys, which left other optional credentials and overrides
(OPEN_ROUTER_API_KEY, S3 and Daytona keys, ELASTIC_PASSWORD, the OIDC
trio, connector client ids, MICROSOFT_AUTHORITY, MCP_OAUTH_REDIRECT_URI)
holding the literal "None" or "" a .env file spells "unset" with, so
truthiness checks and fallbacks downstream saw a value. One rule on the
group base replaces those lists: every Optional[str] field maps "", "None"
and whitespace to None and strips real values. Plain str fields are left
alone. The OIDC required-settings check therefore also rejects those
spellings.
EMBEDDINGS_POOLING is Literal["cls", "mean"] with case-insensitive
parsing; its consumer silently ignored anything else.
Bounds added where the consumer rejects or misbehaves on the value:
SCHEDULE_RUN_OUTPUT_RETENTION_DAYS and MESSAGE_EVENTS_RETENTION_DAYS (the
cleanup repositories raise on <= 0), EMBEDDINGS_DELEGATE_TIMEOUT, the
remote-device idle/pairing/invocation TTLs and CELERY_VISIBILITY_TIMEOUT
(> 0), REMOTE_DEVICE_CMD_QUEUE_TTL_SECONDS (> 605, the documented drain
deadline), GRAPHRAG_MAX_CHUNKS_FOR_EXTRACTION (>= 0; negative would slice
the pending list from the end).
The generated reference now renders generic type arguments
(dict[str, int] rather than dict).
`docsgpt up --native` installs services meant to outlive the shell. Development
wants the opposite, and until now it meant three terminals from the guide:
uvicorn, celery, and vite.
`docsgpt dev` runs this checkout's API and worker as children of one terminal,
both restarting when a file is saved, their output interleaved and labelled, and
Ctrl-C stopping them together. `--ui` adds the Vite dev server, `--mock-llm`
runs the bundled mock model so no API key is needed, and `--no-worker` leaves
the worker to your editor's debugger. Celery has no reloader of its own, so the
worker is wrapped in watchfiles when it is installed, and runs plain when it is
not.
Alongside it, the commands a dev loop keeps reaching for:
- `docsgpt doctor` checks what usually breaks a new setup: PostgreSQL answering
and its schema matching this version, Redis answering, a model provider being
configured, and the port being free.
- `docsgpt restart [api|worker]` bounces services without rewriting settings or
rerunning migrations, which `down` plus `up` did.
- `docsgpt logs -f` follows a native install instead of telling you to run
`tail -f` yourself.
- `docsgpt env set` applies itself to a running native install rather than
asking you to run `docsgpt up` again to change one value.
Two bugs found on the way, both older than this change:
- `docsgpt api --reload` watched the working directory, which in a checkout is
178,425 files: .venv, node_modules, and the indexes/ and inputs/ the app
writes to while ingesting, so the server restarted itself mid-request. It
watches the package now — 1,217 files.
- The VS Code "Flask Debugger" ran `flask run`, which serves only the WSGI app:
/mcp, the SSE streams and artifact downloads 404 under it. The guide warned
about this in prose while the debug config did it anyway. It runs uvicorn on
the ASGI app now, like production.
About 85 call sites read a setting as getattr(settings, "NAME", fallback),
each carrying its own copy of the default. Every one of those names is a
field with a default on the model, so the fallback could never apply to
the real settings object; it only masked drift. Two had drifted:
- OPENAI_PROMPT_CACHE_KEY defaults to True on the model but the reader
fell back to False, and two test stubs relied on that.
- SharePoint's MICROSOFT_AUTHORITY fallback to
https://login.microsoftonline.com/<tenant> never fired, because the
attribute always exists (as None), so MSAL got authority=None. The
connector now derives the tenant authority when the setting is unset,
as its test always assumed.
Four places read EMBEDDINGS_KEY straight from os.environ, skipping the
"None"/"" normalisation the model applies; they read the setting now.
Test stubs that replaced a module's settings with a SimpleNamespace list
every setting the code under test reads.
SAGEMAKER_REGION, SAGEMAKER_ACCESS_KEY and SAGEMAKER_SECRET_KEY survive
only as a fallback for the S3_* credentials. They carry
Field(deprecated=...) now, so any read emits a DeprecationWarning naming
the replacement and the generated reference shows the notice. The S3
store is the one sanctioned reader; it silences that warning locally
because it already logs its own operator-facing one when the fallback
is actually used.
DEFAULT_MAX_HISTORY was referenced nowhere. RETRIEVERS_ENABLED was read by
no code at all, while two docs pages described it as an enforced
allow-list; both the setting and those claims are removed.
The "AUTH_TYPE=oidc requires OIDC_ISSUER, OIDC_CLIENT_ID and
OIDC_FRONTEND_URL" check lived in app.py, so it only ran when the Flask
app was imported; a worker or script with the same misconfiguration
started fine. It is now a model validator on the auth group and runs
wherever Settings is loaded, with the same message.
DEPLOYMENT_TYPE, which app.py read straight from the environment to
decide whether a missing JWT_SECRET_KEY is fatal, is a documented
setting on the server group now, so it shows up in the reference like
every other variable the app reads.
Enum-like settings whose allowed values were only listed in a comment are
now Literal types, so a typo fails at startup with a message naming the
allowed values instead of falling through to a default with a warning
(or, for VECTOR_STORE, failing on first use):
AUTH_TYPE, VECTOR_STORE, STORAGE_TYPE, URL_STRATEGY, OCR_BACKEND,
OCR_ENGINE, SANDBOX_BACKEND, DOC_PARSER_ENGINE, TTS_PROVIDER, STT_PROVIDER
Each keeps a before-validator that strips and lower-cases the value, since
the registries that consume them already lower-cased at the use site, and
AUTH_TYPE maps the "None"/"none"/"" spellings a .env file carries to None
(it was the string "None" before, which only worked because nothing
compared against it). An empty TTS/STT provider still means "off".
LLM_PROVIDER stays a plain str because providers are plugin-extensible.
Containers are typed (dict[str, int], list[str], dict[str, Any]) instead
of bare dict/list, six fields that were Optional with a non-None default
are plain, and integer settings whose description already states a range
carry it as a constraint (ge=0 for "0 disables", ge=1 for counts that
cannot be zero, 0 < threshold <= 1).
The hand-maintained settings page documented 95 of 258 settings and
.env-template 42, and both drifted as fields were added. The field
descriptions now live on the model, so the reference is rendered from it:
python -m docsgpt.core.settings.reference --write
writes docs/content/Deploying/Settings-Reference.mdx, one section per
settings group with each field's type, default, constraints, aliases and
description. --check reports a stale page, and tests/core/test_settings.py
fails when the checked-in page no longer matches the definitions, so a
new setting cannot land undocumented.
test_settings.py also pins the composition contract: every group field is
a flat Settings attribute, no field is defined twice, every field has a
description, and the secret-normalising validator of every group is
applied (the case that a shared method name would silently drop).
The App Configuration page points at the reference instead of at
settings.py, and the reference is listed in the Deploying navigation.
docsgpt/core/settings.py had grown to 258 fields in one 600-line class,
touched by about two commits a week, with related settings scattered
(GitHub ingest caps inside the embeddings block, API keys in four places,
the OpenAI Responses knobs 100 lines from the other OpenAI fields).
It is now a package: one module per domain (auth, llm, embeddings,
retrieval, vectorstores, database, workers, ingestion, ocr, storage,
connectors, server, events, agents, guardrails, scheduler, sandbox,
speech), each a SettingsGroup owning its fields and validators, composed
by multiple inheritance into the same flat Settings class. Every
attribute name, type, default, alias and constraint is unchanged, so
settings.NAME reads, .env files and test monkeypatches all keep working;
the import path docsgpt.core.settings is the package. Settings.normalize_api_key
is kept as a classmethod for callers that reuse it.
The comment above or beside each field became its Field(description=...),
so the definitions are visible to tooling; the next commit generates the
docs reference from them.
Pitfall recorded for future groups: pydantic collects validators by
method name across the MRO, so two groups naming a validator the same
would silently keep only one. Each group's validator has a unique name.
From review of #2800:
- systemd `enable --now` starts nothing when the unit is already active, so a
second `up --native` kept the old ExecStart and left the API on its previous
port. start enables and then restarts, as the launchd path already did by
booting the job out first.
- An explicit `home` now wins over XDG_CONFIG_HOME, which is what callers pass
it for.
- The ExecStart program must be a real executable: when `docsgpt` is not on
PATH, sys.argv[0] is accepted only if it can be run, and otherwise the
failure is raised before any unit is written.
- `up --native` over a directory holding a Docker install now refuses and says
how to proceed, instead of starting native services beside containers that
down, status and uninstall would no longer see.
- Docs: without a terminal only --postgres-uri is required, and the Windows
fallback names `docsgpt beat`, which the worker cannot embed there.
SystemdServices was the least covered part of the module and cannot be run on
this machine, so it now has tests for install, start, stop, remove, is_running
and a failing systemctl.
From review of #2799:
- The archive is created 0600 rather than at the process umask: it holds the
install's data, and --with-settings puts .env and its secrets in it.
- restore validates everything the manifest declares before the stack is
stopped, so a damaged archive fails while DocsGPT is still running rather
than after `compose down` has taken it away.
- Only the volumes a backup is made of are restored. A hand-made manifest can
no longer point import_volume at postgres_data, whose contents it empties.
- psql runs with ON_ERROR_STOP=on, so a restore that fails halfway cannot
start DocsGPT again and call it a success.
- import_volume unpacks into the container's own filesystem first and clears
the live volume only once the tar has come out whole, so a corrupt one
leaves the volume as it was.
- The backend and the worker stop while the archive is made and start again
even if the dump fails, so the dump and the volume tars describe the same
moment instead of drifting apart as ingestion writes.
Run DocsGPT without Docker: the API and the worker each become a service
on the machine itself, a launchd agent on macOS and a systemd user unit
on Linux, pointed at a PostgreSQL and a Redis that already run.
`docsgpt up --native --postgres-uri ... --redis-url ...` writes the same
.env a Docker install uses, applies the migrations and starts both
services. status, logs, down and uninstall work on a native install the
same way they do on a Docker one, and never touch the database or Redis:
they were the user's to begin with.
One Redis URL covers the broker, the result backend and the cache on
three consecutive databases, starting at the one the URL names, so a
Redis that already holds something else can be shared.
Windows has neither service manager, so native mode refuses it and says
what to do instead.
`docsgpt backup` writes one archive holding a pg_dump of the database, a tar
of each data volume and a manifest of what it came from; `docsgpt restore`
puts it back over an install. The settings file is left out unless
--with-settings asks for it, since it holds the install's secrets, and a
backup taken with a newer DocsGPT is refused without --force.
The volume tars go through the image the install already runs, so a backup
pulls nothing extra, and compose calls can now redirect stdout and stdin so
the dump never passes through this process.
- install.sh saves the get.docker.com and uv installers to a file and runs
them only after the download finished, so a cut-off transfer runs nothing.
- Neither installer prints DOCSGPT_PACKAGE, which may be a URL with
credentials.
- The CI step assigns the wheel path before exporting it, so a missing wheel
fails instead of installing from PyPI.
- Docker-Deploying shows one code block per platform; Quickstart names the
/opt/docsgpt home used for root on Linux.
deployment/install.sh (curl | bash) and install.ps1 (irm | iex) check for
Docker, install uv when it is missing or older than 0.8 (pinned 0.12.15 via
Astral's installer), install or upgrade the docsgpt package with
`uv tool install`, and hand the terminal to `docsgpt up` with any arguments.
On Linux without Docker the shell installer offers get.docker.com. Both run
entirely inside a function, so a download cut short runs nothing.
Releases attach both scripts next to the Compose file, which is where
docs.ac/install and docs.ac/install.ps1 will point. installer-lint.yml runs
shellcheck and the PowerShell parser; docker-image-verify.yml now installs
through install.sh. README, Quickstart, Docker-Deploying and the changelog
lead with the one-liner.
- wait_healthy starts no request once the deadline is reached, and neither
its pauses nor the requests after the first run past it; the first attempt
still always runs (status uses a zero timeout).
- envfile writes $ as $$ inside double quotes, which Compose interpolates,
and reads $$ back as $, so a value such as pa$w'rd reaches the container
unchanged.
- Upgrading no longer describes the working directory as the data home.
Docker-Deploying gains a `docsgpt up` section, Pip-Install and Upgrading
describe the ~/.docsgpt/server data home, and the changelog covers both.
docker-image-verify.yml installs the wheel and runs `docsgpt up`, `status`,
a second `up` that must keep the secrets, and `uninstall --purge` against
the image it built. The standalone Compose file maps host.docker.internal
to the host gateway, so a model server on a Linux host is reachable the way
`docsgpt up` suggests.
The backend image builds the web UI with scripts/build_frontend.sh and
serves it through docsgpt/ui.py, so the standalone Compose file drops the
frontend container. UI and API share port 7091, published on 127.0.0.1
unless DOCSGPT_BIND says otherwise. POSTGRES_PASSWORD is configurable, and
an optional https profile puts Caddy in front of a public domain.
docker-image-verify.yml starts the standalone stack on the image it built
and checks the API, the UI, /config.js and a client-side route on one port.
- Ship tiktoken's cl100k_base inside the package and build the encoding
from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
temp dir, and read tokenizer.json and repo metadata from that cache, so
a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
the endpoints return 404, audio files fail to ingest with a clear
message, /api/config reports tts_available/stt_available, and the UI
hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.