The agent and workflow reads now return resource_states to people who may
edit them: every attached tool, source and prompt (and workflow node tool
and source) with active or stopped and the reason: deleted,
owner_lost_access, the sponsor reasons, connection_needs_reconnect,
connection_removed or connector_disabled. Each entry names the sponsor,
someone other than the reader who can fix it, the service for a connection
reason, and whether the reader may take it over or reconnect it. When
something can be taken over, sponsor_audience says who it would reach.
The state comes from the checks the run itself uses (ref_access,
resolve_holder_tool and the tool's connection as the run resolves it), so
the page and the run can't disagree. A run that leaves a resource out logs
resource_stopped with the holder, type, id and reason.
The workflow read gives sponsor details, run state and node resource names
only to people who may edit it, and names only resources the workflow runs,
someone sponsored, or the reader can see. Owner saves and new workflows
now refuse node tools and sources the owner can't use, like editor saves.
Someone who reaches an agent only through its public link was offered the
approval card for writes on the owner's connected accounts, so a stranger
could approve for the owner. Those writes are now refused with a tool
result, like an API-key caller's, unless the owner allowed the action in
the agent's Access details. Team members keep the card, and a tool on the
caller's own account (member mode) is unaffected. The flag survives a
resume, and workflow nodes now follow the run's caller rules (scheduled,
API-key and public-link) instead of starting from none.
A node's tools resolved as the person running the workflow, so a teammate
or public-link user lost every owner tool they could not use themselves.
They now resolve as the workflow owner, then as the editor who attached
them, like an agent's own tools. The runner stays the invoker, so a
member-mode connection still uses their own account.
chunks is a total per request, split across the attached sources, so an
agent with two sources and the default of 2 got a single chunk from each,
and its answers changed with whichever chunk won. 6 gives three per
source for about 4-5k more input tokens per retrieval turn.
Every literal default moves from 2 to 6: the request default, the
retrievers, the internal search tool, workflow agent nodes, scheduled and
headless runs, agent create/update/import, the source retrieval config and
the frontend forms. Existing agents and sources keep what they store; a
source saved with chunks=2 now counts as configured at 2, which is pinned
by a test.
Headless runs also read chunks=0 as unset (`or 2`), so an agent with
retrieval switched off retrieved anyway on scheduled runs; 0 now stays 0.
Paused, denied, skipped and client-run tool calls all go through
trace_unexecuted_tool_call. The research synthesis, continuation and
workflow step spans use with-blocks instead of hand-rolled GeneratorExit
handling. Fixes blank lines and adds a missing hint and docstrings.
Guardrail evaluations are traced when they call a remote check or fire; a
firing guardrail drops every content preview from the stored trace. Each
workflow node and each research phase is a step span, and the workflow run
id links the trace to its Logs row.
sum_tokens_in_range now skips the scheduler's per-run rollup rows, whose
tokens are already on the run's per-call rows, so the per-agent 24h limit
and the admin total no longer count scheduled spend twice.
Workflow node LLMs carry the workflow agent's id, so their usage rows are
attributed to the agent instead of landing with a user id only.
usage_totals returns a user's tokens and cost since a window start, split
by interactive and agent-key traffic.
About 85 call sites read a setting as getattr(settings, "NAME", fallback),
each carrying its own copy of the default. Every one of those names is a
field with a default on the model, so the fallback could never apply to
the real settings object; it only masked drift. Two had drifted:
- OPENAI_PROMPT_CACHE_KEY defaults to True on the model but the reader
fell back to False, and two test stubs relied on that.
- SharePoint's MICROSOFT_AUTHORITY fallback to
https://login.microsoftonline.com/<tenant> never fired, because the
attribute always exists (as None), so MSAL got authority=None. The
connector now derives the tenant authority when the setting is unset,
as its test always assumed.
Four places read EMBEDDINGS_KEY straight from os.environ, skipping the
"None"/"" normalisation the model applies; they read the setting now.
Test stubs that replaced a module's settings with a SimpleNamespace list
every setting the code under test reads.
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.
Kept for one release:
- A top-level application package whose meta-path finder resolves
application.x.y to the already-imported docsgpt.x.y object, so old imports
and entry points (celery -A application.app.celery,
uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
docsgpt.* task on start-up, so messages queued by the previous release still
run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
the previous release wrote are left unread instead of firing twice.
The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.