An agent called with its API key (widget, API) runs as its owner, and
nobody can approve an action there, so a write set to Always allow ran
on the owner's account for anyone holding the key. A caller is external
when the request carries the key and is not signed in as the owner (an
owner previewing their agent keeps full access).
For external callers, write actions on connection-backed tools are
refused unless listed in the agent config's new api_write_allowlist
(tool_id:action), and a missing connection is refused instead of pausing
on a Connect card the widget cannot show. The flag and list survive a
paused-and-resumed stream. Owners pick the allowed actions under Access
details; a team editor's update keeps the stored list.
chunks is a total per request, split across the attached sources, so an
agent with two sources and the default of 2 got a single chunk from each,
and its answers changed with whichever chunk won. 6 gives three per
source for about 4-5k more input tokens per retrieval turn.
Every literal default moves from 2 to 6: the request default, the
retrievers, the internal search tool, workflow agent nodes, scheduled and
headless runs, agent create/update/import, the source retrieval config and
the frontend forms. Existing agents and sources keep what they store; a
source saved with chunks=2 now counts as configured at 2, which is pinned
by a test.
Headless runs also read chunks=0 as unset (`or 2`), so an agent with
retrieval switched off retrieved anyway on scheduled runs; 0 now stays 0.
A turn whose agent raised wrote no user_logs row, so it only surfaced as
the agent's system error row. Every finished turn now writes its chat row,
at level error with the error when it failed, and linked to its trace; the
system row for the same traced activity is no longer listed twice.
The OTel replay and the trace INSERT ran in the stream's finally, so a slow
database held the SSE connection open after the last event. The trace is
still frozen when the stream ends, but written on a small writer pool.
build_agent no longer takes request_id from the request body: it becomes
the primary LLM's usage request id, and quotas count distinct request ids,
so a client could make every call count as one. Requests refused after
setup started (unauthorized, over quota, resume conflict, setup error) now
write their trace, marked error, through an after-request hook; streaming
routes hand the trace to complete_stream instead.
Tool results, denial comments and tool exception text now reach a span only
as a capture-gated preview; span.error is a fixed message, since it is
stored and exported regardless of content settings. A turn whose stream
yields an error event is recorded as failed. Search traces are listed by a
query copied into the small summary column, so the Logs timeline never
reads the spans JSONB. Adds GraphRAG span tests.
The trace lifecycle moves into a decorator so complete_stream keeps its
real parameters for introspection. GET /api/traces is readable with the
analytics:read scope, like the Logs endpoint it complements.
StreamProcessor starts the trace and mints the request id before the
agent is built, so pre-fetch retrieval and compression are inside it and
side-channel LLM calls share the id. complete_stream activates it in the
SSE pump thread, binds the message and conversation, and writes it once
however the stream ends: paused, failed, abandoned or superseded (dropped).
user_logs rows now carry request_id and message_id.
- check_usage treats a request through a keyless (draft) agent as agent
traffic, matching how its usage rows are bucketed and the headless rule
- dashboard edits carry the stored enabled flag instead of re-enabling the
policy; disabled policies are labelled in the Quotas tab and the editor
- quota 429s send x-should-retry: false so OpenAI SDK clients do not retry
a refusal that cannot succeed before the reset
- cached-input and cache-write rates for Anthropic, OpenRouter and Groq
gpt-oss-120b; refresh OpenRouter deepseek-v3.2 list prices
- UsageQuota reuses usagePercent; docs note that a user override needs an
existing user
- A tool continuation refused for usage now releases the resume claim it
took; before, retries got a 409 until the stale claim was reverted.
- Agent traffic is any row with an agent key or an agent id, so keyless
agents and workflow nodes count toward the agent bucket, not direct.
- The user quota modal discards responses for a previously opened user.
- The usage meter shows every limited bucket, not only 'all'.
- Restore the class separator on the analytics stat card that a formatter
run removed, and align the OpenRouter DeepSeek description with its rates.
check_usage now checks the billable user's quota on every request, before
the per-agent 24h limits, which keep applying to traffic through an agent.
Until now a request without an agent key skipped every limit. A refusal is
a 429 with Retry-After and a body naming the budget, usage, limit, the
layer the limit came from and when it resets.
Headless runs check the agent owner's quota before starting. A refused
scheduled run is recorded as budget_exceeded; a refused webhook run returns
a quota_exceeded result instead of raising, so Celery does not retry it.
About 85 call sites read a setting as getattr(settings, "NAME", fallback),
each carrying its own copy of the default. Every one of those names is a
field with a default on the model, so the fallback could never apply to
the real settings object; it only masked drift. Two had drifted:
- OPENAI_PROMPT_CACHE_KEY defaults to True on the model but the reader
fell back to False, and two test stubs relied on that.
- SharePoint's MICROSOFT_AUTHORITY fallback to
https://login.microsoftonline.com/<tenant> never fired, because the
attribute always exists (as None), so MSAL got authority=None. The
connector now derives the tenant authority when the setting is unset,
as its test always assumed.
Four places read EMBEDDINGS_KEY straight from os.environ, skipping the
"None"/"" normalisation the model applies; they read the setting now.
Test stubs that replaced a module's settings with a SimpleNamespace list
every setting the code under test reads.
The sources list used to start with a fake "Default" entry that had no id
and, at run time, meant "no source, skip retrieval". The agent form
pre-selected it, snapped back to it when the last source was deselected,
and refused to publish without it, so a new agent always looked like it
had a knowledge base when it had none.
Backend
- /api/sources returns only ingested sources; no placeholder row.
- Publishing an agent no longer requires a source on create or update.
The legacy "default" value is still accepted and maps to NULL.
Frontend
- The agent source picker starts empty, can be cleared, and shows a hint
that a source-less agent answers from the model and its tools only.
- The picker groups sources into "Your sources" and "Shared with team"
when any team-shared source exists, shows "N sources selected" for a
multi-selection, and gets the same "Go to Sources" / "Upload new"
footer as the chat picker. A source uploaded from the form is selected
when it lands.
- Source selection serialisation and the picker id live in one helper
shared with the chat picker; the four copies in the form are gone.
- The client no longer seeds a placeholder source in the store, and the
dead auto-select of a "default" document is removed.
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.
Kept for one release:
- A top-level application package whose meta-path finder resolves
application.x.y to the already-imported docsgpt.x.y object, so old imports
and entry points (celery -A application.app.celery,
uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
docsgpt.* task on start-up, so messages queued by the previous release still
run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
the previous release wrote are left unread instead of firing twice.
The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.