The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.
Kept for one release:
- A top-level application package whose meta-path finder resolves
application.x.y to the already-imported docsgpt.x.y object, so old imports
and entry points (celery -A application.app.celery,
uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
docsgpt.* task on start-up, so messages queued by the previous release still
run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
the previous release wrote are left unread instead of firing twice.
The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
An unchained request with no system message staged None, and the record
step skipped the commit, so the previous head hash survived a transcript
that never received a head; a later chained request restoring that head
would have omitted it. The staged value is now committed as-is, None
included. Also pins each rejection predicate of is_usable_compression_point
with its own test.
Three review findings on the bounded-chain change.
Mid-execution compression rebuilt the conversation from the in-flight
messages, which after a turn-start reuse hold only the recent turns: the
summary living in the system prompt never reached the compressor, so the
new summary replaced the old one, and the persisted point's query_index
was relative to that shortened list. The summary the agent is running
under now rides into the synthetic conversation as its latest point
(query_index -1, so every in-flight query is new), for both the database
and the in-memory path, and the database path persists the index of the
saved conversation's last row.
Saved points with an empty summary, which earlier versions wrote, were
treated as reusable: get_compressed_context sliced the raw history away
and the effective token count made the conversation look small. Point
selection everywhere now takes the latest usable point (non-blank
summary, positive token count) and falls back to the raw history when
there is none.
The chained system-head hash was committed while building the request,
so a transport failure followed by the same-primary retry omitted a
changed system message. The hash is now staged per request and committed
only when the provider records the response.
- append_compression_point only skips a point when both query_index and
compressed_summary are present and match the last one; points without
those fields (as in the repository tests) were all being treated as
duplicates.
- The incremental compression tail and the orchestrator's "anything new
since the last point" check exclude the visible summary row, which the
prompt already receives through existing_compressions.
- Summary rows carry a persisted metadata marker; replay filters on the
marker, and falls back to the label only for rows written before it that
have no tool calls and no per-turn metadata, so a user who types the
label text keeps their turn.
- The prompt_cache_key is a hash of the user id, never the id itself.
- Describe truncation="auto" as dropping the oldest items.
In store mode every user turn chained onto the previous response, so the
provider's stored transcript grew without bound (measured: 889k prompt
tokens for a 37k-token saved history) while every local guard, the
compression pipeline included, measured the saved history. Each chained
tool round also re-sent the system message, which the server appends rather
than dedupes, and a saved compression point was applied exactly once, in the
turn that made it.
Chaining is now bounded. A turn starts from the saved history when the
previous turn's reported prompt reached the chain budget (default: the
model's context window), when the conversation was compressed after that
turn was produced, or when OPENAI_RESPONSES_CHAIN_ACROSS_TURNS is off.
Chained rounds omit an unchanged system head (hash carried in the persisted
Responses state). truncation="auto" is available behind a setting as a
backstop against a chain that outgrows the model's window.
Compression: a saved point is applied at every turn start; the threshold
counts the summary plus the queries after the point instead of the raw
history; re-compression summarises only the tail on top of the last point;
the mid-execution path marks itself persisted and resets the provider chain
so the rebuilt messages are the context; an empty summary is rejected; the
visible "[Context Compression Summary]" rows are no longer replayed as
history; appending the same point twice is a no-op.
Cache hints: a per-user prompt_cache_key and an optional
prompt_cache_retention on Responses API calls.
Measured on Azure with the same client shape as production (stateless
OpenAI client, server-side tools, PDF part): tokens billed on the sixth turn
fell from 58k to 35k, tool rounds add tens of tokens instead of ~2.8k, the
turn after a compression reused the saved summary in under two seconds
instead of re-summarising, and the round after a mid-execution compression
started from the compressed context instead of the full stored transcript.
Token accounting:
- Drain each tool round's provider stream to exhaustion before running
tools and recursing, so the usage decorator persists exactly one
token_usage row per LLM call, at call end. Previously every round's
generator was abandoned mid-iteration and flushed together at request
teardown, writing N near-identical rows stamped with the final
round's provider counts (duplicate billing).
- Consume the Chat Completions include_usage terminal chunk (it arrives
after finish_reason and was never read) so streamed calls record
provider-exact token counts instead of tiktoken estimates.
- Claim provider-reported usage once per call (_last_usage_claimed) so
a late-finalized generator can never adopt another call's counts.
Oversized-context guards:
- Enforce Responses API function_call/function_call_output pairing in
the input builder (drop unpaired items; bypassed for store-mode
previous_response_id chaining where calls are matched server-side).
- Hard pre-send context gate: shrink oversized tool results and refuse
payloads that cannot fit the model's window before dispatch, so a
hopeless request is never sent or billed.
- Cap a single tool result entering the LLM context
(TOOL_RESULT_MAX_TOKENS, default 20000); the tool journal and
persistence keep the full result. Applied on the resume/continuation
path too.
- Skip the fallback attempt when the payload cannot fit the fallback
model's context window (10% estimation slack).
- Compression: never save a compression point that does not reduce
tokens; bound oversized verbatim fields kept after a compression
point (COMPRESSION_RECENT_FIELD_MAX_TOKENS, default 8000).
Robustness fixes from review:
- Google parallel function calls: complete index-less ToolCalls are no
longer merged into one another (dict arguments raised TypeError on
+=; second call could execute with the first call's arguments).
- Trailing-frame failures after a delivered answer no longer error the
stream or restream the whole answer from the fallback
(_stream_reached_finish).
- In-memory compression falls back to minimal pruning when the summary
is not smaller than the original.
- keep<=0 guard in the middle-truncation helpers (a tiny cap returned
marker + full text).
Frontend: tooltip on the Analytics tokens stat card explaining that
agent tool loops re-send conversation context on every step.