A turn whose agent raised wrote no user_logs row, so it only surfaced as
the agent's system error row. Every finished turn now writes its chat row,
at level error with the error when it failed, and linked to its trace; the
system row for the same traced activity is no longer listed twice.
The OTel replay and the trace INSERT ran in the stream's finally, so a slow
database held the SSE connection open after the last event. The trace is
still frozen when the stream ends, but written on a small writer pool.
A failed trace-summary lookup now leaves the Logs page intact without
chips. Tool-call counts include only calls that ran, not their paused,
denied or skipped records. A local guardrail that fires unchanged on every
streamed segment is recorded once, so it cannot use up the span cap.
build_agent no longer takes request_id from the request body: it becomes
the primary LLM's usage request id, and quotas count distinct request ids,
so a client could make every call count as one. Requests refused after
setup started (unauthorized, over quota, resume conflict, setup error) now
write their trace, marked error, through an after-request hook; streaming
routes hand the trace to complete_stream instead.
Tool results, denial comments and tool exception text now reach a span only
as a capture-gated preview; span.error is a fixed message, since it is
stored and exported regardless of content settings. A turn whose stream
yields an error event is recorded as failed. Search traces are listed by a
query copied into the small summary column, so the Logs timeline never
reads the spans JSONB. Adds GraphRAG span tests.
GET /api/traces returns a Logs row's stored traces, scoped to the caller or
to an agent they own. get_user_logs rows now carry the ids that find their
trace and a merged summary (duration, LLM and tool calls, tokens, retrieval
time) from one batched lookup per page; searches, MCP searches and graph
builds, which have no log row of their own, are listed from their traces.
StreamProcessor starts the trace and mints the request id before the
agent is built, so pre-fetch retrieval and compression are inside it and
side-channel LLM calls share the id. complete_stream activates it in the
SSE pump thread, binds the message and conversation, and writes it once
however the stream ends: paused, failed, abandoned or superseded (dropped).
user_logs rows now carry request_id and message_id.
One row per execution holding its span tree as JSONB, linked to Logs rows
by request, message, activity and workflow-run ids. message_id cascades so
deleting or truncating a conversation drops its traces; a daily beat task
enforces TRACES_RETENTION_DAYS.
`float("inf") > 0` is true, so a stored `"inf"` was preserved as if it
were a real count and reached the UI as "∞". Require a finite positive
value, which also covers `nan` explicitly rather than by accident.
Chunk cards in the source viewer read `metadata.token_count` and print a
bare "-" whenever it is absent. Ingestion records the count, but chunks
indexed before it was written -- and any path that rebuilds a chunk's
metadata from scratch -- reach the UI without it, so the whole file shows
"-" with no way to recover the number short of a re-ingest.
Fill the key in on the way out when it is missing or unusable (empty,
non-numeric, zero or negative), counting only the page being returned so
the cost is bounded by `per_page`. A count that is already recorded is
left untouched, including a numeric string, since stores round-trip
metadata types differently.
The fallback counts in cl100k rather than the embedding model's
tokenizer: loading that tokenizer can reach for a Hugging Face download,
which does not belong in a request path.
Correctness
- /api/remote never recorded source.created, so URL, GitHub and connector
sources had a source.deleted with no matching creation. All three creation
paths now go through one _audit_source_created helper.
- The prompt-cache rate divided cached tokens by a whole bucket's prompt
tokens. A bucket is a day and mixes calls whose provider reports a cache
breakdown with calls whose provider does not, so filtering buckets in the
client could not separate them and the rate was understated by however much
traffic ran on a non-reporting provider. The denominator is now computed in
SQL over the reporting rows.
- The outcome pill matched values nothing writes. Guardrails emit triggered /
not_evaluated and the device feed emits dispatched; the map had blocked /
denied / allowed, so a guardrail that fired rendered neutral grey -- the one
signal the merged feed exists to surface. Fixtures were seeding the
fictional values, so the tests passed on it too.
- Stream duration_ms timed the consumer. stream_token_usage is a generator,
so start-to-exhaustion includes the agent loop's tool handling and the SSE
client's pace; a slow browser recorded ~30s for a sub-second call. It now
accumulates only the time spent inside next().
Safety
- Activity filters failed open: an unknown facet or unparseable timestamp was
dropped, and no filter means every row, so a typo widened an audit view and
on the export streamed the full history. Both are now a 400.
- The search term was interpolated into an ILIKE pattern, so "100%" matched
everything and "q1_report" matched more than it should. Escaped.
- 0034 set actor_id NOT NULL with no default. A previous-release process
inserting mid-rollout would raise, and in admin/routes.py that insert shares
the request transaction, so a role grant beside it would roll back too.
Noise and dead code
- The per-user panel is a security panel: data-plane events file under the
actor, so an active account's routine deletes pushed a denied login out of
the 20-row window. It now excludes them; the Activity tab shows everything.
- device_audit_log had no created_at-leading index, so the merged feed
sequentially scanned that branch every page (migration 0036).
- conversation.deleted_all no longer records when nothing was deleted, and
agent.updated no longer records an empty field list.
- Dropped by_model from /admin/usage (no consumer; an extra aggregate per page
load), the duplicate filter surface on AuthEventsRepository that nothing
called, and the unreachable FLOW_LABELS.schedule entry.
- Type hints on record_event's conn and the remaining unannotated helpers.
- 0034's backfill set target_id to the acting admin for instance- and
team-scoped quota policy changes, which are filed under the actor and have
no user target. It now returns NULL for those, matching what quotas.py
records going forward, with a test covering all three scopes.
- The CSV export wrote every cell verbatim. user_agent is an attacker's raw
header, recorded without authenticating on a denied login, and the export
is opened in a spreadsheet by an admin -- a cell starting "=" would be
evaluated there. Every cell now has a leading formula trigger neutralized.
- The activity feed applied every response it received, so a slow reply for
an old filter could overwrite the current one. Guarded by a request id,
the same way settings/Analytics already does.
- The activity detail expander and the top-user drill-down were row onClick
handlers, unreachable without a mouse. Both are real buttons now, the
expander carrying aria-expanded and a label naming its event.
- FLOW_LABELS was applied to both breakdowns in the spend modal, so a model
named "fallback" or "workflow" rendered as a flow description. Only the
flow table maps keys now.
- Changing the range in that modal left the previous range's totals and
chart on screen while the new request was in flight.
Review pass over the preceding commits.
- The activity export paged by OFFSET. The journals are append-only and the
feed is newest-first, so rows written mid-export shift the window down and
repeat rows already emitted. Pages by keyset on (created_at, feed, id) now.
- top_token_users and the per-user breakdown counted run-level rollup rows.
A scheduled run already has a row per LLM call, so its spend was billed
twice -- invisible while the column was tokens, obvious once it was dollars.
Both now exclude them, matching every other spend query.
- total_cost was summed from the per-model split, which drops rows with no
model_id and so undercounted the figure printed above the series. Summed
from the series instead.
- The CSV export encoded its detail cell without the fallback the NDJSON
branch had, so a non-JSON-native value would have failed the stream.
- The activity view reset pagination in an effect, which fetched the stale
page against the new filters before fetching again. Reset in the setters.
- The audit taxonomy is no longer re-exported from the helper module; the one
caller that wanted it imports from where it lives.
Three gaps in one view.
Cost: token_usage.cost has been written on every call since quotas landed, and
tokens_by_model already returned it, but bucketed_totals and top_token_users
selected tokens only. An admin could set a USD quota in /admin/quotas and had
no way to see the spend it was capping. Buckets and top users now carry cost,
the endpoint reports a window total and a per-model split, and the chart takes
a Tokens/Cost toggle with a currency axis.
Group-by: the endpoint has supported group_by=model|agent|source from the
start and the UI only ever sent bucket=day. The selector is now wired, and a
grouped series pads missing buckets so a model that was idle on Tuesday plots
a zero instead of shifting its whole row one bar left.
Latency: surfaced from the columns the previous commit added, as p50/p95 with
a median time-to-first-token underneath. Unmeasured rows are excluded rather
than counted as zero, and the card says so when nothing was measured.
Also surfaces the prompt-cache hit rate, computed only over rows whose
provider reported a cache breakdown -- NULL means "not reported", and folding
those in as 0% would understate it.
Per-user drill-down: GET /api/admin/users/<id>/usage returns a daily series
plus splits by model and by flow, reachable from the top-users table and from
the user detail dialog. The detail dialog previously showed one tokens_30d
number, which answers neither "what is this person costing" nor "what is
driving it".
DocsGPT keeps three append-only journals -- auth_events, device_audit_log and
guardrail_events -- each readable only on its own terms, and the second of
those had an admin endpoint that no UI ever called. An operator reviewing an
incident wants one timeline, not three.
Adds GET /api/admin/activity: all three journals projected onto a common row
shape and merged, ordered newest first, with facets for journal, category,
event name, actor, affected user, a time window and free-text search over the
detail payload. A category filter that only one journal can satisfy drops the
others from the union rather than scanning and discarding them.
GET /api/admin/activity/events returns the distinct (event, category) pairs
the instance has actually recorded, so the filter UI can offer real choices
instead of asking an operator to type "oidc_login_denied" from memory.
GET /api/admin/activity/export streams the same filtered feed as CSV or
NDJSON. Streamed rather than buffered, and capped, so a compliance export of a
busy instance is bounded work.
guardrail_events.api_key (a raw agent key) and matched_value (unredacted
source text) are excluded from the projection, mirroring the exclusion the
per-agent guardrail view already makes. Admin-gating is not a reason to widen
what a list response carries.
Logins, role grants and provisioning were audited from the first release;
creating and deleting sources, agents, agent keys and conversations were not.
An operator reviewing the trail could see who signed in but not who deleted
the source they were asking about.
Adds docsgpt/api/audit.py: one helper that records an event inside the
caller's transaction, in a savepoint, swallowing failures -- an audit write
must never be able to fail the action it describes -- and tolerating the
absence of a Flask request so a Celery task can record too.
Events added: source.created (upload and wiki), source.deleted,
source.reingested, agent.created, agent.updated, agent.deleted,
agent.key_regenerated, conversation.deleted, conversation.deleted_all.
agent.updated records field names only, never values, which can carry prompts
and credentials.
The same module carries the event -> category map (identity / access / config
/ data) that the admin activity feed filters on.
Validation problems are returned as values rather than raised and echoed
with str(exc), and a huge integer limit is rejected as out of range instead
of overflowing. Tests no longer call mutating endpoints inside asserts.
The unpriced-model notice asked the live registry whether a model has a
price, so a priced model whose provider was later disabled showed up as
unpriced. It now lists models whose calls this period were all recorded at
$0. Token limits are serialized as integers.
Admins read and set the instance default, team allowances and user
overrides under /api/admin/quotas. A user's endpoint also returns the
limits those layers resolve to, the layer each came from and the usage
against them. The overview lists catalog models used this period that have
no price, since a cost limit cannot see them. Every write is audited.
GET /api/user/quota gives a user their own limited buckets, usage and reset
time without naming the policies behind them; any valid token may call it.
check_usage now checks the billable user's quota on every request, before
the per-agent 24h limits, which keep applying to traffic through an agent.
Until now a request without an agent key skipped every limit. A refusal is
a 429 with Retry-After and a body naming the budget, usage, limit, the
layer the limit came from and when it resets.
Headless runs check the agent owner's quota before starting. A refused
scheduled run is recorded as budget_exceeded; a refused webhook run returns
a quota_exceeded result instead of raising, so Celery does not retry it.
POST /api/user/tokens/<id>/regenerate swaps the secret of an existing token
in place: name, scopes and restrictions stay, the old secret stops matching
at once, and the expiry is reset. The lifetime defaults to the one the token
was last issued with (clamped to today's policy) or to expires_in_days when
given. An expired token can be renewed this way; a revoked one cannot. It is
session only like the rest of token management, and writes a pat_regenerated
audit event. regenerated_at records the rotation.
A resource restriction was checked on ids in the request, not on what the
addressed row pulls in or belongs to. Closed:
- Workflow writes for tokens restricted on sources, tools or prompts (a graph
names those inside its nodes), and attaching a workflow to an agent unless
the token is restricted on workflows too.
- Chat for tools-restricted tokens (chat executes tools; rejected at token
creation as well), and agent-less chat for tokens restricted on prompts or
workflows.
- conversation_id on chat: it must belong to the agent being run, or to no
agent for agent-less chat. Otherwise the server continued, appended to, or
resumed pending tool calls of another agent's conversation.
- Schedules for tokens restricted on anything but agents; schedule-id routes
for every restricted token.
- Conversations and analytics for every restricted token, not only
agent-restricted ones.
Also: create, first publish and adopt return the agent API key masked to a
token without agents:keys; token ids must be canonical UUIDs (urn:uuid: gave
a 500); an expired token is reported as expired; token creation takes a
per-user advisory lock so the cap cannot be raced; admin revoke-sessions
writes a pat_revoked event per token. The UI drops a row whose revoke returns
404 and does not offer a tools restriction next to chat:run.
The ASGI message events route now accepts the same scopes as its Flask
sibling (conversations:read or chat:run) through a shared constant. A token
request that fails routing gets Flask's 404/405 instead of a 403. Creating a
token retires an expired token that still held the name. allowed_ids uses
is_pat instead of a bare literal. PAT_ENABLED is documented as the master
switch it is: turning it off stops every existing token from authenticating.
The docs explain that sources are matched by name (oldest wins) and point CI
flows at sources upload --replace.
Message tail and the ASGI reconnect stream cannot tie a message to an
allowlist, so any token with a resource filter is refused there. The tools
listing now applies the allowlist to default and builtin rows as well. Token
creation answers 400 instead of 500 for a JSON body that is not an object.
A central rule table maps each route and method to the scope a token needs;
a route that is not listed cannot be called with a token, and a test fails
when a registered route is left unclassified. Token management, admin, team
management, sign-in, device pairing and OAuth handshakes are never token
reachable, and a token never carries the admin role.
A token restricted to specific agents, sources, prompts, tools or workflows
is held to its allowlist: ids are checked wherever a route carries them,
listings are filtered, creation is refused, and routes whose rows cannot be
tied to the allowlist are closed. Agent import checks the resolved target.
handle_auth resolves a dgpt_pat_ bearer against the database instead of
decoding it as a JWT, for both the Flask and the ASGI routes. Scopes and the
resource filter always come from the token row, and the claims that mark a
PAT are stripped from decoded JWTs so a session token cannot pose as one.
ASGI routes reject tokens unless they name the scope that admits them, and
/api/user/me reports what the calling token may do.
Returning a deferred marker traded one wrong signal for a worse one. A
redelivery reuses the original task id — Context.as_execution_options carries
task_id into the retry — so the duplicate's return marked the very id the
client polls as SUCCESS. /api/task_status reports celery's state verbatim and
the UI maps SUCCESS to "done", so the GraphRAG enable modal would announce a
finished build, rendered from a payload with no counts, while the run holding
the lease was still extracting.
Raise Ignore instead: celery records no state for the duplicate, so the task
id keeps whatever the holder sets and the poller keeps waiting. The autoretry
wrapper re-raises Ignore ahead of autoretry_for, so the wider
autoretry_for=(Exception,) on these tasks cannot turn it back into a retry.
A task that runs longer than the broker's visibility timeout is redelivered
while its first run is still going. The idempotency lease correctly stops the
duplicate from doing the work, but the duplicate then re-queued itself once
per LEASE_TTL until celery ran out of retries and raised
MaxRetriesExceededError — so a perfectly healthy long task (a large graph
extraction is the one that found this) reported a task failure, with a
traceback, while the real run was still making progress next to it.
Catch the exhaustion and return a "deferred" result instead. A normal
deferral still re-queues: only the give-up path changes, and the lease
holder's dedup row is left untouched so its own completion still records.
- Ship tiktoken's cl100k_base inside the package and build the encoding
from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
temp dir, and read tokenizer.json and repo metadata from that cache, so
a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
the endpoints return 404, audio files fail to ingest with a clear
message, /api/config reports tts_available/stt_available, and the UI
hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
The connector OAuth popup posted the session token to window.opener with
a '*' target origin, so any page that opened the popup received it. An
attacker with an account on a multi-user deployment could start a flow for
their own pending session, get a victim to finish the provider consent, and
receive a token backed by the victim's Drive/SharePoint/Confluence tokens.
- Post popup results only to allowed frontend origins: the callback origin,
OIDC_FRONTEND_URL, the new CONNECTOR_ALLOWED_ORIGINS, and localhost:5173
when the callback runs on a loopback host.
- Render the success page from the callback itself so the token never
appears in a URL; callback-status ignores session_token/user_email params.
- ConnectorAuth accepts messages only from the popup it opened, on the
callback origin reported by /api/connectors/auth.
- /api/connectors/disconnect requires auth and only deletes the caller's
session.
- /api/connectors/sync and /api/remote reject session tokens the caller
does not own.
Fixes#2766
- Per-user SSE cap uses per-connection leases in a sorted set instead of
a shared INCR/DECR counter with a TTL. A stream that outlived the TTL
could let the counter expire, then decrement another stream's slot or
drive the count negative. Leases refresh while a stream sends frames
and age out when a stream dies without cleanup.
- Shielded cleanup is bounded per step: stream close and on_close in
ClosingStreamingResponse, unsubscribe and close in AsyncTopic, lease
release, and the replay-budget check, so a dead Redis connection can't
hold a request or a graceful shutdown.
- Artifact downloads send an ASCII filename plus an RFC 5987 filename*
for non-ASCII names, build headers before opening the file, and close
disk handles off the event loop.
- Type hints and docstrings on the new helpers.
Move GET /api/events, the remote-device command stream and artifact
downloads from Flask to Starlette routes mounted ahead of the Flask
catch-all. On Flask each held an a2wsgi threadpool slot for as long as
its response stayed open, and because uvicorn drops writes after a
client disconnects, a closed tab never released it.
- asgi_auth: one JWT/OIDC gate for Starlette routes; the chat reconnect
reader uses it too
- ClosingStreamingResponse closes the body iterator and releases the
SSE slot or file handle even when the client leaves before the first
frame
- AsyncTopic liveness probe replaces the sync client's socket_timeout
guard against half-open pub/sub sockets
- ASYNC_REDIS_MAX_CONNECTIONS sizes the async Redis pool; every open
stream holds a connection and redis-py defaults to 100
PinAgent looked agents up with no ownership, share or team predicate, and
PinnedAgents returned those rows through a hand-rolled dict that bypassed
the blanking _format_agent_output applies. Anyone holding an agent id --
a share-link recipient learns the real UUID -- could pin it and keep
reading its name, description, prompt, tools, type, status and masked key
after the share was revoked.
Add _user_may_pin: the owner, a team grantee, a recipient of a still-live
share link, or a system template. PinAgent checks it before pinning and
PinnedAgents re-checks every row on read, so a revoked grant drops out of
the list instead of persisting. Unpinning stays open regardless of current
access, or a revoked share would strand a pin the user cannot clear. The
masked key is now owner-only, matching _format_agent_output.
Adds the first cross-user pin tests: every existing one pinned the
caller's own agent.
Review follow-ups:
- The agents list and the pinned list dropped agents that had neither a
source nor a retriever. Publishing such an agent is now allowed, so it
vanished from the list right after publish. Both filters are gone; a
source-less agent skips retrieval and is still runnable. Tests cover
both listings.
- The agent form selected an uploaded source by diffing the source
catalog before and after the upload, which would also pick up any
source created or shared in the meantime. Upload now passes the id of
the source it created to onSuccessfulUpload, and the form selects only
that id.
The sources list used to start with a fake "Default" entry that had no id
and, at run time, meant "no source, skip retrieval". The agent form
pre-selected it, snapped back to it when the last source was deselected,
and refused to publish without it, so a new agent always looked like it
had a knowledge base when it had none.
Backend
- /api/sources returns only ingested sources; no placeholder row.
- Publishing an agent no longer requires a source on create or update.
The legacy "default" value is still accepted and maps to NULL.
Frontend
- The agent source picker starts empty, can be cleared, and shows a hint
that a source-less agent answers from the model and its tools only.
- The picker groups sources into "Your sources" and "Shared with team"
when any team-shared source exists, shows "N sources selected" for a
multi-selection, and gets the same "Go to Sources" / "Upload new"
footer as the chat picker. A source uploaded from the form is selected
when it lands.
- Source selection serialisation and the picker id live in one helper
shared with the chat picker; the four copies in the form are gone.
- The client no longer seeds a placeholder source in the store, and the
dead auto-select of a "default" document is removed.
- CI installs the backend requirements from docsgpt/; the old cd into
application/ silently installed nothing.
- The root .dockerignore re-admits only application/__init__.py. An upgraded
checkout may still hold gitignored application/{inputs,indexes,vectors,.env}
from the old layout, and the directory rule shipped them into the image.
- The compose files keep the host bind mounts on application/{indexes,inputs,
vectors}, so an upgrade does not start with empty data. The move comes with
the packaging work, together with an upgrade note.
- The alias loader puts the real docsgpt spec back on the shared module object
after import (the import machinery stamped the alias spec on it, which made
importlib.reload rename the module and skip re-execution) and delegates
get_code/get_source/get_filename to the target loader, so
python -m application.<name> runs.
- Each legacy application.* task name is registered as its own task object,
a subclass carrying the old name. Registering the same object under two
keys made Celery's tracer log every run under whichever name it built last.
- The redbeat key prefix stays redbeat:docsgpt:; the three schedule_syncs
entries get stable names instead. redbeat tracks its static entries and
deletes the ones that vanish from beat_schedule at start-up, and rewrites the
task path of named entries in place, so neither a prefix bump nor a cleanup
pass is needed (checked against redbeat 2.4.2 with a seeded Redis).
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.
Kept for one release:
- A top-level application package whose meta-path finder resolves
application.x.y to the already-imported docsgpt.x.y object, so old imports
and entry points (celery -A application.app.celery,
uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
docsgpt.* task on start-up, so messages queued by the previous release still
run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
the previous release wrote are left unread instead of firing twice.
The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
An unchained request with no system message staged None, and the record
step skipped the commit, so the previous head hash survived a transcript
that never received a head; a later chained request restoring that head
would have omitted it. The staged value is now committed as-is, None
included. Also pins each rejection predicate of is_usable_compression_point
with its own test.
Three review findings on the bounded-chain change.
Mid-execution compression rebuilt the conversation from the in-flight
messages, which after a turn-start reuse hold only the recent turns: the
summary living in the system prompt never reached the compressor, so the
new summary replaced the old one, and the persisted point's query_index
was relative to that shortened list. The summary the agent is running
under now rides into the synthetic conversation as its latest point
(query_index -1, so every in-flight query is new), for both the database
and the in-memory path, and the database path persists the index of the
saved conversation's last row.
Saved points with an empty summary, which earlier versions wrote, were
treated as reusable: get_compressed_context sliced the raw history away
and the effective token count made the conversation look small. Point
selection everywhere now takes the latest usable point (non-blank
summary, positive token count) and falls back to the raw history when
there is none.
The chained system-head hash was committed while building the request,
so a transport failure followed by the same-primary retry omitted a
changed system message. The hash is now staged per request and committed
only when the provider records the response.
- append_compression_point only skips a point when both query_index and
compressed_summary are present and match the last one; points without
those fields (as in the repository tests) were all being treated as
duplicates.
- The incremental compression tail and the orchestrator's "anything new
since the last point" check exclude the visible summary row, which the
prompt already receives through existing_compressions.
- Summary rows carry a persisted metadata marker; replay filters on the
marker, and falls back to the label only for rows written before it that
have no tool calls and no per-turn metadata, so a user who types the
label text keeps their turn.
- The prompt_cache_key is a hash of the user id, never the id itself.
- Describe truncation="auto" as dropping the oldest items.
In store mode every user turn chained onto the previous response, so the
provider's stored transcript grew without bound (measured: 889k prompt
tokens for a 37k-token saved history) while every local guard, the
compression pipeline included, measured the saved history. Each chained
tool round also re-sent the system message, which the server appends rather
than dedupes, and a saved compression point was applied exactly once, in the
turn that made it.
Chaining is now bounded. A turn starts from the saved history when the
previous turn's reported prompt reached the chain budget (default: the
model's context window), when the conversation was compressed after that
turn was produced, or when OPENAI_RESPONSES_CHAIN_ACROSS_TURNS is off.
Chained rounds omit an unchanged system head (hash carried in the persisted
Responses state). truncation="auto" is available behind a setting as a
backstop against a chain that outgrows the model's window.
Compression: a saved point is applied at every turn start; the threshold
counts the summary plus the queries after the point instead of the raw
history; re-compression summarises only the tail on top of the last point;
the mid-execution path marks itself persisted and resets the provider chain
so the rebuilt messages are the context; an empty summary is rejected; the
visible "[Context Compression Summary]" rows are no longer replayed as
history; appending the same point twice is a no-op.
Cache hints: a per-user prompt_cache_key and an optional
prompt_cache_retention on Responses API calls.
Measured on Azure with the same client shape as production (stateless
OpenAI client, server-side tools, PDF part): tokens billed on the sixth turn
fell from 58k to 35k, tool rounds add tens of tokens instead of ~2.8k, the
turn after a compression reused the saved summary in under two seconds
instead of re-summarising, and the round after a mid-execution compression
started from the compressed context instead of the full stored transcript.