10 Commits
Author SHA1 Message Date
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00
Alex 427d85d737 fix(llm): commit the staged head hash unconditionally on a recorded response
An unchained request with no system message staged None, and the record
step skipped the commit, so the previous head hash survived a transcript
that never received a head; a later chained request restoring that head
would have omitted it. The staged value is now committed as-is, None
included. Also pins each rejection predicate of is_usable_compression_point
with its own test.
2026-09-05 11:26:25 +01:00
Alex dbee30a048 fix(compression,llm): keep the summary across mid-execution compression, ignore empty saved points, commit the head hash on success
Three review findings on the bounded-chain change.

Mid-execution compression rebuilt the conversation from the in-flight
messages, which after a turn-start reuse hold only the recent turns: the
summary living in the system prompt never reached the compressor, so the
new summary replaced the old one, and the persisted point's query_index
was relative to that shortened list. The summary the agent is running
under now rides into the synthetic conversation as its latest point
(query_index -1, so every in-flight query is new), for both the database
and the in-memory path, and the database path persists the index of the
saved conversation's last row.

Saved points with an empty summary, which earlier versions wrote, were
treated as reusable: get_compressed_context sliced the raw history away
and the effective token count made the conversation look small. Point
selection everywhere now takes the latest usable point (non-blank
summary, positive token count) and falls back to the raw history when
there is none.

The chained system-head hash was committed while building the request,
so a transport failure followed by the same-primary retry omitted a
changed system message. The hash is now staged per request and committed
only when the provider records the response.
2026-09-05 10:39:58 +01:00
Alex 1f86139b9c fix(compression,llm): address review — exact point dedupe, marked summary rows, opaque cache key
- append_compression_point only skips a point when both query_index and
  compressed_summary are present and match the last one; points without
  those fields (as in the repository tests) were all being treated as
  duplicates.
- The incremental compression tail and the orchestrator's "anything new
  since the last point" check exclude the visible summary row, which the
  prompt already receives through existing_compressions.
- Summary rows carry a persisted metadata marker; replay filters on the
  marker, and falls back to the label only for rows written before it that
  have no tool calls and no per-turn metadata, so a user who types the
  label text keeps their turn.
- The prompt_cache_key is a hash of the user id, never the id itself.
- Describe truncation="auto" as dropping the oldest items.
2026-09-05 09:30:13 +01:00
Alex 04d358ab69 fix(llm,compression): bound cross-turn Responses chaining and make compression stick
In store mode every user turn chained onto the previous response, so the
provider's stored transcript grew without bound (measured: 889k prompt
tokens for a 37k-token saved history) while every local guard, the
compression pipeline included, measured the saved history. Each chained
tool round also re-sent the system message, which the server appends rather
than dedupes, and a saved compression point was applied exactly once, in the
turn that made it.

Chaining is now bounded. A turn starts from the saved history when the
previous turn's reported prompt reached the chain budget (default: the
model's context window), when the conversation was compressed after that
turn was produced, or when OPENAI_RESPONSES_CHAIN_ACROSS_TURNS is off.
Chained rounds omit an unchanged system head (hash carried in the persisted
Responses state). truncation="auto" is available behind a setting as a
backstop against a chain that outgrows the model's window.

Compression: a saved point is applied at every turn start; the threshold
counts the summary plus the queries after the point instead of the raw
history; re-compression summarises only the tail on top of the last point;
the mid-execution path marks itself persisted and resets the provider chain
so the rebuilt messages are the context; an empty summary is rejected; the
visible "[Context Compression Summary]" rows are no longer replayed as
history; appending the same point twice is a no-op.

Cache hints: a per-user prompt_cache_key and an optional
prompt_cache_retention on Responses API calls.

Measured on Azure with the same client shape as production (stateless
OpenAI client, server-side tools, PDF part): tokens billed on the sixth turn
fell from 58k to 35k, tool rounds add tens of tokens instead of ~2.8k, the
turn after a compression reused the saved summary in under two seconds
instead of re-summarising, and the round after a mid-execution compression
started from the compressed context instead of the full stored transcript.
2026-09-05 00:28:51 +01:00
Alex ca5f80995d fix: accurate per-call token usage and oversized-context guards
Token accounting:
- Drain each tool round's provider stream to exhaustion before running
  tools and recursing, so the usage decorator persists exactly one
  token_usage row per LLM call, at call end. Previously every round's
  generator was abandoned mid-iteration and flushed together at request
  teardown, writing N near-identical rows stamped with the final
  round's provider counts (duplicate billing).
- Consume the Chat Completions include_usage terminal chunk (it arrives
  after finish_reason and was never read) so streamed calls record
  provider-exact token counts instead of tiktoken estimates.
- Claim provider-reported usage once per call (_last_usage_claimed) so
  a late-finalized generator can never adopt another call's counts.

Oversized-context guards:
- Enforce Responses API function_call/function_call_output pairing in
  the input builder (drop unpaired items; bypassed for store-mode
  previous_response_id chaining where calls are matched server-side).
- Hard pre-send context gate: shrink oversized tool results and refuse
  payloads that cannot fit the model's window before dispatch, so a
  hopeless request is never sent or billed.
- Cap a single tool result entering the LLM context
  (TOOL_RESULT_MAX_TOKENS, default 20000); the tool journal and
  persistence keep the full result. Applied on the resume/continuation
  path too.
- Skip the fallback attempt when the payload cannot fit the fallback
  model's context window (10% estimation slack).
- Compression: never save a compression point that does not reduce
  tokens; bound oversized verbatim fields kept after a compression
  point (COMPRESSION_RECENT_FIELD_MAX_TOKENS, default 8000).

Robustness fixes from review:
- Google parallel function calls: complete index-less ToolCalls are no
  longer merged into one another (dict arguments raised TypeError on
  +=; second call could execute with the first call's arguments).
- Trailing-frame failures after a delivered answer no longer error the
  stream or restream the whole answer from the fallback
  (_stream_reached_finish).
- In-memory compression falls back to minimal pruning when the summary
  is not smaller than the original.
- keep<=0 guard in the middle-truncation helpers (a tiny cap returned
  marker + full text).

Frontend: tooltip on the Analytics tokens stat card explaining that
agent tool loops re-send conversation context on every step.
2026-07-15 21:20:39 +01:00
Alex 318de18d43 feat: BYOM (#2433) 2026-04-27 22:09:33 +01:00
Alex 73256389cf feat: client side tools 2026-03-31 22:20:55 +01:00
Alex d5c0322e2a chore: more tests 2026-03-30 16:13:08 +01:00
Alex fe185e5b8d chore: api and tool tests 2026-03-28 21:51:47 +00:00