Commit Graph
146 Commits
Author SHA1 Message Date
Alex 02f2fa198a fix: disabled STT also covers live finish; clarify model cache location 2026-09-15 19:28:19 +01:00
Alex 7da46c2bea feat: air-gapped deployment guide, no implicit downloads
- Ship tiktoken's cl100k_base inside the package and build the encoding
  from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
  temp dir, and read tokenizer.json and repo metadata from that cache, so
  a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
  the endpoints return 404, audio files fail to ingest with a clear
  message, /api/config reports tts_available/stt_available, and the UI
  hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
  packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
2026-09-15 17:54:24 +01:00
Alex 3dfe3b6112 fix: slight hardening 2026-09-15 08:44:37 +01:00
Alex e8305ef137 fix: more mini connector hardening 2026-09-14 22:49:34 +01:00
Alex 08e8de7370 fix: mini connector fixes 2026-09-14 22:29:26 +01:00
Alex 41b3afed14 fix: stop leaking connector OAuth session tokens to other origins
The connector OAuth popup posted the session token to window.opener with
a '*' target origin, so any page that opened the popup received it. An
attacker with an account on a multi-user deployment could start a flow for
their own pending session, get a victim to finish the provider consent, and
receive a token backed by the victim's Drive/SharePoint/Confluence tokens.

- Post popup results only to allowed frontend origins: the callback origin,
  OIDC_FRONTEND_URL, the new CONNECTOR_ALLOWED_ORIGINS, and localhost:5173
  when the callback runs on a loopback host.
- Render the success page from the callback itself so the token never
  appears in a URL; callback-status ignores session_token/user_email params.
- ConnectorAuth accepts messages only from the popup it opened, on the
  callback origin reported by /api/connectors/auth.
- /api/connectors/disconnect requires auth and only deletes the caller's
  session.
- /api/connectors/sync and /api/remote reject session tokens the caller
  does not own.

Fixes #2766
2026-09-14 21:46:09 +01:00
Alex fb66a0b230 fix: stability improvements on event loop 2026-09-14 12:52:15 +01:00
Alex e889392eba fix(api): address review on the event-loop streams
- Per-user SSE cap uses per-connection leases in a sorted set instead of
  a shared INCR/DECR counter with a TTL. A stream that outlived the TTL
  could let the counter expire, then decrement another stream's slot or
  drive the count negative. Leases refresh while a stream sends frames
  and age out when a stream dies without cleanup.
- Shielded cleanup is bounded per step: stream close and on_close in
  ClosingStreamingResponse, unsubscribe and close in AsyncTopic, lease
  release, and the replay-budget check, so a dead Redis connection can't
  hold a request or a graceful shutdown.
- Artifact downloads send an ASCII filename plus an RFC 5987 filename*
  for non-ASCII names, build headers before opening the file, and close
  disk handles off the event loop.
- Type hints and docstrings on the new helpers.
2026-09-14 10:19:59 +01:00
Alex 54b540ea42 feat(api): serve long-lived streams on the event loop
Move GET /api/events, the remote-device command stream and artifact
downloads from Flask to Starlette routes mounted ahead of the Flask
catch-all. On Flask each held an a2wsgi threadpool slot for as long as
its response stayed open, and because uvicorn drops writes after a
client disconnects, a closed tab never released it.

- asgi_auth: one JWT/OIDC gate for Starlette routes; the chat reconnect
  reader uses it too
- ClosingStreamingResponse closes the body iterator and releases the
  SSE slot or file handle even when the client leaves before the first
  frame
- AsyncTopic liveness probe replaces the sync client's socket_timeout
  guard against half-open pub/sub sockets
- ASYNC_REDIS_MAX_CONNECTIONS sizes the async Redis pool; every open
  stream holds a connection and redis-py defaults to 100
2026-09-14 08:33:46 +01:00
Alex c26b8a3403 fix(agents): gate agent pinning on visibility
PinAgent looked agents up with no ownership, share or team predicate, and
PinnedAgents returned those rows through a hand-rolled dict that bypassed
the blanking _format_agent_output applies. Anyone holding an agent id --
a share-link recipient learns the real UUID -- could pin it and keep
reading its name, description, prompt, tools, type, status and masked key
after the share was revoked.

Add _user_may_pin: the owner, a team grantee, a recipient of a still-live
share link, or a system template. PinAgent checks it before pinning and
PinnedAgents re-checks every row on read, so a revoked grant drops out of
the list instead of persisting. Unpinning stays open regardless of current
access, or a revoked share would strand a pin the user cannot clear. The
masked key is now owner-only, matching _format_agent_output.

Adds the first cross-user pin tests: every existing one pinned the
caller's own agent.
2026-09-11 11:28:14 +01:00
Alex 0fe059db12 fix(agents): list source-less agents and select only the uploaded source
Review follow-ups:

- The agents list and the pinned list dropped agents that had neither a
  source nor a retriever. Publishing such an agent is now allowed, so it
  vanished from the list right after publish. Both filters are gone; a
  source-less agent skips retrieval and is still runnable. Tests cover
  both listings.
- The agent form selected an uploaded source by diffing the source
  catalog before and after the upload, which would also pick up any
  source created or shared in the meantime. Upload now passes the id of
  the source it created to onSuccessfulUpload, and the form selects only
  that id.
2026-09-09 18:39:09 +01:00
Alex 79d418f32a feat(agents): make sources optional and drop the synthetic "Default" source
The sources list used to start with a fake "Default" entry that had no id
and, at run time, meant "no source, skip retrieval". The agent form
pre-selected it, snapped back to it when the last source was deselected,
and refused to publish without it, so a new agent always looked like it
had a knowledge base when it had none.

Backend
- /api/sources returns only ingested sources; no placeholder row.
- Publishing an agent no longer requires a source on create or update.
  The legacy "default" value is still accepted and maps to NULL.

Frontend
- The agent source picker starts empty, can be cleared, and shows a hint
  that a source-less agent answers from the model and its tools only.
- The picker groups sources into "Your sources" and "Shared with team"
  when any team-shared source exists, shows "N sources selected" for a
  multi-selection, and gets the same "Go to Sources" / "Upload new"
  footer as the chat picker. A source uploaded from the form is selected
  when it lands.
- Source selection serialisation and the picker id live in one helper
  shared with the chat picker; the four copies in the form are gone.
- The client no longer seeds a placeholder source in the store, and the
  dead auto-select of a "default" document is removed.
2026-09-09 17:59:08 +01:00
Alex adb6963523 fix(rename): address review on the package rename
- CI installs the backend requirements from docsgpt/; the old cd into
  application/ silently installed nothing.
- The root .dockerignore re-admits only application/__init__.py. An upgraded
  checkout may still hold gitignored application/{inputs,indexes,vectors,.env}
  from the old layout, and the directory rule shipped them into the image.
- The compose files keep the host bind mounts on application/{indexes,inputs,
  vectors}, so an upgrade does not start with empty data. The move comes with
  the packaging work, together with an upgrade note.
- The alias loader puts the real docsgpt spec back on the shared module object
  after import (the import machinery stamped the alias spec on it, which made
  importlib.reload rename the module and skip re-execution) and delegates
  get_code/get_source/get_filename to the target loader, so
  python -m application.<name> runs.
- Each legacy application.* task name is registered as its own task object,
  a subclass carrying the old name. Registering the same object under two
  keys made Celery's tracer log every run under whichever name it built last.
- The redbeat key prefix stays redbeat:docsgpt:; the three schedule_syncs
  entries get stable names instead. redbeat tracks its static entries and
  deletes the ones that vanish from beat_schedule at start-up, and rewrites the
  task path of named entries in place, so neither a prefix bump nor a cleanup
  pass is needed (checked against redbeat 2.4.2 with a seeded Redis).
2026-09-07 12:02:07 +01:00
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00
Alex 427d85d737 fix(llm): commit the staged head hash unconditionally on a recorded response
An unchained request with no system message staged None, and the record
step skipped the commit, so the previous head hash survived a transcript
that never received a head; a later chained request restoring that head
would have omitted it. The staged value is now committed as-is, None
included. Also pins each rejection predicate of is_usable_compression_point
with its own test.
2026-09-05 11:26:25 +01:00
Alex dbee30a048 fix(compression,llm): keep the summary across mid-execution compression, ignore empty saved points, commit the head hash on success
Three review findings on the bounded-chain change.

Mid-execution compression rebuilt the conversation from the in-flight
messages, which after a turn-start reuse hold only the recent turns: the
summary living in the system prompt never reached the compressor, so the
new summary replaced the old one, and the persisted point's query_index
was relative to that shortened list. The summary the agent is running
under now rides into the synthetic conversation as its latest point
(query_index -1, so every in-flight query is new), for both the database
and the in-memory path, and the database path persists the index of the
saved conversation's last row.

Saved points with an empty summary, which earlier versions wrote, were
treated as reusable: get_compressed_context sliced the raw history away
and the effective token count made the conversation look small. Point
selection everywhere now takes the latest usable point (non-blank
summary, positive token count) and falls back to the raw history when
there is none.

The chained system-head hash was committed while building the request,
so a transport failure followed by the same-primary retry omitted a
changed system message. The hash is now staged per request and committed
only when the provider records the response.
2026-09-05 10:39:58 +01:00
Alex 1f86139b9c fix(compression,llm): address review — exact point dedupe, marked summary rows, opaque cache key
- append_compression_point only skips a point when both query_index and
  compressed_summary are present and match the last one; points without
  those fields (as in the repository tests) were all being treated as
  duplicates.
- The incremental compression tail and the orchestrator's "anything new
  since the last point" check exclude the visible summary row, which the
  prompt already receives through existing_compressions.
- Summary rows carry a persisted metadata marker; replay filters on the
  marker, and falls back to the label only for rows written before it that
  have no tool calls and no per-turn metadata, so a user who types the
  label text keeps their turn.
- The prompt_cache_key is a hash of the user id, never the id itself.
- Describe truncation="auto" as dropping the oldest items.
2026-09-05 09:30:13 +01:00
Alex 04d358ab69 fix(llm,compression): bound cross-turn Responses chaining and make compression stick
In store mode every user turn chained onto the previous response, so the
provider's stored transcript grew without bound (measured: 889k prompt
tokens for a 37k-token saved history) while every local guard, the
compression pipeline included, measured the saved history. Each chained
tool round also re-sent the system message, which the server appends rather
than dedupes, and a saved compression point was applied exactly once, in the
turn that made it.

Chaining is now bounded. A turn starts from the saved history when the
previous turn's reported prompt reached the chain budget (default: the
model's context window), when the conversation was compressed after that
turn was produced, or when OPENAI_RESPONSES_CHAIN_ACROSS_TURNS is off.
Chained rounds omit an unchanged system head (hash carried in the persisted
Responses state). truncation="auto" is available behind a setting as a
backstop against a chain that outgrows the model's window.

Compression: a saved point is applied at every turn start; the threshold
counts the summary plus the queries after the point instead of the raw
history; re-compression summarises only the tail on top of the last point;
the mid-execution path marks itself persisted and resets the provider chain
so the rebuilt messages are the context; an empty summary is rejected; the
visible "[Context Compression Summary]" rows are no longer replayed as
history; appending the same point twice is a no-op.

Cache hints: a per-user prompt_cache_key and an optional
prompt_cache_retention on Responses API calls.

Measured on Azure with the same client shape as production (stateless
OpenAI client, server-side tools, PDF part): tokens billed on the sixth turn
fell from 58k to 35k, tool rounds add tens of tokens instead of ~2.8k, the
turn after a compression reused the saved summary in under two seconds
instead of re-summarising, and the round after a mid-execution compression
started from the compressed context instead of the full stored transcript.
2026-09-05 00:28:51 +01:00
Alex f7cd94668f fix(attachments): close the .txt gap in the parseability gate
Review follow-ups on the attachment gate.

.txt was listed as parser-backed, but it has no parser — it *is* the
plain-text fallthrough. That let it skip the content check, so renaming a
video to notes.txt walked straight back into the bug the gate exists for
(verified: 5132 chars of binary "extracted" and stored). The list is now
exactly the file extractor's keys, .txt included in the content check like
any other unparsed suffix, and the drift test asserts equality rather than
containment. The sniff now recognises a UTF-16/32 BOM as text, so a
Notepad "Unicode" .txt is not caught by the NUL-byte rule.

The picker's accept filter listed parser-backed suffixes only, hiding .txt,
.py and .log — files the gate reads happily — behind "All files". It now
carries text/* as well, so it can never be narrower than what the upload
accepts.

A rejected batch carries one errors entry per file, but the non-200 branch
applied the top-level message to every chip, so two files failing for two
reasons both reported the first one. Reasons are now matched by
upload_index, with the top-level message as fallback.

_get_store_attachment_user_error no longer reads str(exc): the
unsupported-type message is rebuilt from the filename, so no exception
state can reach a response body (CodeQL py/stack-trace-exposure).
2026-09-03 09:13:24 +01:00
Alex 96cad1f8e2 fix(attachments): refuse unparseable chat attachments
A chat attachment with no parser fell through to SimpleDirectoryReader's
plain-text open(), so a phone-uploaded video was "extracted" into megabytes
of binary garbage, truncated, and stored with extraction.status == "ok".

Gate attachments in two tiers instead. A suffix with a dedicated parser is
admitted on its name — a PDF is binary and parses fine. Anything else has to
read as text: the first 8KB are sampled and refused on a NUL byte or too many
other control bytes. That keeps source, config and log files working through
the plain-text fallthrough, and keeps out videos, archives and renamed
binaries alike. The route checks the staged spool before anything is stored
or queued; the worker repeats the check where the local file exists, raising
the non-retryable AttachmentRejectedError.

SUPPORTED_ATTACHMENT_EXTENSIONS gains the parser-backed suffixes it was
missing (.tiff, .tif, .bmp, .webp, .vtt, .xml) and is now exactly the file
extractor's keys plus .txt, with a test asserting the two agree. The composer
applies the same rule client-side, so an unsupported file is named before it
costs an upload, and a test pins the frontend list to the backend one.

Attachment failures now show their reason inline under the chips rather than
only in a hover tooltip, which a touch user can never see, and only after a
send was attempted. Dropped `accept` from the dropzone: it discarded rejected
drops with no feedback and disagreed with the server about text files.
2026-09-03 08:52:12 +01:00
ManishMadan2882 57e6836ec0 fix(feedback): accept widget api_key on /api/feedback 2026-08-30 06:25:28 +05:30
Alex 12dd7c7eab fix: stop the re-embed migration from destroying the index it rebuilds
Five defects from a review of the embeddings work, four of them silent.

- Write local files atomically. `LocalStorage.save_file` streamed straight onto
  the destination, so an interrupted write left a truncated file. `reembed`
  rewrites every index it touches, and a half-written `index.faiss` loads at
  neither the old width nor the new one -- the source was unrecoverable, with
  no backup and no temp file left behind. Bytes now land beside the destination
  and move into place with `os.replace`. S3 was already safe (single PUT).

- Read pgvector chunks a page at a time. `reembed_pgvector` materialised every
  `(id, text)` row for a source before embedding -- ~1.6 GB at 200k chunks and
  several times that for non-Latin scripts, with the `PGresult` held alongside
  until the cursor closed. Inside the shipped 4Gi limit, while also holding the
  model, that is an OOMKill -- which is exactly the SIGKILL the point above
  turned into a destroyed index. It now walks the source by keyset.

- Bound the first wave of delegated embeds. The failure cooldown is only latched
  once the first `get()` returns, so every request already in flight paid the
  full EMBEDDINGS_DELEGATE_TIMEOUT: measured 64 threads all timing out together,
  and at the shipped 60s across a 96-thread WSGI pool that is an API serving
  nothing at all, health checks included. One caller now probes while the rest
  fail fast; after a single success the gate leaves the path entirely.

- Ship EMBEDDINGS_NAME commented in .env-template. The comment directly above it
  says to leave it commented when upgrading, and the line shipped set. Any value
  reaching `.env` lands in `model_fields_set`, which makes `resolve_embeddings_pin`
  bail -- so a template-derived `.env` disabled the legacy pin outright and
  repointed a populated index at a different 768-dim model, where no width check
  fires. The pin already picks granite for a fresh install and mpnet for an
  existing one, so nothing needs to be set by hand.

- Stamp `sources.model` on wiki sources. They were created with the column NULL
  and then embedded like any other source, and the boot check reads NULL as
  "pre-dates the column, therefore the legacy model" -- reporting a correctly
  embedded source as stale on every startup of every process. Stamped at
  creation, and again on each page re-embed so existing rows heal.

The two docs that promised the FAISS index survives a failed run said so of the
embed only; both now describe the write, and upgrading.mdx says to stop ingest
for the duration.
2026-08-28 16:08:15 +01:00
Alex 5c55d2610b Merge pull request #2690 from arc53/workflow-export
Workflow export
2026-08-23 11:46:26 +01:00
Alex b47bfa8a37 fix: mini hardening 2026-08-23 11:22:12 +01:00
Alex 29f661f3c7 fix: more test fixes and additions, fix loss on resume 2026-08-22 12:57:58 +01:00
Alex e96ff8658c fix: more stability for durable tasks, retry strategy, refactor dead
code
2026-08-22 09:44:09 +01:00
Alex 114585cd7d fix: little more tool call hardening 2026-08-21 15:01:45 +01:00
Pavel 33d93ba0ca fixes 3 2026-08-20 23:07:15 +02:00
Pavel 7baf33aea0 edits v1 2026-08-20 11:37:29 +02:00
Alex 0b36257202 feat: parser improvements and fixes 2026-08-19 23:49:26 +01:00
Pavel 1040a7efcc tool content loss fixes 2026-08-19 23:08:46 +02:00
Pavel 3c78540931 Workflow export 2026-08-18 17:10:26 +02:00
Alex cf075790ab fix: k8s config and silent source ingests 2026-08-13 11:25:56 +01:00
Alex e4b06609b1 fix: adopted agent image loss 2026-08-13 10:05:17 +01:00
Alex ee552620eb fix: mini nits 2026-08-12 17:05:57 +01:00
Alex 663869868f fix: better zip protection 2026-08-12 16:55:33 +01:00
Alex 4b1bc17c77 feat: image refactor 2026-08-12 12:36:43 +01:00
Alex 9aa7a01949 fix: e2e tests 2026-08-11 11:47:09 +01:00
Alex 64a6b81fbb feat: guardrails init 2026-08-11 00:09:56 +01:00
Alex 26038da499 feat: auto cancel on retry or edit during reconciliation 2026-08-10 15:26:03 +01:00
Alex ca345e3bf8 feat: improve tool call durability on long calls 2026-08-10 14:32:30 +01:00
Alex a8f1e0959e fix: track agent creation errors more 2026-08-09 11:20:35 +01:00
Alex c5d2329a61 fix: more test fixes 2026-08-08 12:28:20 +01:00
Alex e0649d25cf fix: mini issues 2026-08-08 11:59:23 +01:00
Alex 795e39a6bc fix: source authorization, silent retrieval failures, and prompt structure
Source access control
---------------------
`active_docs` is client-supplied and reached the retriever unchecked, and the
retriever queries `WHERE source_id = <id>` with no owner predicate — so any
caller could pass any source id to /stream or /api/answer and have another
tenant's documents quoted back, while /api/sources/<id>/search correctly
refused the same id. Gate it through `can_access`, the helper the guarded
endpoints already use, and filter `self.source` down to the authorized set.
Fails closed: no principal, or a check that errors, drops the source.

Three sibling paths had the same gap:

- workflow agent nodes: `AgentNodeConfig.sources` is written verbatim from
  client JSON at save time and nothing validated it, so a node could name any
  tenant's source. Gate against the workflow owner, so shared workflows keep
  reading their owner's sources like shared agents do.
- /api/share: `_resolve_source_pg_id` resolved any id with no ownership
  predicate and baked it into the agent the share creates; /api/search then
  searched it. Authorize before attaching.
- search_service: re-resolve the ids stored on an agent row instead of
  trusting them, so a row written by any future path with the same gap cannot
  be read back.

Team grantees previously lost their source's retrieval config: the post-check
read was still owner-scoped, so it missed and fell back to defaults (an
`agentic_tool` source was bulk-prefetched for every grantee). Read unscoped
after `can_access` passes.

Retrieval
---------
`PGVectorStore._ensure_table_exists` created an IVFFlat index on the empty
table it had just created. IVFFlat computes centroids at build time, so those
centroids were random, and combined with the `source_id` post-filter a source
with hundreds of embedded chunks returned zero rows — retrieval reported no
documents, the model answered from memory, and nothing was logged. Stop
creating the index (exact search is correct and fast well past the sizes most
deployments reach); raise `ivfflat.probes` to sqrt(lists) where an index still
exists; and re-run a short indexed search exactly, since post-filtering means
no index setting can guarantee a full result. `graphrag` had the same
empty-table index with no fallback at all.

Also: bound `chunks` to 0-500 on both the request and agent paths (0 still
means "skip retrieval"), let a source's configured `retrieval.chunks` outrank
the request body, and cap ClassicRAG's per-source floor at
max(top_k, n_sources) so attaching sources cannot inflate the result set.

Silent failures
---------------
An empty retrieval was invisible to both the model and the client: the `source`
event was suppressed when the list was empty, so "searched and found nothing"
looked identical to "no source attached", and the prompt said nothing at all.
Emit the event always, and tell the model when a search ran and returned
nothing. A file that parses to nothing now fails ingest with a message naming
the cause instead of storing an embedding of the empty string. `score_threshold`
returns warnings when the active store or retriever cannot honour it.

Prompt structure
----------------
Retrieved documents move from the system prompt into the user turn, with the
injection guard restated next to them: they change every turn (defeating prefix
caching), they are third-party text that should not carry system authority, and
routing them through the query budget makes them truncatable rather than
silently crowding it out. Documents are shed lowest-ranked-first before the
question is touched.

The six chat presets (3 tones x 2 retrieval modes) differed only in their
Answering section; they are now composed from single-source fragments at load
time, not through Jinja inheritance, which would have opened a file-read
surface in the template sandbox and broken the tool-prefetch parser. Per-tool
guidance moves out of the prompt into tool schemas, so it travels with the tool
and cannot render when the tool is absent. A plain-text custom prompt is staged
as a persona value inside the skeleton instead of replacing it wholesale — it
used to silently lose the injection guard, platform block, memory and
attachments, and its braces are now inert.

Other fixes
-----------
- agents/base: an oversized system prompt drove the query budget negative and
  dispatched a full-price request with an empty question; raise instead.
- llm/anthropic: migrate off the retired Text Completions API. It flattened
  history to first+last message and ignored tools entirely. Adds the missing
  Anthropic handler, without which every tool call was silently dropped.
- sources/upload: `sitemap` had no branch, so every sitemap ingest died on a
  TypeError; `validate_url` now rejects a falsy URL cleanly.
- workflow nodes: retrieved documents never reached the node agent, so a
  classic node with a source and an ordinary prompt answered "I have no
  documents" while the run reported completed.
- parser/bulk: copy the metadata dict, or every chunk reports the last chunk's
  token_count.
- crawler_loader: carry the page title, or citations render the whole chunk
  body as the label.
2026-08-08 10:21:52 +01:00
Alex 9f19bc9fc7 Merge pull request #2637 from arc53/fix/docling-parse-error-and-schedule-status
Fix/docling parse error and schedule status
2026-08-06 12:55:54 +01:00
Alex 0ecb421954 fix: error type fixes and docling parsing improvements 2026-08-06 12:29:51 +01:00
Alex bda11242e1 fix(stream): persist a turn that errored without an answer as failed
WorkflowEngine reports node failures by *yielding* `{"type": "error"}`
rather than raising, so complete_stream's generator returns normally and
the except handler never runs. The turn was finalized `status="complete"`
with an empty response.

Live, the client renders an error bubble with a Retry button. On reload
it does not: mapServerQueryToClient only surfaces `metadata.error` for
`failed` rows, so history showed a blank message with no error and no way
to retry. A user hitting this re-sends the same prompt into new
conversations, which is exactly what the 2026-08-01 report shows — nine
blank first messages in seven hours.

Tracks a `stream_error` flag alongside the existing `paused` machinery
and finalizes `failed` when the turn produced no answer, recording the
user-facing message in `metadata.error`. An error arriving *after* output
keeps `complete` so partial text is not discarded; structured answers
count as output too, since they live in `structured_chunks` rather than
`response_full`. The flag is recorded before the pause branches so those
paths cannot lose it.

save_conversation grew a `status` parameter (default `complete`) for the
non-WAL branch, which took no status and so landed on the column default
— the same blank-complete row on a path the WAL fix did not cover.
Title generation now also runs for failed turns: _maybe_generate_title
only regenerates while the name is still the question-prefix fallback, so
skipping it would strand a conversation whose first turn failed with the
raw prompt as its name forever.

logging.py counts a yielded error toward `activity_finished.status`.
These failures previously logged `status=ok` with `answer_length=0`,
which is why user-visible blank answers never appeared in error metrics.
2026-08-05 11:25:11 +01:00
Pavel f878a0f902 Redundant test 2026-07-28 15:41:25 +02:00
Pavel 6c63c91643 The conversations fix 2026-07-27 22:10:04 +02:00