Token accounting:
- Drain each tool round's provider stream to exhaustion before running
tools and recursing, so the usage decorator persists exactly one
token_usage row per LLM call, at call end. Previously every round's
generator was abandoned mid-iteration and flushed together at request
teardown, writing N near-identical rows stamped with the final
round's provider counts (duplicate billing).
- Consume the Chat Completions include_usage terminal chunk (it arrives
after finish_reason and was never read) so streamed calls record
provider-exact token counts instead of tiktoken estimates.
- Claim provider-reported usage once per call (_last_usage_claimed) so
a late-finalized generator can never adopt another call's counts.
Oversized-context guards:
- Enforce Responses API function_call/function_call_output pairing in
the input builder (drop unpaired items; bypassed for store-mode
previous_response_id chaining where calls are matched server-side).
- Hard pre-send context gate: shrink oversized tool results and refuse
payloads that cannot fit the model's window before dispatch, so a
hopeless request is never sent or billed.
- Cap a single tool result entering the LLM context
(TOOL_RESULT_MAX_TOKENS, default 20000); the tool journal and
persistence keep the full result. Applied on the resume/continuation
path too.
- Skip the fallback attempt when the payload cannot fit the fallback
model's context window (10% estimation slack).
- Compression: never save a compression point that does not reduce
tokens; bound oversized verbatim fields kept after a compression
point (COMPRESSION_RECENT_FIELD_MAX_TOKENS, default 8000).
Robustness fixes from review:
- Google parallel function calls: complete index-less ToolCalls are no
longer merged into one another (dict arguments raised TypeError on
+=; second call could execute with the first call's arguments).
- Trailing-frame failures after a delivered answer no longer error the
stream or restream the whole answer from the fallback
(_stream_reached_finish).
- In-memory compression falls back to minimal pruning when the summary
is not smaller than the original.
- keep<=0 guard in the middle-truncation helpers (a tiny cap returned
marker + full text).
Frontend: tooltip on the Analytics tokens stat card explaining that
agent tool loops re-send conversation context on every step.
Enable code_executor and artifact_generator by default in chats (added to
DEFAULT_CHAT_TOOLS) — they load via the synthetic-id path user- and
conversation-scoped like scheduler, and persist artifacts (no user_tools FK).
Make read_document an internal agent-builtin flagged workflow_only so it appears
only in the workflow builder, not the classic agent picker or the Add-Tool
catalog (reuses the builtin synthetic-id path; still run-scoped and authz-gated).
Per decision, default-on code_executor runs sandboxed code without an approval
prompt (the sandbox is the trust boundary); this is documented in settings and
the threat model, and multi-tenant deployments should add per-tenant isolation
(Daytona / gVisor / egress policy).
Replace the sandbox Docling extractor with read_document, backed by the in-process
backend parser (the same one ingestion uses) and offloaded to a dedicated
'parsing' Celery queue so it can run on GPU-capable workers with predictable RAM.
The tool resolves the input ref under the run-scoped gate, enqueues the parse,
and awaits it with a timeout (degrading to an error rather than hanging); the
worker independently re-resolves the artifact through the same gate and never
trusts a raw path. Untrusted files get the upload path's safeguards (extension
whitelist, size cap, sanitized temp file, cleanup). Options: output
(markdown/text/structured/chunks), ocr, pages, engine, max_chars, include_tables,
persist, json_schema. The workflow native-file 'extract' fallback now uses the
same worker path, so document parsing no longer needs the sandbox and works on
every backend.
Also fixes the branch's periodic-task test (the sandbox reaper made it 12) and
points the dev and e2e Celery workers at the parsing queue.
The classic prompt renderer now passes artifact_parent={conversation_id} so a
normal agent's prompt can resolve a prior-turn artifact with
{{ artifacts.artifact(id) }}, scoped to its own conversation (parent-derived
authz; a missing conversation_id safely yields an empty lookup). This makes the
artifacts variable surfaced in the agent builder functional, and is the parent
wiring an agent-as-tool / subagent feature would also rely on.
Let workflow runs consume and produce documents end to end: bridge uploaded
attachments into run-scoped artifacts so nodes receive the input documents
(with a per-run cap and server-computed size/sha256, and the run row pre-created
so produced artifacts are authorized during the run); emit the run id to the
client and add a builder panel that lists, previews, and downloads a run's
artifacts; and allow attaching documents to a Preview run via the existing
upload flow.
Also fixes issues a compliance workflow surfaced: attachment ownership now keys
on the raw identity instead of a sanitized one (the sanitized form could not be
read back and could collide across users); workflow code nodes read prior state
from a state.json data file instead of templating it into the program, so
untrusted document content can never be interpolated into executed code;
structured node output wrapped in code fences is recovered; and the live
speech-to-text ownership check compares the raw identity.
Add runtime governance for the sandbox and artifact store: a per-process
concurrent-session cap with least-recently-used eviction of idle sessions and a
periodic idle reaper (Celery beat), so sandbox kernels do not accumulate. The
session manager performs all backend start/stop outside its lock and tears down
the captured handle, so eviction never closes a concurrently re-opened session.
Add per-user artifact quotas (count, total bytes, and per-file size) enforced at
persistence time as a soft cap, and best-effort cleanup of per-render scratch
directories in the sandbox workspace.
Add a code workflow node that runs code in the run-scoped sandbox session and
writes produced files as artifact references into workflow state, passing them
by reference (only id and metadata, never bytes) so downstream nodes and CEL
conditions can branch on them. Add an artifacts.* templating namespace that
resolves those references to metadata via a run-scoped lookup, available to
both the workflow engine and the prompt renderer. Extract the sandbox-to-
artifact persistence into a shared helper reused by the code node and the
code_executor tool.
Serve artifact metadata and bytes over HTTP with parent-derived
authorization: conversation-parented artifacts inherit conversation
access (owner, shared_with, or a public share token whose conversation
matches the parent), workflow-run artifacts check run ownership, and
access fails closed when the parent is missing or deleted.
Adds list/get/versions/download/restore routes, an authenticated
storage-agnostic download (sanitized Content-Disposition, 302 to a
short-lived private S3 presigned URL when that strategy is configured),
a generate_presigned_url primitive on the storage base and S3 backend,
and generalizes the tools artifact endpoint for documents/files. Shared
authorization helpers live in a dedicated module used by both surfaces.
GraphStore.get_graph_overview (top-N by degree, bounded) + get_node_detail (description + linked chunks). GET /api/sources/<id>/graph + /graph/node/<id> (read-access gated, node scoped to source_id; empty graph -> {nodes:[],edges:[]}). Frontend GraphView (react-force-graph-2d) wired to config.kind=graphrag via card-click/'View graph'; node-click -> description + chunks; tooltip renders untrusted names as text (no innerHTML XSS). All read-only GraphStore methods rollback their txn (no idle-in-transaction lock). Unit G7. (Also a pre-existing prettier fix in WorkflowPreview to keep lint green.)
graph_enabled() sets kind=graphrag + retriever=graphrag. extract_graph_worker fetches the source's pgvector chunks and runs G3 extraction (graphrag_available guard, empty no-op). extract_graph durable+idempotent task; key varies with source updated_at so re-ingest/re-enable re-run incrementally (G3 checkpoint skips done chunks) while concurrent same-state enqueues dedup. The 4 ingest paths enqueue after embed when kind=graphrag (isolated in try/except so a broker hiccup can't fail the ingest). POST /api/sources/<id>/graphrag/enable: pgvector+GRAPHRAG_ENABLED gated, owner/editor write-authz; PATCH config still can't flip kind->graphrag. Unit G4.
sources.tokens was only set at ingest, so wikis showed no/stale token count on the card (blank wikis showed nothing; converted ones showed the stale original count). rebuild_wiki_directory_structure (called after every wiki mutation) now also sets sources.tokens to the sum of wiki_pages.token_count; blank wiki create sets tokens=0. So the card reflects live wiki content.
Unit 6 of F-Wiki (D22 + D24). Migration 0024 adds wiki_pages.updated_via; set to 'agent' by WikiTool + convert, 'human' by the edit/seed endpoints (content-hash short-circuit preserves it). WikiViewer gains a markdown editor (Save via PUT with expected_version; 409 reloads latest while keeping the draft; 403 graceful) and a per-page provenance stamp (editor/when/version). Edit gated on write access client-side; backend remains the real authz.
Unit 5 of F-Wiki (D20-D23). convert_source_to_wiki task reuses reingest's storage file-load + parser to materialize files->wiki_pages (one page/file), skips/reports non-text, re-embeds per page, and flips kind=wiki + exposure=agentic_tool only when pages were created. POST /wiki/convert (explicit, write-authz; blank source enables inline, fileful enqueues the task; rejects mid-ingest). PUT /wiki/page for human edits (write-authz, optimistic version -> 409, re-embed). PATCH /config preserves kind (kind changes only via convert).
Unit 4 of F-Wiki. POST /api/sources/wiki creates a type=wiki/kind=wiki source with no ingest task (optional seed page enqueues reembed). GET /wiki/pages + /wiki/page serve the tree + fresh page content, read-access gated (owner or team grant), path-validated. Frontend: 'Create Wiki' ingestor entry, read-only WikiViewer (FileTree + react-markdown, no raw HTML), sync/reingest hidden for wiki sources.
Unit 3 of F-Wiki. WikiTool (view/create/str_replace/insert/delete/rename) over wiki_pages: exact-case unique str_replace (no silent multi-replace), optimistic version on edits (WikiPageConflict), reads served fresh from Postgres, untrusted-content fencing on reads, 1MB page cap. Injected via add_wiki_tool only for writable wiki sources (effective_write_owner: owner/team-editor; viewers get nothing), scoped to one source_id. Each mutation enqueues reembed_wiki_page (owner as user, per-page idempotency key) and rebuilds directory_structure. Shared validate_tool_path extracted from MemoryTool.
Per-page re-embed (Unit 2 of F-Wiki): targeted delete of the page's old chunks, re-chunk via the source's chunking config, add_chunk with reingest-matching metadata (source=path), set embed_status embedded/failed. Durable + idempotent (key=content_hash), mirroring reingest_source_task. Missing page => purge only.
Introduces a per-source config contract that makes RAG behavior strategy-dispatched instead of a single hardcoded path. Every source gains a validated JSONB config; an empty/absent config reproduces current behavior byte-for-byte, and the whole path is gated by PER_SOURCE_RETRIEVAL_ENABLED.
Foundation: sources.config JSONB column + migration 0022_source_config; SourceConfig/ChunkingConfig/RetrievalConfig pydantic models (strict on write, lenient on read); ChunkerCreator and RetrieverCreator.register registries; config threaded through the upload routes, ingest/remote/connector workers, and reingest.
Retrieval: a Dispatcher groups sources by retriever key (all-classic collapses to today's single ClassicRAG under one shared token budget; non-classic retrievers get their own instance), removing the previous single-global-retriever collapse in stream_processor. Per-source chunks, score_threshold (honored for pgvector/mongodb, safely ignored elsewhere), and rephrase_query toggle. New PATCH /api/sources/<id>/config with team-aware (effective_write_owner) authz and a requires_reingest signal.
Chunking strategies: recursive, markdown, parent_child (selectable per source; re-ingest to apply). Search exposure: per-source prefetch vs agentic_tool for agentic/research agents. Map-reduce prescreen: optional LLM relevance pre-filter implemented as a composable post-retrieval stage that wraps any retriever.
Backend and frontend (shared Retrieval options panel + edit modal) with tests; backend suite and frontend vitest green. Excludes the wiki and GraphRAG flagships.