Token accounting:
- Drain each tool round's provider stream to exhaustion before running
tools and recursing, so the usage decorator persists exactly one
token_usage row per LLM call, at call end. Previously every round's
generator was abandoned mid-iteration and flushed together at request
teardown, writing N near-identical rows stamped with the final
round's provider counts (duplicate billing).
- Consume the Chat Completions include_usage terminal chunk (it arrives
after finish_reason and was never read) so streamed calls record
provider-exact token counts instead of tiktoken estimates.
- Claim provider-reported usage once per call (_last_usage_claimed) so
a late-finalized generator can never adopt another call's counts.
Oversized-context guards:
- Enforce Responses API function_call/function_call_output pairing in
the input builder (drop unpaired items; bypassed for store-mode
previous_response_id chaining where calls are matched server-side).
- Hard pre-send context gate: shrink oversized tool results and refuse
payloads that cannot fit the model's window before dispatch, so a
hopeless request is never sent or billed.
- Cap a single tool result entering the LLM context
(TOOL_RESULT_MAX_TOKENS, default 20000); the tool journal and
persistence keep the full result. Applied on the resume/continuation
path too.
- Skip the fallback attempt when the payload cannot fit the fallback
model's context window (10% estimation slack).
- Compression: never save a compression point that does not reduce
tokens; bound oversized verbatim fields kept after a compression
point (COMPRESSION_RECENT_FIELD_MAX_TOKENS, default 8000).
Robustness fixes from review:
- Google parallel function calls: complete index-less ToolCalls are no
longer merged into one another (dict arguments raised TypeError on
+=; second call could execute with the first call's arguments).
- Trailing-frame failures after a delivered answer no longer error the
stream or restream the whole answer from the fallback
(_stream_reached_finish).
- In-memory compression falls back to minimal pruning when the summary
is not smaller than the original.
- keep<=0 guard in the middle-truncation helpers (a tiny cap returned
marker + full text).
Frontend: tooltip on the Analytics tokens stat card explaining that
agent tool loops re-send conversation context on every step.
Wrap a storage read in a context manager (close the file handle), add a comment
on the expected SandboxCapacityError in the churn test, and unify the
artifacts-routes test on a single import style. Also annotate the intentional
per-session sandbox workspace path with # nosec B108 (controlled dir inside the
runner container, not an insecure shared temp file).
Replace the sandbox Docling extractor with read_document, backed by the in-process
backend parser (the same one ingestion uses) and offloaded to a dedicated
'parsing' Celery queue so it can run on GPU-capable workers with predictable RAM.
The tool resolves the input ref under the run-scoped gate, enqueues the parse,
and awaits it with a timeout (degrading to an error rather than hanging); the
worker independently re-resolves the artifact through the same gate and never
trusts a raw path. Untrusted files get the upload path's safeguards (extension
whitelist, size cap, sanitized temp file, cleanup). Options: output
(markdown/text/structured/chunks), ocr, pages, engine, max_chars, include_tables,
persist, json_schema. The workflow native-file 'extract' fallback now uses the
same worker path, so document parsing no longer needs the sandbox and works on
every backend.
Also fixes the branch's periodic-task test (the sandbox reaper made it 12) and
points the dev and e2e Celery workers at the parsing queue.
The classic prompt renderer now passes artifact_parent={conversation_id} so a
normal agent's prompt can resolve a prior-turn artifact with
{{ artifacts.artifact(id) }}, scoped to its own conversation (parent-derived
authz; a missing conversation_id safely yields an empty lookup). This makes the
artifacts variable surfaced in the agent builder functional, and is the parent
wiring an agent-as-tool / subagent feature would also rely on.
Let workflow runs consume and produce documents end to end: bridge uploaded
attachments into run-scoped artifacts so nodes receive the input documents
(with a per-run cap and server-computed size/sha256, and the run row pre-created
so produced artifacts are authorized during the run); emit the run id to the
client and add a builder panel that lists, previews, and downloads a run's
artifacts; and allow attaching documents to a Preview run via the existing
upload flow.
Also fixes issues a compliance workflow surfaced: attachment ownership now keys
on the raw identity instead of a sanitized one (the sanitized form could not be
read back and could collide across users); workflow code nodes read prior state
from a state.json data file instead of templating it into the program, so
untrusted document content can never be interpolated into executed code;
structured node output wrapped in code fences is recovered; and the live
speech-to-text ownership check compares the raw identity.
Serve artifact metadata and bytes over HTTP with parent-derived
authorization: conversation-parented artifacts inherit conversation
access (owner, shared_with, or a public share token whose conversation
matches the parent), workflow-run artifacts check run ownership, and
access fails closed when the parent is missing or deleted.
Adds list/get/versions/download/restore routes, an authenticated
storage-agnostic download (sanitized Content-Disposition, 302 to a
short-lived private S3 presigned URL when that strategy is configured),
a generate_presigned_url primitive on the storage base and S3 backend,
and generalizes the tools artifact endpoint for documents/files. Shared
authorization helpers live in a dedicated module used by both surfaces.
GraphStore.get_graph_overview (top-N by degree, bounded) + get_node_detail (description + linked chunks). GET /api/sources/<id>/graph + /graph/node/<id> (read-access gated, node scoped to source_id; empty graph -> {nodes:[],edges:[]}). Frontend GraphView (react-force-graph-2d) wired to config.kind=graphrag via card-click/'View graph'; node-click -> description + chunks; tooltip renders untrusted names as text (no innerHTML XSS). All read-only GraphStore methods rollback their txn (no idle-in-transaction lock). Unit G7. (Also a pre-existing prettier fix in WorkflowPreview to keep lint green.)
graph_enabled() sets kind=graphrag + retriever=graphrag. extract_graph_worker fetches the source's pgvector chunks and runs G3 extraction (graphrag_available guard, empty no-op). extract_graph durable+idempotent task; key varies with source updated_at so re-ingest/re-enable re-run incrementally (G3 checkpoint skips done chunks) while concurrent same-state enqueues dedup. The 4 ingest paths enqueue after embed when kind=graphrag (isolated in try/except so a broker hiccup can't fail the ingest). POST /api/sources/<id>/graphrag/enable: pgvector+GRAPHRAG_ENABLED gated, owner/editor write-authz; PATCH config still can't flip kind->graphrag. Unit G4.
sources.tokens was only set at ingest, so wikis showed no/stale token count on the card (blank wikis showed nothing; converted ones showed the stale original count). rebuild_wiki_directory_structure (called after every wiki mutation) now also sets sources.tokens to the sum of wiki_pages.token_count; blank wiki create sets tokens=0. So the card reflects live wiki content.
Unit 6 of F-Wiki (D22 + D24). Migration 0024 adds wiki_pages.updated_via; set to 'agent' by WikiTool + convert, 'human' by the edit/seed endpoints (content-hash short-circuit preserves it). WikiViewer gains a markdown editor (Save via PUT with expected_version; 409 reloads latest while keeping the draft; 403 graceful) and a per-page provenance stamp (editor/when/version). Edit gated on write access client-side; backend remains the real authz.
Unit 5 of F-Wiki (D20-D23). convert_source_to_wiki task reuses reingest's storage file-load + parser to materialize files->wiki_pages (one page/file), skips/reports non-text, re-embeds per page, and flips kind=wiki + exposure=agentic_tool only when pages were created. POST /wiki/convert (explicit, write-authz; blank source enables inline, fileful enqueues the task; rejects mid-ingest). PUT /wiki/page for human edits (write-authz, optimistic version -> 409, re-embed). PATCH /config preserves kind (kind changes only via convert).
Unit 4 of F-Wiki. POST /api/sources/wiki creates a type=wiki/kind=wiki source with no ingest task (optional seed page enqueues reembed). GET /wiki/pages + /wiki/page serve the tree + fresh page content, read-access gated (owner or team grant), path-validated. Frontend: 'Create Wiki' ingestor entry, read-only WikiViewer (FileTree + react-markdown, no raw HTML), sync/reingest hidden for wiki sources.
Introduces a per-source config contract that makes RAG behavior strategy-dispatched instead of a single hardcoded path. Every source gains a validated JSONB config; an empty/absent config reproduces current behavior byte-for-byte, and the whole path is gated by PER_SOURCE_RETRIEVAL_ENABLED.
Foundation: sources.config JSONB column + migration 0022_source_config; SourceConfig/ChunkingConfig/RetrievalConfig pydantic models (strict on write, lenient on read); ChunkerCreator and RetrieverCreator.register registries; config threaded through the upload routes, ingest/remote/connector workers, and reingest.
Retrieval: a Dispatcher groups sources by retriever key (all-classic collapses to today's single ClassicRAG under one shared token budget; non-classic retrievers get their own instance), removing the previous single-global-retriever collapse in stream_processor. Per-source chunks, score_threshold (honored for pgvector/mongodb, safely ignored elsewhere), and rephrase_query toggle. New PATCH /api/sources/<id>/config with team-aware (effective_write_owner) authz and a requires_reingest signal.
Chunking strategies: recursive, markdown, parent_child (selectable per source; re-ingest to apply). Search exposure: per-source prefetch vs agentic_tool for agentic/research agents. Map-reduce prescreen: optional LLM relevance pre-filter implemented as a composable post-retrieval stage that wraps any retriever.
Backend and frontend (shared Retrieval options panel + edit modal) with tests; backend suite and frontend vitest green. Excludes the wiki and GraphRAG flagships.
main added migration 0018_tool_attempts_attribution (revises 0017_oidc_scim), which collided with the feature's 0018_agent_slug. Renumbered the agent-slug migration to 0019 (revises 0018_tool_attempts_attribution) so the Alembic chain stays linear (single head). Auto-merge was conflict-free — models.py, the frontend API layer, and all 7 locale files merged additively.
Rewrite the default/creative/strict presets (classic + agentic) into
structured sections: grounding and cite-by-title guidance, insufficient-
context behavior, current date, respond-in-user-language, scoped mermaid
usage, an untrusted-content guardrail, and a conditional XML-tagged
document context block. A memory directory listing is injected at render
time via the template prefetch mechanism so the model starts oriented
without burning a tool call.
Fixes along the way:
- Agentic preset swap was dead code: _get_prompt_content cached the
classic preset before create_agent's swap check ran, so agentic and
research agents always got the classic preset. The swap now happens
inside _get_prompt_content.
- Jinja autoescape corrupted document content in custom prompts
(< -> <); prompts are not HTML, autoescape is now off.
- Literal {summaries} leaked into the prompt when no docs were
retrieved; the placeholder is now stripped.
- Agentic/research prompts referenced tool names from a dropped naming
scheme (search_internal, reason_think); they now reference the real
names (search, reason).
- The strict preset told the model to "be very creative and use your
imagination" right after "never make up information".
- extract_tool_usages recorded intermediate attribute chains as
bare-tool usages, which meant "run all actions" at prefetch; only
maximal chains are recorded now.
- Headless runs retrieved docs but never rendered them into the
prompt; the prompt is now rendered like the streaming path.
- Default tools were unreachable by name in prompt templates
(prefetch results were keyed by synthetic id only); defaults now
claim the name key unless an explicit row shadows it.
Tool layer: memory/notes/todo actions are namespaced (memory_view,
note_overwrite, todo_create, ...) with legacy unprefixed names still
accepted via prefix stripping; duplicate action names across tools are
disambiguated with the owning tool's name instead of numeric suffixes;
thin tool descriptions rewritten (brave, duckduckgo, telegram, ntfy,
cryptoprice, read_webpage, internal_search, think).
Docs are now wrapped per chunk in <document index>/<source>/<content>
tags for citation-by-title support.
- Move analytics endpoints to Postgres with agent filtering that matches both stamps (api_key for external traffic, agent_id for owner/headless)
- Add tool & schedule analytics, token grouping (model/agent/source) and side-channel toggle
- Merge chat/system/webhook/workflow/schedule events into one logs timeline with level/type/search filters
- Stamp user/agent on tool_call_attempts at propose time (migration 0018) and backfill via parent message
- Frontend: revamped Analytics charts and Logs page