GraphStore.get_graph_overview (top-N by degree, bounded) + get_node_detail (description + linked chunks). GET /api/sources/<id>/graph + /graph/node/<id> (read-access gated, node scoped to source_id; empty graph -> {nodes:[],edges:[]}). Frontend GraphView (react-force-graph-2d) wired to config.kind=graphrag via card-click/'View graph'; node-click -> description + chunks; tooltip renders untrusted names as text (no innerHTML XSS). All read-only GraphStore methods rollback their txn (no idle-in-transaction lock). Unit G7. (Also a pre-existing prettier fix in WorkflowPreview to keep lint green.)
graph_enabled() sets kind=graphrag + retriever=graphrag. extract_graph_worker fetches the source's pgvector chunks and runs G3 extraction (graphrag_available guard, empty no-op). extract_graph durable+idempotent task; key varies with source updated_at so re-ingest/re-enable re-run incrementally (G3 checkpoint skips done chunks) while concurrent same-state enqueues dedup. The 4 ingest paths enqueue after embed when kind=graphrag (isolated in try/except so a broker hiccup can't fail the ingest). POST /api/sources/<id>/graphrag/enable: pgvector+GRAPHRAG_ENABLED gated, owner/editor write-authz; PATCH config still can't flip kind->graphrag. Unit G4.
sources.tokens was only set at ingest, so wikis showed no/stale token count on the card (blank wikis showed nothing; converted ones showed the stale original count). rebuild_wiki_directory_structure (called after every wiki mutation) now also sets sources.tokens to the sum of wiki_pages.token_count; blank wiki create sets tokens=0. So the card reflects live wiki content.
Unit 6 of F-Wiki (D22 + D24). Migration 0024 adds wiki_pages.updated_via; set to 'agent' by WikiTool + convert, 'human' by the edit/seed endpoints (content-hash short-circuit preserves it). WikiViewer gains a markdown editor (Save via PUT with expected_version; 409 reloads latest while keeping the draft; 403 graceful) and a per-page provenance stamp (editor/when/version). Edit gated on write access client-side; backend remains the real authz.
Unit 5 of F-Wiki (D20-D23). convert_source_to_wiki task reuses reingest's storage file-load + parser to materialize files->wiki_pages (one page/file), skips/reports non-text, re-embeds per page, and flips kind=wiki + exposure=agentic_tool only when pages were created. POST /wiki/convert (explicit, write-authz; blank source enables inline, fileful enqueues the task; rejects mid-ingest). PUT /wiki/page for human edits (write-authz, optimistic version -> 409, re-embed). PATCH /config preserves kind (kind changes only via convert).
Unit 4 of F-Wiki. POST /api/sources/wiki creates a type=wiki/kind=wiki source with no ingest task (optional seed page enqueues reembed). GET /wiki/pages + /wiki/page serve the tree + fresh page content, read-access gated (owner or team grant), path-validated. Frontend: 'Create Wiki' ingestor entry, read-only WikiViewer (FileTree + react-markdown, no raw HTML), sync/reingest hidden for wiki sources.
Unit 3 of F-Wiki. WikiTool (view/create/str_replace/insert/delete/rename) over wiki_pages: exact-case unique str_replace (no silent multi-replace), optimistic version on edits (WikiPageConflict), reads served fresh from Postgres, untrusted-content fencing on reads, 1MB page cap. Injected via add_wiki_tool only for writable wiki sources (effective_write_owner: owner/team-editor; viewers get nothing), scoped to one source_id. Each mutation enqueues reembed_wiki_page (owner as user, per-page idempotency key) and rebuilds directory_structure. Shared validate_tool_path extracted from MemoryTool.
Per-page re-embed (Unit 2 of F-Wiki): targeted delete of the page's old chunks, re-chunk via the source's chunking config, add_chunk with reingest-matching metadata (source=path), set embed_status embedded/failed. Durable + idempotent (key=content_hash), mirroring reingest_source_task. Missing page => purge only.
Introduces a per-source config contract that makes RAG behavior strategy-dispatched instead of a single hardcoded path. Every source gains a validated JSONB config; an empty/absent config reproduces current behavior byte-for-byte, and the whole path is gated by PER_SOURCE_RETRIEVAL_ENABLED.
Foundation: sources.config JSONB column + migration 0022_source_config; SourceConfig/ChunkingConfig/RetrievalConfig pydantic models (strict on write, lenient on read); ChunkerCreator and RetrieverCreator.register registries; config threaded through the upload routes, ingest/remote/connector workers, and reingest.
Retrieval: a Dispatcher groups sources by retriever key (all-classic collapses to today's single ClassicRAG under one shared token budget; non-classic retrievers get their own instance), removing the previous single-global-retriever collapse in stream_processor. Per-source chunks, score_threshold (honored for pgvector/mongodb, safely ignored elsewhere), and rephrase_query toggle. New PATCH /api/sources/<id>/config with team-aware (effective_write_owner) authz and a requires_reingest signal.
Chunking strategies: recursive, markdown, parent_child (selectable per source; re-ingest to apply). Search exposure: per-source prefetch vs agentic_tool for agentic/research agents. Map-reduce prescreen: optional LLM relevance pre-filter implemented as a composable post-retrieval stage that wraps any retriever.
Backend and frontend (shared Retrieval options panel + edit modal) with tests; backend suite and frontend vitest green. Excludes the wiki and GraphRAG flagships.
main added migration 0018_tool_attempts_attribution (revises 0017_oidc_scim), which collided with the feature's 0018_agent_slug. Renumbered the agent-slug migration to 0019 (revises 0018_tool_attempts_attribution) so the Alembic chain stays linear (single head). Auto-merge was conflict-free — models.py, the frontend API layer, and all 7 locale files merged additively.
Rewrite the default/creative/strict presets (classic + agentic) into
structured sections: grounding and cite-by-title guidance, insufficient-
context behavior, current date, respond-in-user-language, scoped mermaid
usage, an untrusted-content guardrail, and a conditional XML-tagged
document context block. A memory directory listing is injected at render
time via the template prefetch mechanism so the model starts oriented
without burning a tool call.
Fixes along the way:
- Agentic preset swap was dead code: _get_prompt_content cached the
classic preset before create_agent's swap check ran, so agentic and
research agents always got the classic preset. The swap now happens
inside _get_prompt_content.
- Jinja autoescape corrupted document content in custom prompts
(< -> <); prompts are not HTML, autoescape is now off.
- Literal {summaries} leaked into the prompt when no docs were
retrieved; the placeholder is now stripped.
- Agentic/research prompts referenced tool names from a dropped naming
scheme (search_internal, reason_think); they now reference the real
names (search, reason).
- The strict preset told the model to "be very creative and use your
imagination" right after "never make up information".
- extract_tool_usages recorded intermediate attribute chains as
bare-tool usages, which meant "run all actions" at prefetch; only
maximal chains are recorded now.
- Headless runs retrieved docs but never rendered them into the
prompt; the prompt is now rendered like the streaming path.
- Default tools were unreachable by name in prompt templates
(prefetch results were keyed by synthetic id only); defaults now
claim the name key unless an explicit row shadows it.
Tool layer: memory/notes/todo actions are namespaced (memory_view,
note_overwrite, todo_create, ...) with legacy unprefixed names still
accepted via prefix stripping; duplicate action names across tools are
disambiguated with the owning tool's name instead of numeric suffixes;
thin tool descriptions rewritten (brave, duckduckgo, telegram, ntfy,
cryptoprice, read_webpage, internal_search, think).
Docs are now wrapped per chunk in <document index>/<source>/<content>
tags for citation-by-title support.
- Move analytics endpoints to Postgres with agent filtering that matches both stamps (api_key for external traffic, agent_id for owner/headless)
- Add tool & schedule analytics, token grouping (model/agent/source) and side-channel toggle
- Merge chat/system/webhook/workflow/schedule events into one logs timeline with level/type/search filters
- Stamp user/agent on tool_call_attempts at propose time (migration 0018) and backfill via parent message
- Frontend: revamped Analytics charts and Logs page
Follow-up to the OIDC security hardening, from a max-effort re-review:
- Refresh now re-checks the denylist immediately before minting, against the
(possibly remapped) identity but anchored on the original session `iat` — so a
back-channel logout / SCIM deny that lands during the IdP grant, or one
targeting the refreshed sub/sid, still blocks renewal instead of being escaped
by the renewed token's fresh iat. Completes the watermark revocation fix.
- SCIM PUT `userName` immutability check is now case-insensitive, matching the
case-insensitive list/create — a differently-cased userName echo no longer
400s "userName is immutable" and blocks deprovision. Completes the SCIM
case-insensitivity fix.
- Drop the unconditional state-cookie deletion on every callback exit: it let
one tab's callback clear another in-flight tab's cookie, breaking concurrent
logins. The cookie self-expires (max_age) and the Redis state is single-use,
so the delete wasn't needed.
Tests added for the refresh revocation re-check and the case-insensitive PUT.
Address the high/medium correctness findings on the OIDC/SCIM PR:
- Login CSRF / session fixation: bind `state` to a Secure/HttpOnly/SameSite=Lax
cookie at login and require the callback to echo it, so a code+state captured
from another browser can't silently sign a victim into the attacker's account.
- Require `exp` on session JWTs under AUTH_TYPE=oidc (require_exp), so an
exp-less HS256 token signed with JWT_SECRET_KEY can't authenticate forever or
outlive the denylist.
- Denylist now keys revocation on an `iat` watermark instead of a deletable
flag: a fresh login (newer iat) self-supersedes a revocation without clearing
it, so sessions revoked on other devices stay revoked. Drops the
login/SCIM-reactivation denylist-clearing paths (allow_user/allow_idp_sub).
- Refresh: gate the disabled-account check on the post-grant identity (not just
the old sub); attempt the IdP grant before consuming the refresh token and
return a retryable 503 (restoring the token) on transient IdP errors instead
of force-logging-out a live session.
- Gate the oidc blueprint at request time on AUTH_TYPE=oidc, so non-oidc
deployments cleanly 404 these routes instead of 500-ing on an unset
OIDC_ISSUER (mirrors SCIM_ENABLED).
- Surface revocation write failures: back-channel logout returns 502, and SCIM
deactivation rolls back and returns 503, when the denylist write fails — so
the IdP retries instead of recording a logout/deprovision that didn't revoke.
- Back-channel logout: require `jti`, run the replay check unconditionally, and
reject stale `iat` beyond the replay-cache window.
- Make migration 0017 idempotent (IF NOT EXISTS) so re-apply can't wedge startup.
- SCIM userName matching is case-insensitive (caseExact=false) for the list
filter and create-dedup.
Tests added/updated across test_oidc.py, test_scim.py, test_auth.py,
test_app_routes.py and the SCIM integration test.
The "Tool approval needed" toast (and several sibling surfaces) could
linger after the state they represent was already gone. User-scoped SSE
events (tool.approval.required, schedule.autopaused, attachment.queued,
…) are durable and replayed on reconnect, but no terminal path emitted a
matching clearing event and the reconciler only wrote operator-facing
stack_logs — so a failed/expired message replayed its approval prompt
with nothing to act on, and the toast trusted event presence over the
actual message state.
Backend — emit a user-facing event on every terminal path:
- reconciler deletes pending_tool_state and publishes
tool.approval.cleared when a stuck message is failed;
cleanup_pending_tool_state does the same for TTL-reaped rows
- reconciler now publishes source.ingest.failed (stalled ingest),
schedule.run.failed (timeout/pending) and schedule.completed (once)
- schedules PATCH-resume / DELETE publish schedule.resumed / .cancelled
so a stale schedule.autopaused can't outvote them on replay
- store_attachment gains the on_poison terminal-event hook the ingest
tasks already have; mcp_oauth_task gains a soft/hard time limit so a
hung flow self-reports mcp.oauth.failed
- tighten the reconciler exemption to the (conversation_id, user_id)
composite key
Frontend — stop trusting event presence over truth:
- notificationsSlice.resolveToolApproval evicts the matching
tool.approval.required and persists its (stable) id dismissed so the
backlog replay stays suppressed
- ToolApprovalToast drops approval events older than the resumable TTL
window as a backstop for a lost clearing event
- schedulesSlice handles schedule.resumed / .cancelled / .completed
- upload dismissals now outlive the SSE backlog retention window
Tests cover the new clearing events at the reconciler, slice and
dispatch layers.
- Drop multimodal content for non-OpenAI-family providers (Google/Anthropic
raised on OpenAI image_url parts) so multimodal requests degrade to text
instead of returning 500.
- Stateless continuations (no conversation_id) default save_conversation to
False, avoiding an orphan conversation with an empty question on every tool
round (stateful continuations still default True).
- Only strip leaked reasoning from content for structured requests
(response_format / json_schema / json_object); legitimate answers that mention
the marker text are no longer corrupted.
- Forward sampling params (temperature, max_tokens, ...) on continuation turns.
- An explicit json_object request clears an agent-configured json_schema so it
isn't silently overridden.
- Drop max_tokens when max_completion_tokens is also sent (OpenAI rejects both).
- Don't build a chatcmpl-None completion id from the placeholder "None" id event.
OpenAI-compatible clients send multimodal user turns as a `content` array of
typed parts. translate_request previously assigned the array straight to the
question, breaking the string-only retrieval / token-budgeting / history paths
(HTTP 500). Now:
- content_to_text() extracts text from content arrays for the question,
history and system prompt, so the string paths work unchanged.
- The full content array (text + image_url parts) is preserved as
`multimodal_content`, threaded to the agent and emitted as the final user
message so images reach the model. Token budgeting uses the text only.
The content array (incl. image_url) now reaches the LLM call intact; images
render for vision-capable models. A text-only upstream model will reject the
image_url variant, as expected.
- Stateless tool continuation. OpenAI-compatible clients (opencode, etc.)
resend the full messages array — system, user, assistant(tool_calls),
tool(results) — but no conversation_id, so the prior
"conversation_id required for tool continuation" 400 broke every tool call.
When no conversation_id is present, rebuild the agent + pending tool calls +
tool results directly from the resent messages
(StreamProcessor.build_continuation_from_messages) instead of loading
server-side pending_tool_state, and call gen_continuation.
- Forward OpenAI sampling params (temperature, max_tokens,
max_completion_tokens, top_p, frequency_penalty, presence_penalty, stop,
seed) from the request to the LLM gen call; the agent otherwise uses its
configured defaults.
Strict OpenAI clients (e.g. opencode / the Vercel AI SDK) validate every
streaming `data:` frame as a chat.completion.chunk. The id / source /
tool_call / tool_calls_pending events were emitted as bare `{"docsgpt": ...}`
objects with no `choices`, which those clients reject ("expected array,
received undefined" on `choices`).
Wrap the extension in an otherwise-empty chunk
(choices:[{index:0,delta:{},finish_reason:null}]) with a top-level `docsgpt`
field that OpenAI clients ignore, and skip the placeholder "None"
conversation_id frame emitted when the call is not persisted.
Make the OpenAI-compatible Chat Completions endpoint honor per-request
Structured Outputs and keep it OpenAI-compatible.
- translate_request now forwards the request's `response_format` (json_schema)
or a `response_schema` convenience field to the agent as its json_schema,
overriding the agent-configured schema for that request.
- Honor `response_format.json_schema.strict` (default true); strict:false
passes the schema through without forcing additionalProperties:false /
all-required (OpenAI's lenient mode).
- Support `response_format {"type":"json_object"}` via the provider's native
JSON mode.
- Keep `content` clean: some models echo their reasoning into content as
stringified `{'type': 'thought', ...}` reprs when response_format is set.
Strip those from content and reroute them to `reasoning_content` (OpenAI
never puts reasoning in content), for both streaming and non-streaming.
Flow: translate_request -> StreamProcessor._configure_agent ->
Agent.json_schema / json_schema_strict / json_object -> _llm_gen ->
prepare_structured_output_format(schema, strict).