Commit Graph
518 Commits
Author SHA1 Message Date
Alex 6e4480b61b fix: mini compression issues 2026-06-23 19:15:22 +01:00
Alex ba5c9a8910 fix: graph rag ui improvements 2026-06-23 17:11:49 +01:00
Alex 2594bb0bed feat(graphrag): graph-view endpoints + network visualization
GraphStore.get_graph_overview (top-N by degree, bounded) + get_node_detail (description + linked chunks). GET /api/sources/<id>/graph + /graph/node/<id> (read-access gated, node scoped to source_id; empty graph -> {nodes:[],edges:[]}). Frontend GraphView (react-force-graph-2d) wired to config.kind=graphrag via card-click/'View graph'; node-click -> description + chunks; tooltip renders untrusted names as text (no innerHTML XSS). All read-only GraphStore methods rollback their txn (no idle-in-transaction lock). Unit G7. (Also a pre-existing prettier fix in WorkflowPreview to keep lint green.)
2026-06-23 02:48:18 +01:00
Alex 4742aec4c6 feat(graphrag): durable extract_graph task + ingest-path enqueue + enable route
graph_enabled() sets kind=graphrag + retriever=graphrag. extract_graph_worker fetches the source's pgvector chunks and runs G3 extraction (graphrag_available guard, empty no-op). extract_graph durable+idempotent task; key varies with source updated_at so re-ingest/re-enable re-run incrementally (G3 checkpoint skips done chunks) while concurrent same-state enqueues dedup. The 4 ingest paths enqueue after embed when kind=graphrag (isolated in try/except so a broker hiccup can't fail the ingest). POST /api/sources/<id>/graphrag/enable: pgvector+GRAPHRAG_ENABLED gated, owner/editor write-authz; PATCH config still can't flip kind->graphrag. Unit G4.
2026-06-23 01:29:51 +01:00
Alex ef44459984 fix: small wiki fixes 2026-06-23 00:19:56 +01:00
Alex f809f660f4 fix(wiki): maintain sources.tokens for wikis (sum of page token_counts)
sources.tokens was only set at ingest, so wikis showed no/stale token count on the card (blank wikis showed nothing; converted ones showed the stale original count). rebuild_wiki_directory_structure (called after every wiki mutation) now also sets sources.tokens to the sum of wiki_pages.token_count; blank wiki create sets tokens=0. So the card reflects live wiki content.
2026-06-22 23:33:47 +01:00
Alex 919fe10ab9 feat(wiki): human page editing + provenance stamps
Unit 6 of F-Wiki (D22 + D24). Migration 0024 adds wiki_pages.updated_via; set to 'agent' by WikiTool + convert, 'human' by the edit/seed endpoints (content-hash short-circuit preserves it). WikiViewer gains a markdown editor (Save via PUT with expected_version; 409 reloads latest while keeping the draft; 403 graceful) and a per-page provenance stamp (editor/when/version). Edit gated on write access client-side; backend remains the real authz.
2026-06-22 20:42:37 +01:00
Alex 0c9e0313bd feat(wiki): convert existing source to wiki + human page edit endpoint
Unit 5 of F-Wiki (D20-D23). convert_source_to_wiki task reuses reingest's storage file-load + parser to materialize files->wiki_pages (one page/file), skips/reports non-text, re-embeds per page, and flips kind=wiki + exposure=agentic_tool only when pages were created. POST /wiki/convert (explicit, write-authz; blank source enables inline, fileful enqueues the task; rejects mid-ingest). PUT /wiki/page for human edits (write-authz, optimistic version -> 409, re-embed). PATCH /config preserves kind (kind changes only via convert).
2026-06-22 20:16:18 +01:00
Alex 5110189143 feat(wiki): create-wiki route + read endpoints + minimal read-only UI
Unit 4 of F-Wiki. POST /api/sources/wiki creates a type=wiki/kind=wiki source with no ingest task (optional seed page enqueues reembed). GET /wiki/pages + /wiki/page serve the tree + fresh page content, read-access gated (owner or team grant), path-validated. Frontend: 'Create Wiki' ingestor entry, read-only WikiViewer (FileTree + react-markdown, no raw HTML), sync/reingest hidden for wiki sources.
2026-06-22 18:41:54 +01:00
Alex 98aa8db953 feat(wiki): WikiTool with injection + write authz
Unit 3 of F-Wiki. WikiTool (view/create/str_replace/insert/delete/rename) over wiki_pages: exact-case unique str_replace (no silent multi-replace), optimistic version on edits (WikiPageConflict), reads served fresh from Postgres, untrusted-content fencing on reads, 1MB page cap. Injected via add_wiki_tool only for writable wiki sources (effective_write_owner: owner/team-editor; viewers get nothing), scoped to one source_id. Each mutation enqueues reembed_wiki_page (owner as user, per-page idempotency key) and rebuilds directory_structure. Shared validate_tool_path extracted from MemoryTool.
2026-06-22 18:18:04 +01:00
Alex d68be86244 feat(wiki): reembed_wiki_page durable task
Per-page re-embed (Unit 2 of F-Wiki): targeted delete of the page's old chunks, re-chunk via the source's chunking config, add_chunk with reingest-matching metadata (source=path), set embed_status embedded/failed. Durable + idempotent (key=content_hash), mirroring reingest_source_task. Missing page => purge only.
2026-06-22 17:49:04 +01:00
Alex 057b510adf fix: better error logging 2026-06-22 13:47:09 +01:00
Alex a9068faefa feat: small mixes on mixes sources 2026-06-22 13:13:51 +01:00
Alex babc067aa4 feat: remove old chunk management 2026-06-22 10:40:33 +01:00
Alex f6400cd736 feat: per-source RAG configuration (retrieval strategies, chunking, exposure, prescreen)
Introduces a per-source config contract that makes RAG behavior strategy-dispatched instead of a single hardcoded path. Every source gains a validated JSONB config; an empty/absent config reproduces current behavior byte-for-byte, and the whole path is gated by PER_SOURCE_RETRIEVAL_ENABLED.

Foundation: sources.config JSONB column + migration 0022_source_config; SourceConfig/ChunkingConfig/RetrievalConfig pydantic models (strict on write, lenient on read); ChunkerCreator and RetrieverCreator.register registries; config threaded through the upload routes, ingest/remote/connector workers, and reingest.

Retrieval: a Dispatcher groups sources by retriever key (all-classic collapses to today's single ClassicRAG under one shared token budget; non-classic retrievers get their own instance), removing the previous single-global-retriever collapse in stream_processor. Per-source chunks, score_threshold (honored for pgvector/mongodb, safely ignored elsewhere), and rephrase_query toggle. New PATCH /api/sources/<id>/config with team-aware (effective_write_owner) authz and a requires_reingest signal.

Chunking strategies: recursive, markdown, parent_child (selectable per source; re-ingest to apply). Search exposure: per-source prefetch vs agentic_tool for agentic/research agents. Map-reduce prescreen: optional LLM relevance pre-filter implemented as a composable post-retrieval stage that wraps any retriever.

Backend and frontend (shared Retrieval options panel + edit modal) with tests; backend suite and frontend vitest green. Excludes the wiki and GraphRAG flagships.
2026-06-20 21:54:23 +01:00
Alex 1b7f84d981 Merge remote-tracking branch 'origin/main' into team 2026-06-16 17:34:27 +01:00
Alex 26d741ab94 feat: notifications on share 2026-06-16 14:50:55 +01:00
Alex d37e6bec51 feat: better tags and dropdowns for sharing 2026-06-16 12:14:28 +01:00
Alex d21f4758be feat: teams 2026-06-16 09:21:21 +01:00
Alex c772ebc60f fix: namespace 2026-06-15 11:44:02 +01:00
Alex 9a349401f4 feat: admin dashboard 2026-06-15 11:30:02 +01:00
Alex 82cd7b7d49 fix: tests 2026-06-14 22:38:28 +01:00
Alex 42678e95cd feat: admin role rbac 2026-06-14 21:36:07 +01:00
Alex ee4abf1038 fix: keep agent published on update 2026-06-14 14:44:27 +01:00
Alex ab3337ced8 fix: stale numbers and slug reassignments 2026-06-14 14:24:18 +01:00
Alex 3e968be6cc Merge origin/main into agent-exports; renumber agent-slug migration to 0019
main added migration 0018_tool_attempts_attribution (revises 0017_oidc_scim), which collided with the feature's 0018_agent_slug. Renumbered the agent-slug migration to 0019 (revises 0018_tool_attempts_attribution) so the Alembic chain stays linear (single head). Auto-merge was conflict-free — models.py, the frontend API layer, and all 7 locale files merged additively.
2026-06-13 14:42:50 +01:00
Alex f8e2fb715e Merge pull request #2534 from arc53/logs-and-anal-revamp
revamp analytics & logs — per-agent attribution, unified log timeline
2026-06-13 14:25:32 +01:00
Alex cd6f40471a fix: sanitise on request 2026-06-13 13:55:13 +01:00
Alex 5f29a384e2 feat: agent import / export 2026-06-13 13:31:46 +01:00
Pavel 9fb59785b2 Second fixes batch 2026-06-13 15:39:16 +04:00
Pavel 59704b0f73 sec fixes 2026-06-12 18:19:15 +04:00
Alex b8eaa5a68d fix: mini tool call issues 2026-06-11 16:49:52 +01:00
Alex 636d4d35d1 feat(prompts): revamp preset prompts, tool naming, and prompt templating
Rewrite the default/creative/strict presets (classic + agentic) into
structured sections: grounding and cite-by-title guidance, insufficient-
context behavior, current date, respond-in-user-language, scoped mermaid
usage, an untrusted-content guardrail, and a conditional XML-tagged
document context block. A memory directory listing is injected at render
time via the template prefetch mechanism so the model starts oriented
without burning a tool call.

Fixes along the way:
- Agentic preset swap was dead code: _get_prompt_content cached the
  classic preset before create_agent's swap check ran, so agentic and
  research agents always got the classic preset. The swap now happens
  inside _get_prompt_content.
- Jinja autoescape corrupted document content in custom prompts
  (< -> &lt;); prompts are not HTML, autoescape is now off.
- Literal {summaries} leaked into the prompt when no docs were
  retrieved; the placeholder is now stripped.
- Agentic/research prompts referenced tool names from a dropped naming
  scheme (search_internal, reason_think); they now reference the real
  names (search, reason).
- The strict preset told the model to "be very creative and use your
  imagination" right after "never make up information".
- extract_tool_usages recorded intermediate attribute chains as
  bare-tool usages, which meant "run all actions" at prefetch; only
  maximal chains are recorded now.
- Headless runs retrieved docs but never rendered them into the
  prompt; the prompt is now rendered like the streaming path.
- Default tools were unreachable by name in prompt templates
  (prefetch results were keyed by synthetic id only); defaults now
  claim the name key unless an explicit row shadows it.

Tool layer: memory/notes/todo actions are namespaced (memory_view,
note_overwrite, todo_create, ...) with legacy unprefixed names still
accepted via prefix stripping; duplicate action names across tools are
disambiguated with the owning tool's name instead of numeric suffixes;
thin tool descriptions rewritten (brave, duckduckgo, telegram, ntfy,
cryptoprice, read_webpage, internal_search, think).

Docs are now wrapped per chunk in <document index>/<source>/<content>
tags for citation-by-title support.
2026-06-11 11:18:46 +01:00
Pavel 64e19f4d11 revamp analytics & logs — per-agent attribution, unified log timeline
- Move analytics endpoints to Postgres with agent filtering that matches both stamps (api_key for external traffic, agent_id for owner/headless)
- Add tool & schedule analytics, token grouping (model/agent/source) and side-channel toggle
- Merge chat/system/webhook/workflow/schedule events into one logs timeline with level/type/search filters
- Stamp user/agent on tool_call_attempts at propose time (migration 0018) and backfill via parent message
- Frontend: revamped Analytics charts and Logs page
2026-06-11 13:35:57 +04:00
Alex 10cc16eab0 Merge pull request #2530 from arc53/oidc-login
feat: oidc login
2026-06-10 17:36:36 +01:00
Alex 321d88d5be fix(oidc): close re-review gaps in the hardening commit
Follow-up to the OIDC security hardening, from a max-effort re-review:

- Refresh now re-checks the denylist immediately before minting, against the
  (possibly remapped) identity but anchored on the original session `iat` — so a
  back-channel logout / SCIM deny that lands during the IdP grant, or one
  targeting the refreshed sub/sid, still blocks renewal instead of being escaped
  by the renewed token's fresh iat. Completes the watermark revocation fix.
- SCIM PUT `userName` immutability check is now case-insensitive, matching the
  case-insensitive list/create — a differently-cased userName echo no longer
  400s "userName is immutable" and blocks deprovision. Completes the SCIM
  case-insensitivity fix.
- Drop the unconditional state-cookie deletion on every callback exit: it let
  one tab's callback clear another in-flight tab's cookie, breaking concurrent
  logins. The cookie self-expires (max_age) and the Redis state is single-use,
  so the delete wasn't needed.

Tests added for the refresh revocation re-check and the case-insensitive PUT.
2026-06-10 14:15:32 +01:00
Alex a3d09e8890 fix(oidc): harden session security from PR review (CSRF, revocation, refresh)
Address the high/medium correctness findings on the OIDC/SCIM PR:

- Login CSRF / session fixation: bind `state` to a Secure/HttpOnly/SameSite=Lax
  cookie at login and require the callback to echo it, so a code+state captured
  from another browser can't silently sign a victim into the attacker's account.
- Require `exp` on session JWTs under AUTH_TYPE=oidc (require_exp), so an
  exp-less HS256 token signed with JWT_SECRET_KEY can't authenticate forever or
  outlive the denylist.
- Denylist now keys revocation on an `iat` watermark instead of a deletable
  flag: a fresh login (newer iat) self-supersedes a revocation without clearing
  it, so sessions revoked on other devices stay revoked. Drops the
  login/SCIM-reactivation denylist-clearing paths (allow_user/allow_idp_sub).
- Refresh: gate the disabled-account check on the post-grant identity (not just
  the old sub); attempt the IdP grant before consuming the refresh token and
  return a retryable 503 (restoring the token) on transient IdP errors instead
  of force-logging-out a live session.
- Gate the oidc blueprint at request time on AUTH_TYPE=oidc, so non-oidc
  deployments cleanly 404 these routes instead of 500-ing on an unset
  OIDC_ISSUER (mirrors SCIM_ENABLED).
- Surface revocation write failures: back-channel logout returns 502, and SCIM
  deactivation rolls back and returns 503, when the denylist write fails — so
  the IdP retries instead of recording a logout/deprovision that didn't revoke.
- Back-channel logout: require `jti`, run the replay check unconditionally, and
  reject stale `iat` beyond the replay-cache window.
- Make migration 0017 idempotent (IF NOT EXISTS) so re-apply can't wedge startup.
- SCIM userName matching is case-insensitive (caseExact=false) for the list
  filter and create-dedup.

Tests added/updated across test_oidc.py, test_scim.py, test_auth.py,
test_app_routes.py and the SCIM integration test.
2026-06-10 13:58:58 +01:00
Alex 160736f075 feat: visibility settings 2026-06-10 01:40:53 +01:00
Alex 054f0f1b8b feat: oidc renewal, groups, logout, scim 2026-06-10 00:57:18 +01:00
Alex 8bc777a428 feat: oidc login 2026-06-09 23:24:44 +01:00
Alex 28d62f4199 feat: better tool call followup on compat mode 2026-06-09 18:20:14 +01:00
Alex 82bacfdb2b feat: conv visibility 2026-06-08 23:46:28 +01:00
Alex 73ed1cf607 feat: async events on asgi 2026-06-08 22:02:21 +01:00
Alex 1f06f31aa3 feat: better shutdowns 2026-06-08 15:07:31 +01:00
Alex 59e9523666 fix: revoke stale UI surfaces when durable state goes terminal
The "Tool approval needed" toast (and several sibling surfaces) could
linger after the state they represent was already gone. User-scoped SSE
events (tool.approval.required, schedule.autopaused, attachment.queued,
…) are durable and replayed on reconnect, but no terminal path emitted a
matching clearing event and the reconciler only wrote operator-facing
stack_logs — so a failed/expired message replayed its approval prompt
with nothing to act on, and the toast trusted event presence over the
actual message state.

Backend — emit a user-facing event on every terminal path:
- reconciler deletes pending_tool_state and publishes
  tool.approval.cleared when a stuck message is failed;
  cleanup_pending_tool_state does the same for TTL-reaped rows
- reconciler now publishes source.ingest.failed (stalled ingest),
  schedule.run.failed (timeout/pending) and schedule.completed (once)
- schedules PATCH-resume / DELETE publish schedule.resumed / .cancelled
  so a stale schedule.autopaused can't outvote them on replay
- store_attachment gains the on_poison terminal-event hook the ingest
  tasks already have; mcp_oauth_task gains a soft/hard time limit so a
  hung flow self-reports mcp.oauth.failed
- tighten the reconciler exemption to the (conversation_id, user_id)
  composite key

Frontend — stop trusting event presence over truth:
- notificationsSlice.resolveToolApproval evicts the matching
  tool.approval.required and persists its (stable) id dismissed so the
  backlog replay stays suppressed
- ToolApprovalToast drops approval events older than the resumable TTL
  window as a backstop for a lost clearing event
- schedulesSlice handles schedule.resumed / .cancelled / .completed
- upload dismissals now outlive the SSE backlog retention window

Tests cover the new clearing events at the reconciler, slice and
dispatch layers.
2026-06-07 13:17:50 +01:00
Alex 37417f99c8 fix(v1): address review feedback (multimodal provider gate, save default, leak scope, continuation sampling)
- Drop multimodal content for non-OpenAI-family providers (Google/Anthropic
  raised on OpenAI image_url parts) so multimodal requests degrade to text
  instead of returning 500.
- Stateless continuations (no conversation_id) default save_conversation to
  False, avoiding an orphan conversation with an empty question on every tool
  round (stateful continuations still default True).
- Only strip leaked reasoning from content for structured requests
  (response_format / json_schema / json_object); legitimate answers that mention
  the marker text are no longer corrupted.
- Forward sampling params (temperature, max_tokens, ...) on continuation turns.
- An explicit json_object request clears an agent-configured json_schema so it
  isn't silently overridden.
- Drop max_tokens when max_completion_tokens is also sent (OpenAI rejects both).
- Don't build a chatcmpl-None completion id from the placeholder "None" id event.
2026-06-04 12:55:47 +01:00
Alex 9eb7878ad3 feat(v1): support multimodal content arrays (text + image_url)
OpenAI-compatible clients send multimodal user turns as a `content` array of
typed parts. translate_request previously assigned the array straight to the
question, breaking the string-only retrieval / token-budgeting / history paths
(HTTP 500). Now:

- content_to_text() extracts text from content arrays for the question,
  history and system prompt, so the string paths work unchanged.
- The full content array (text + image_url parts) is preserved as
  `multimodal_content`, threaded to the agent and emitted as the final user
  message so images reach the model. Token budgeting uses the text only.

The content array (incl. image_url) now reaches the LLM call intact; images
render for vision-capable models. A text-only upstream model will reject the
image_url variant, as expected.
2026-06-04 12:23:08 +01:00
Alex 534ffc4e5b feat(v1): stateless tool continuation + sampling-param passthrough
- Stateless tool continuation. OpenAI-compatible clients (opencode, etc.)
  resend the full messages array — system, user, assistant(tool_calls),
  tool(results) — but no conversation_id, so the prior
  "conversation_id required for tool continuation" 400 broke every tool call.
  When no conversation_id is present, rebuild the agent + pending tool calls +
  tool results directly from the resent messages
  (StreamProcessor.build_continuation_from_messages) instead of loading
  server-side pending_tool_state, and call gen_continuation.

- Forward OpenAI sampling params (temperature, max_tokens,
  max_completion_tokens, top_p, frequency_penalty, presence_penalty, stop,
  seed) from the request to the LLM gen call; the agent otherwise uses its
  configured defaults.
2026-06-04 12:17:27 +01:00
Alex b1f220ec05 fix(v1): make docsgpt extension SSE frames valid chat.completion.chunks
Strict OpenAI clients (e.g. opencode / the Vercel AI SDK) validate every
streaming `data:` frame as a chat.completion.chunk. The id / source /
tool_call / tool_calls_pending events were emitted as bare `{"docsgpt": ...}`
objects with no `choices`, which those clients reject ("expected array,
received undefined" on `choices`).

Wrap the extension in an otherwise-empty chunk
(choices:[{index:0,delta:{},finish_reason:null}]) with a top-level `docsgpt`
field that OpenAI clients ignore, and skip the placeholder "None"
conversation_id frame emitted when the call is not persisted.
2026-06-04 12:07:30 +01:00
Alex 472918f577 feat(v1): support response_format / response_schema on /v1/chat/completions
Make the OpenAI-compatible Chat Completions endpoint honor per-request
Structured Outputs and keep it OpenAI-compatible.

- translate_request now forwards the request's `response_format` (json_schema)
  or a `response_schema` convenience field to the agent as its json_schema,
  overriding the agent-configured schema for that request.
- Honor `response_format.json_schema.strict` (default true); strict:false
  passes the schema through without forcing additionalProperties:false /
  all-required (OpenAI's lenient mode).
- Support `response_format {"type":"json_object"}` via the provider's native
  JSON mode.
- Keep `content` clean: some models echo their reasoning into content as
  stringified `{'type': 'thought', ...}` reprs when response_format is set.
  Strip those from content and reroute them to `reasoning_content` (OpenAI
  never puts reasoning in content), for both streaming and non-streaming.

Flow: translate_request -> StreamProcessor._configure_agent ->
Agent.json_schema / json_schema_strict / json_object -> _llm_gen ->
prepare_structured_output_format(schema, strict).
2026-06-04 11:51:13 +01:00