Source access control
---------------------
`active_docs` is client-supplied and reached the retriever unchecked, and the
retriever queries `WHERE source_id = <id>` with no owner predicate — so any
caller could pass any source id to /stream or /api/answer and have another
tenant's documents quoted back, while /api/sources/<id>/search correctly
refused the same id. Gate it through `can_access`, the helper the guarded
endpoints already use, and filter `self.source` down to the authorized set.
Fails closed: no principal, or a check that errors, drops the source.
Three sibling paths had the same gap:
- workflow agent nodes: `AgentNodeConfig.sources` is written verbatim from
client JSON at save time and nothing validated it, so a node could name any
tenant's source. Gate against the workflow owner, so shared workflows keep
reading their owner's sources like shared agents do.
- /api/share: `_resolve_source_pg_id` resolved any id with no ownership
predicate and baked it into the agent the share creates; /api/search then
searched it. Authorize before attaching.
- search_service: re-resolve the ids stored on an agent row instead of
trusting them, so a row written by any future path with the same gap cannot
be read back.
Team grantees previously lost their source's retrieval config: the post-check
read was still owner-scoped, so it missed and fell back to defaults (an
`agentic_tool` source was bulk-prefetched for every grantee). Read unscoped
after `can_access` passes.
Retrieval
---------
`PGVectorStore._ensure_table_exists` created an IVFFlat index on the empty
table it had just created. IVFFlat computes centroids at build time, so those
centroids were random, and combined with the `source_id` post-filter a source
with hundreds of embedded chunks returned zero rows — retrieval reported no
documents, the model answered from memory, and nothing was logged. Stop
creating the index (exact search is correct and fast well past the sizes most
deployments reach); raise `ivfflat.probes` to sqrt(lists) where an index still
exists; and re-run a short indexed search exactly, since post-filtering means
no index setting can guarantee a full result. `graphrag` had the same
empty-table index with no fallback at all.
Also: bound `chunks` to 0-500 on both the request and agent paths (0 still
means "skip retrieval"), let a source's configured `retrieval.chunks` outrank
the request body, and cap ClassicRAG's per-source floor at
max(top_k, n_sources) so attaching sources cannot inflate the result set.
Silent failures
---------------
An empty retrieval was invisible to both the model and the client: the `source`
event was suppressed when the list was empty, so "searched and found nothing"
looked identical to "no source attached", and the prompt said nothing at all.
Emit the event always, and tell the model when a search ran and returned
nothing. A file that parses to nothing now fails ingest with a message naming
the cause instead of storing an embedding of the empty string. `score_threshold`
returns warnings when the active store or retriever cannot honour it.
Prompt structure
----------------
Retrieved documents move from the system prompt into the user turn, with the
injection guard restated next to them: they change every turn (defeating prefix
caching), they are third-party text that should not carry system authority, and
routing them through the query budget makes them truncatable rather than
silently crowding it out. Documents are shed lowest-ranked-first before the
question is touched.
The six chat presets (3 tones x 2 retrieval modes) differed only in their
Answering section; they are now composed from single-source fragments at load
time, not through Jinja inheritance, which would have opened a file-read
surface in the template sandbox and broken the tool-prefetch parser. Per-tool
guidance moves out of the prompt into tool schemas, so it travels with the tool
and cannot render when the tool is absent. A plain-text custom prompt is staged
as a persona value inside the skeleton instead of replacing it wholesale — it
used to silently lose the injection guard, platform block, memory and
attachments, and its braces are now inert.
Other fixes
-----------
- agents/base: an oversized system prompt drove the query budget negative and
dispatched a full-price request with an empty question; raise instead.
- llm/anthropic: migrate off the retired Text Completions API. It flattened
history to first+last message and ignored tools entirely. Adds the missing
Anthropic handler, without which every tool call was silently dropped.
- sources/upload: `sitemap` had no branch, so every sitemap ingest died on a
TypeError; `validate_url` now rejects a falsy URL cleanly.
- workflow nodes: retrieved documents never reached the node agent, so a
classic node with a source and an ordinary prompt answered "I have no
documents" while the run reported completed.
- parser/bulk: copy the metadata dict, or every chunk reports the last chunk's
token_count.
- crawler_loader: carry the page title, or citations render the whole chunk
body as the label.
The classic prompt renderer now passes artifact_parent={conversation_id} so a
normal agent's prompt can resolve a prior-turn artifact with
{{ artifacts.artifact(id) }}, scoped to its own conversation (parent-derived
authz; a missing conversation_id safely yields an empty lookup). This makes the
artifacts variable surfaced in the agent builder functional, and is the parent
wiring an agent-as-tool / subagent feature would also rely on.
Let workflow runs consume and produce documents end to end: bridge uploaded
attachments into run-scoped artifacts so nodes receive the input documents
(with a per-run cap and server-computed size/sha256, and the run row pre-created
so produced artifacts are authorized during the run); emit the run id to the
client and add a builder panel that lists, previews, and downloads a run's
artifacts; and allow attaching documents to a Preview run via the existing
upload flow.
Also fixes issues a compliance workflow surfaced: attachment ownership now keys
on the raw identity instead of a sanitized one (the sanitized form could not be
read back and could collide across users); workflow code nodes read prior state
from a state.json data file instead of templating it into the program, so
untrusted document content can never be interpolated into executed code;
structured node output wrapped in code fences is recovered; and the live
speech-to-text ownership check compares the raw identity.
Unit 3 of F-Wiki. WikiTool (view/create/str_replace/insert/delete/rename) over wiki_pages: exact-case unique str_replace (no silent multi-replace), optimistic version on edits (WikiPageConflict), reads served fresh from Postgres, untrusted-content fencing on reads, 1MB page cap. Injected via add_wiki_tool only for writable wiki sources (effective_write_owner: owner/team-editor; viewers get nothing), scoped to one source_id. Each mutation enqueues reembed_wiki_page (owner as user, per-page idempotency key) and rebuilds directory_structure. Shared validate_tool_path extracted from MemoryTool.
Introduces a per-source config contract that makes RAG behavior strategy-dispatched instead of a single hardcoded path. Every source gains a validated JSONB config; an empty/absent config reproduces current behavior byte-for-byte, and the whole path is gated by PER_SOURCE_RETRIEVAL_ENABLED.
Foundation: sources.config JSONB column + migration 0022_source_config; SourceConfig/ChunkingConfig/RetrievalConfig pydantic models (strict on write, lenient on read); ChunkerCreator and RetrieverCreator.register registries; config threaded through the upload routes, ingest/remote/connector workers, and reingest.
Retrieval: a Dispatcher groups sources by retriever key (all-classic collapses to today's single ClassicRAG under one shared token budget; non-classic retrievers get their own instance), removing the previous single-global-retriever collapse in stream_processor. Per-source chunks, score_threshold (honored for pgvector/mongodb, safely ignored elsewhere), and rephrase_query toggle. New PATCH /api/sources/<id>/config with team-aware (effective_write_owner) authz and a requires_reingest signal.
Chunking strategies: recursive, markdown, parent_child (selectable per source; re-ingest to apply). Search exposure: per-source prefetch vs agentic_tool for agentic/research agents. Map-reduce prescreen: optional LLM relevance pre-filter implemented as a composable post-retrieval stage that wraps any retriever.
Backend and frontend (shared Retrieval options panel + edit modal) with tests; backend suite and frontend vitest green. Excludes the wiki and GraphRAG flagships.
Rewrite the default/creative/strict presets (classic + agentic) into
structured sections: grounding and cite-by-title guidance, insufficient-
context behavior, current date, respond-in-user-language, scoped mermaid
usage, an untrusted-content guardrail, and a conditional XML-tagged
document context block. A memory directory listing is injected at render
time via the template prefetch mechanism so the model starts oriented
without burning a tool call.
Fixes along the way:
- Agentic preset swap was dead code: _get_prompt_content cached the
classic preset before create_agent's swap check ran, so agentic and
research agents always got the classic preset. The swap now happens
inside _get_prompt_content.
- Jinja autoescape corrupted document content in custom prompts
(< -> <); prompts are not HTML, autoescape is now off.
- Literal {summaries} leaked into the prompt when no docs were
retrieved; the placeholder is now stripped.
- Agentic/research prompts referenced tool names from a dropped naming
scheme (search_internal, reason_think); they now reference the real
names (search, reason).
- The strict preset told the model to "be very creative and use your
imagination" right after "never make up information".
- extract_tool_usages recorded intermediate attribute chains as
bare-tool usages, which meant "run all actions" at prefetch; only
maximal chains are recorded now.
- Headless runs retrieved docs but never rendered them into the
prompt; the prompt is now rendered like the streaming path.
- Default tools were unreachable by name in prompt templates
(prefetch results were keyed by synthetic id only); defaults now
claim the name key unless an explicit row shadows it.
Tool layer: memory/notes/todo actions are namespaced (memory_view,
note_overwrite, todo_create, ...) with legacy unprefixed names still
accepted via prefix stripping; duplicate action names across tools are
disambiguated with the owning tool's name instead of numeric suffixes;
thin tool descriptions rewritten (brave, duckduckgo, telegram, ntfy,
cryptoprice, read_webpage, internal_search, think).
Docs are now wrapped per chunk in <document index>/<source>/<content>
tags for citation-by-title support.
- Drop multimodal content for non-OpenAI-family providers (Google/Anthropic
raised on OpenAI image_url parts) so multimodal requests degrade to text
instead of returning 500.
- Stateless continuations (no conversation_id) default save_conversation to
False, avoiding an orphan conversation with an empty question on every tool
round (stateful continuations still default True).
- Only strip leaked reasoning from content for structured requests
(response_format / json_schema / json_object); legitimate answers that mention
the marker text are no longer corrupted.
- Forward sampling params (temperature, max_tokens, ...) on continuation turns.
- An explicit json_object request clears an agent-configured json_schema so it
isn't silently overridden.
- Drop max_tokens when max_completion_tokens is also sent (OpenAI rejects both).
- Don't build a chatcmpl-None completion id from the placeholder "None" id event.
OpenAI-compatible clients send multimodal user turns as a `content` array of
typed parts. translate_request previously assigned the array straight to the
question, breaking the string-only retrieval / token-budgeting / history paths
(HTTP 500). Now:
- content_to_text() extracts text from content arrays for the question,
history and system prompt, so the string paths work unchanged.
- The full content array (text + image_url parts) is preserved as
`multimodal_content`, threaded to the agent and emitted as the final user
message so images reach the model. Token budgeting uses the text only.
The content array (incl. image_url) now reaches the LLM call intact; images
render for vision-capable models. A text-only upstream model will reject the
image_url variant, as expected.
- Stateless tool continuation. OpenAI-compatible clients (opencode, etc.)
resend the full messages array — system, user, assistant(tool_calls),
tool(results) — but no conversation_id, so the prior
"conversation_id required for tool continuation" 400 broke every tool call.
When no conversation_id is present, rebuild the agent + pending tool calls +
tool results directly from the resent messages
(StreamProcessor.build_continuation_from_messages) instead of loading
server-side pending_tool_state, and call gen_continuation.
- Forward OpenAI sampling params (temperature, max_tokens,
max_completion_tokens, top_p, frequency_penalty, presence_penalty, stop,
seed) from the request to the LLM gen call; the agent otherwise uses its
configured defaults.
Make the OpenAI-compatible Chat Completions endpoint honor per-request
Structured Outputs and keep it OpenAI-compatible.
- translate_request now forwards the request's `response_format` (json_schema)
or a `response_schema` convenience field to the agent as its json_schema,
overriding the agent-configured schema for that request.
- Honor `response_format.json_schema.strict` (default true); strict:false
passes the schema through without forcing additionalProperties:false /
all-required (OpenAI's lenient mode).
- Support `response_format {"type":"json_object"}` via the provider's native
JSON mode.
- Keep `content` clean: some models echo their reasoning into content as
stringified `{'type': 'thought', ...}` reprs when response_format is set.
Strip those from content and reroute them to `reasoning_content` (OpenAI
never puts reasoning in content), for both streaming and non-streaming.
Flow: translate_request -> StreamProcessor._configure_agent ->
Agent.json_schema / json_schema_strict / json_object -> _llm_gen ->
prepare_structured_output_format(schema, strict).
* feat: implement WorkflowAgent and GraphExecutor for workflow management and execution
* refactor: workflow schemas and introduce WorkflowEngine
- Updated schemas in `schemas.py` to include new agent types and configurations.
- Created `WorkflowEngine` class in `workflow_engine.py` to manage workflow execution.
- Enhanced `StreamProcessor` to handle workflow-related data.
- Added new routes and utilities for managing workflows in the user API.
- Implemented validation and serialization functions for workflows.
- Established MongoDB collections and indexes for workflows and related entities.
* refactor: improve WorkflowAgent documentation and update type hints in WorkflowEngine
* feat: workflow builder and managing in frontend
- Added new endpoints for workflows in `endpoints.ts`.
- Implemented `getWorkflow`, `createWorkflow`, and `updateWorkflow` methods in `userService.ts`.
- Introduced new UI components for alerts, buttons, commands, dialogs, multi-select, popovers, and selects.
- Enhanced styling in `index.css` with new theme variables and animations.
- Refactored modal components for better layout and styling.
- Configured TypeScript paths and Vite aliases for cleaner imports.
* feat: add workflow preview component and related state management
- Implemented WorkflowPreview component for displaying workflow execution.
- Created WorkflowPreviewSlice for managing workflow preview state, including queries and execution steps.
- Added WorkflowMiniMap for visual representation of workflow nodes and their statuses.
- Integrated conversation handling with the ability to fetch answers and manage query states.
- Introduced reusable Sheet component for UI overlays.
- Updated Redux store to include workflowPreview reducer.
* feat: enhance workflow execution details and state management in WorkflowEngine and WorkflowPreview
* feat: enhance workflow components with improved UI and functionality
- Updated WorkflowPreview to allow text truncation for better display of long names.
- Enhanced BaseNode with connectable handles and improved styling for better visibility.
- Added MobileBlocker component to inform users about desktop requirements for the Workflow Builder.
- Introduced PromptTextArea component for improved variable insertion and search functionality, including upstream variable extraction and context addition.
* feat(workflow): add owner validation and graph version support
* fix: ruff lint
---------
Co-authored-by: Alex <a@tushynski.me>
* feat: Implement model registry and capabilities for multi-provider support
- Added ModelRegistry to manage available models and their capabilities.
- Introduced ModelProvider enum for different LLM providers.
- Created ModelCapabilities dataclass to define model features.
- Implemented methods to load models based on API keys and settings.
- Added utility functions for model management in model_utils.py.
- Updated settings.py to include provider-specific API keys.
- Refactored LLM classes (Anthropic, OpenAI, Google, etc.) to utilize new model registry.
- Enhanced utility functions to handle token limits and model validation.
- Improved code structure and logging for better maintainability.
* feat: Add model selection feature with API integration and UI component
* feat: Add model selection and default model functionality in agent management
* test: Update assertions and formatting in stream processing tests
* refactor(llm): Standardize model identifier to model_id
* fix tests
---------
Co-authored-by: Alex <a@tushynski.me>