Commit Graph
49 Commits
Author SHA1 Message Date
Alex a166796f6c feat: prune some deps 2026-08-10 12:10:54 +01:00
Alex c5b372df1e fix: pin rapidocr and onnxruntime 2026-08-10 10:37:43 +01:00
Alex 795e39a6bc fix: source authorization, silent retrieval failures, and prompt structure
Source access control
---------------------
`active_docs` is client-supplied and reached the retriever unchecked, and the
retriever queries `WHERE source_id = <id>` with no owner predicate — so any
caller could pass any source id to /stream or /api/answer and have another
tenant's documents quoted back, while /api/sources/<id>/search correctly
refused the same id. Gate it through `can_access`, the helper the guarded
endpoints already use, and filter `self.source` down to the authorized set.
Fails closed: no principal, or a check that errors, drops the source.

Three sibling paths had the same gap:

- workflow agent nodes: `AgentNodeConfig.sources` is written verbatim from
  client JSON at save time and nothing validated it, so a node could name any
  tenant's source. Gate against the workflow owner, so shared workflows keep
  reading their owner's sources like shared agents do.
- /api/share: `_resolve_source_pg_id` resolved any id with no ownership
  predicate and baked it into the agent the share creates; /api/search then
  searched it. Authorize before attaching.
- search_service: re-resolve the ids stored on an agent row instead of
  trusting them, so a row written by any future path with the same gap cannot
  be read back.

Team grantees previously lost their source's retrieval config: the post-check
read was still owner-scoped, so it missed and fell back to defaults (an
`agentic_tool` source was bulk-prefetched for every grantee). Read unscoped
after `can_access` passes.

Retrieval
---------
`PGVectorStore._ensure_table_exists` created an IVFFlat index on the empty
table it had just created. IVFFlat computes centroids at build time, so those
centroids were random, and combined with the `source_id` post-filter a source
with hundreds of embedded chunks returned zero rows — retrieval reported no
documents, the model answered from memory, and nothing was logged. Stop
creating the index (exact search is correct and fast well past the sizes most
deployments reach); raise `ivfflat.probes` to sqrt(lists) where an index still
exists; and re-run a short indexed search exactly, since post-filtering means
no index setting can guarantee a full result. `graphrag` had the same
empty-table index with no fallback at all.

Also: bound `chunks` to 0-500 on both the request and agent paths (0 still
means "skip retrieval"), let a source's configured `retrieval.chunks` outrank
the request body, and cap ClassicRAG's per-source floor at
max(top_k, n_sources) so attaching sources cannot inflate the result set.

Silent failures
---------------
An empty retrieval was invisible to both the model and the client: the `source`
event was suppressed when the list was empty, so "searched and found nothing"
looked identical to "no source attached", and the prompt said nothing at all.
Emit the event always, and tell the model when a search ran and returned
nothing. A file that parses to nothing now fails ingest with a message naming
the cause instead of storing an embedding of the empty string. `score_threshold`
returns warnings when the active store or retriever cannot honour it.

Prompt structure
----------------
Retrieved documents move from the system prompt into the user turn, with the
injection guard restated next to them: they change every turn (defeating prefix
caching), they are third-party text that should not carry system authority, and
routing them through the query budget makes them truncatable rather than
silently crowding it out. Documents are shed lowest-ranked-first before the
question is touched.

The six chat presets (3 tones x 2 retrieval modes) differed only in their
Answering section; they are now composed from single-source fragments at load
time, not through Jinja inheritance, which would have opened a file-read
surface in the template sandbox and broken the tool-prefetch parser. Per-tool
guidance moves out of the prompt into tool schemas, so it travels with the tool
and cannot render when the tool is absent. A plain-text custom prompt is staged
as a persona value inside the skeleton instead of replacing it wholesale — it
used to silently lose the injection guard, platform block, memory and
attachments, and its braces are now inert.

Other fixes
-----------
- agents/base: an oversized system prompt drove the query budget negative and
  dispatched a full-price request with an empty question; raise instead.
- llm/anthropic: migrate off the retired Text Completions API. It flattened
  history to first+last message and ignored tools entirely. Adds the missing
  Anthropic handler, without which every tool call was silently dropped.
- sources/upload: `sitemap` had no branch, so every sitemap ingest died on a
  TypeError; `validate_url` now rejects a falsy URL cleanly.
- workflow nodes: retrieved documents never reached the node agent, so a
  classic node with a source and an ordinary prompt answered "I have no
  documents" while the run reported completed.
- parser/bulk: copy the metadata dict, or every chunk reports the last chunk's
  token_count.
- crawler_loader: carry the page title, or citations render the whole chunk
  body as the label.
2026-08-08 10:21:52 +01:00
PavelandAlex aaad51f951 Harden protection with pinned requests and path-param encoding (#2486)
* Harden protection with pinned requests and path-param encoding

* fix: domain pinning

* fix: tests

* fix: html test 2

---------

Co-authored-by: Alex <a@tushynski.me>
2026-05-23 02:33:29 +01:00
Alex e167cf8247 fix: broken syncs (#2480)
* fix: broken syncs

* fix: mini fixes
2026-05-17 23:58:28 +01:00
81b6ee5daa Pg 4 (#2390)
* feat: postgres tests

* feat: mongo cutoff

* feat: mongo cutoff

* feat: adjust docs and compose files

* fix: mini code mongo removals

* fix: tests and k8s mongo stuff

* feat: test fixes

* fix: ruff

* fix: vale

* Potential fix for pull request finding 'CodeQL / Clear-text logging of sensitive information'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* fix: mini suggestions

* vale lint fix 2

* fix: codeql columns thing

* fix: test mongo

* fix: tests coverage

* feat: better tests 4

* feat: more tests

* feat: decent coverage

* fix: ruff fixes

* fix: remove mongo mock

* feat: enhance workflow engine and API routes; add document retrieval and source handling

* feat: e2e tests

* fix: mcp, mongo and more

* fix: mini codeql warning

* fix: agent chunk view

* fix: mini issues

* fix: more pg fixes

* feat: postgres prep on start

* feat: qa tests

* fix: mini improvements

* fix: tests

---------

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
Co-authored-by: Siddhant Rai <siddhant.rai.5686@gmail.com>
2026-04-18 13:13:57 +01:00
Alex a1efea81d0 patch: sitemap and web loader 2026-04-15 23:39:44 +01:00
Alex 502819ae52 feat: pg migration, more tables 2026-04-12 12:15:59 +01:00
Alex df57053613 feat: improve crawlers and update chunk filtering (#2250) 2026-01-06 00:52:12 +02:00
Alex 9e7f1ad1c0 Add Amazon S3 support and synchronization features (#2244)
* Add Amazon S3 support and synchronization features

* refactor: remove unused variable in load_data test
2025-12-30 20:26:51 +02:00
Alex aef3e0b4bb chore: update workflow permissions and fix paths in settings (#2227)
* chore: update workflow permissions and fix paths in settings

* dep

* dep upgraes
2025-12-25 14:26:01 +02:00
Alex 98e949d2fd Patches (#2218)
* feat: implement URL validation to prevent SSRF

* feat: add zip extraction security

* ruff fixes
2025-12-24 17:05:35 +02:00
Alex e0a9f08632 refactor and deps (#2184) 2025-12-10 23:53:59 +02:00
Alex 03452ffd9f feat: add GitHub access token support and fix file content fetching logic (#2032) 2025-10-07 16:53:14 +03:00
ManishMadan2882 f09f1433a9 (feat:connectors) separate layer 2025-08-26 01:38:36 +05:30
ManishMadan2882 2410bd8654 (fix:driveLoader) folder ingesting 2025-08-22 19:07:52 +05:30
ManishMadan2882 92d6ae54c3 (fix:google-oauth) no explicit datetime compare 2025-08-22 13:35:03 +05:30
ManishMadan2882 8c3f75e3e2 (feat:ingestion) google drive loader 2025-08-22 13:32:40 +05:30
ManishMadan2882 b2b04268e9 (feat:drive) oauth flow 2025-08-21 02:46:32 +05:30
Alex 481df4d604 fix: enhance error logging with exception info across multiple modules 2025-05-05 13:12:39 +01:00
Pavel fddee69f92 web loader fix
Changes web loader to the correct output.
2025-01-17 19:13:23 +03:00
Pavel 13fcbe3e74 scraper with markdownify 2025-01-15 01:08:09 +03:00
Alex 2245f4690e fix: reddit loader validation 2024-11-15 11:02:27 +00:00
devendra.parihar d3238de8ab fix: lint error 2024-10-18 12:23:17 +05:30
devendra.parihar 09a2705311 fix:GitHubLoader to Handle Binary Files 2024-10-18 12:08:08 +05:30
devendra.parihar a4c0861cf4 fix:GitHubLoader to Handle Binary Files 2024-10-18 12:07:44 +05:30
Alex 6932c7e3e9 feat: add filename to the top 2024-10-05 21:56:47 +01:00
Alex c04687fdd1 fix: github loader metadata clickable 2024-10-05 21:53:30 +01:00
Alex 7717242112 fix(lint): ruff var 2024-10-05 21:37:55 +01:00
Alex 1ad82c22d9 fix: headers 2024-10-05 21:36:04 +01:00
Alex 8fa88175c1 fix: translation + auth 2024-10-05 21:33:58 +01:00
Alex 2611550ffd 2024-10-02 23:44:29 +01:00
Alex c49b7613e0 fix: langchain warning 2024-08-31 12:53:37 +01:00
Siddhant Rai 53e86205ad fix: added more headers from default 2024-05-03 18:47:30 +05:30
Siddhant Rai aa670efe3a fix: connection aborted in WebBaseLoader 2024-05-03 18:25:01 +05:30
Siddhant Rai e01071426f feat: field to pass number of posts as a parameter 2024-03-27 19:20:55 +05:30
Siddhant Rai eed1bfbe50 feat: fields to handle reddit loader + minor changes 2024-03-26 16:07:44 +05:30
Siddhant Rai 60cfea1126 feat: added reddit loader 2024-03-16 20:22:05 +05:30
Pavel 54d187a0ad Fixing ingestion metadata grouping 2024-02-28 19:52:58 +03:00
Alex 0cb3d12d94 Refactor loader classes to accept inputs directly 2024-02-14 15:17:56 +00:00
Pavel 381a2740ee change input 2023-10-13 21:52:56 +04:00
Pavel 024674eef3 List check 2023-10-13 11:42:42 +04:00
Pavel b7d88b4c0f fix wrong link 2023-10-12 19:45:36 +04:00
Pavel 719ca63ec1 fixes 2023-10-12 19:40:23 +04:00
Pavel 2cfb416fd0 Desc loader 2023-10-12 13:44:32 +04:00
Pavel 50f07f9ef5 limit crawler 2023-10-12 12:53:33 +04:00
Pavel c517bdd2e1 Crawler + sitemap 2023-10-12 12:35:26 +04:00
Pavel 658867cb46 No crawler, no sitemap 2023-10-12 01:03:40 +04:00
Alex 8f2ad38503 tests 2023-10-11 10:13:51 +01:00