78 Commits
Author SHA1 Message Date
arc53-machine 26bd4e2a3f Keep the GitHub loader tests off the network
With GITHUB_ACCESS_TOKEN set in the environment, load_data checked the
repository's visibility against the real GitHub API. The tests now clear
the instance token unless they set one.
2026-09-29 15:09:50 +01:00
arc53-machine f2349edb4c Keep a Linear sync going when one issue or document cannot be read
An issue deleted after it was listed, or comments the token cannot read,
now lose that detail with a warning instead of failing the whole sync; an
unreadable document is skipped.
2026-09-29 12:14:10 +01:00
arc53-machine 03d1b10897 Sync Linear issues and documents into Knowledge with the Linear sign-in
The Linear connector now syncs as well as giving agents its tools, from
one connection. Linear's MCP server is its own OAuth issuer, so its
tokens are read through the same MCP tools the agents use (list_issues,
get_issue, list_comments, list_documents) rather than Linear's GraphQL
API, and no OAuth app has to be registered.

A source picks teams and projects, with comments (on by default) and the
projects' documents. Each issue becomes one document with its state,
assignee, priority, labels, description and comments, filed under its
team and citing its Linear URL. Each sync reads up to 500 issues and 100
documents again. /api/connections/<id>/linear lists the teams and
projects to pick from. Sources sync on their schedule with the owner's
connection, and pause when the sign-in needs reconnecting.
2026-09-29 11:32:40 +01:00
arc53-machine 7b516ade82 Add a built-in GitHub connector for repository sync and read-only tools
One GitHub connection feeds both a Knowledge source and an agent tool.
Users connect with a personal access token, checked against GitHub and
named after the account, or, when an admin registers a GitHub App
(GITHUB_CLIENT_ID, GITHUB_CLIENT_SECRET, GITHUB_APP_SLUG), with Sign in
with GitHub. App tokens expire after eight hours and are refreshed before
each sync or tool call; tokens with no expiry are never treated as expired.

The catalog gains an optional second sign-in method (oauth_settings,
exposed as sign_in_methods) without changing any other connector.

Sync lists the repositories the connection can read (the token's own, or
the App installations') and ingests one with the connection's token. The
tool is GitHub's read-only MCP server (api.githubcopilot.com/mcp/readonly):
setup discovers its actions and creates it bound to the connection, and
the executor sends the connection's token only to that server.
2026-09-29 10:41:03 +01:00
arc53-machine 3572d968ec Read private repositories only with the user's own GitHub token
GITHUB_ACCESS_TOKEN belongs to the server, but any user could ingest any
repository it can see, private ones included. It is now used only for
public repositories; the loader checks the repository's visibility first
and asks for a GitHub connection otherwise.

The loader also takes a token from the connection a source syncs from.
Merging a connection's keys into a plain repository URL no longer fails
on json.loads, manual Sync now passes the source's connection, and a
token GitHub rejects pauses the connection's sources for reconnect.
2026-09-29 10:41:03 +01:00
arc53-machine 6b4220d401 Address review: MCP token scope, API-key identity, reconnect lock, setup retries
- An MCP connection's tokens only go to its own server, and a
  connection_id can no longer come from a client: the MCP test/save
  routes drop it and the tool executor uses only the resolved one.
- API-key connections are reused only when they hold the same
  credentials. Two keys that share a hint or a label get separate
  connections ("…abcd (2)"), at runtime and in the 0038 backfill.
- Reconnecting a connection whose secrets cannot be decrypted no longer
  flags it from a second transaction while holding its row lock.
- Connection setup validates the sync request before claiming its
  Idempotency-Key, so a corrected retry is queued.
- A request that names no connection resolves to none without opening
  a database connection; fixes two tests that reached a real database.
2026-09-28 18:45:27 +01:00
arc53-machine 3e11442a83 Merge remote-tracking branch 'origin/main' into connectors
# Conflicts:
#	frontend/DESIGN.md
#	frontend/src/locale/de.json
#	frontend/src/locale/en.json
#	frontend/src/locale/es.json
#	frontend/src/locale/jp.json
#	frontend/src/locale/ru.json
#	frontend/src/locale/zh-TW.json
#	frontend/src/locale/zh.json
#	frontend/src/upload/Upload.tsx
2026-09-28 18:07:12 +01:00
arc53-machine 3511004f50 Connections own their credentials
Migration 0038 moves every stored secret (OAuth tokens, MCP OAuth
tokens and client registrations, API keys) into the connection's
encrypted envelope, links API-key tools to one connection per distinct
credential, allows several accounts per provider, and adds
credential_mode to sources and tools. OAuth MCP tools keep resolving
each member's own token, as they did before.

docsgpt.connectors.service is now the only reader of OAuth tokens:
get_valid_token_info refreshes under a row lock and persists rotated
refresh tokens, and a revoked grant flags the connection, pauses its
sources and notifies the owner. Loaders build from a connection
(BaseConnectorLoader.from_connection), so scheduled sync covers Drive,
SharePoint and Confluence sources with no browser. S3 and Reddit keys
stay on the connection instead of in remote_data.

New endpoints: POST /api/connections, /setup, /reconnect,
/picker-token, /claim, DELETE /api/connections/<id>, per-action
permissions and MCP refresh-tools. Upload, file listing, sync and
validate-session take a connection_id; session tokens keep working for
this release. The tool executor reads credentials from the resolved
connection (owner or member mode) and pauses on a Connect card when a
connection needs signing in. docsgpt connectors reencrypt rewrites
stored credentials after a key rotation.
2026-09-28 17:03:09 +01:00
Pavel f4331cd3a3 rabbit fixes 2026-09-28 18:40:49 +04:00
Pavel 706a0cb2b2 Big source revamp 2026-09-28 17:35:16 +04:00
arc53-machine f882ef49a7 refactor: read settings directly instead of getattr with a second default
About 85 call sites read a setting as getattr(settings, "NAME", fallback),
each carrying its own copy of the default. Every one of those names is a
field with a default on the model, so the fallback could never apply to
the real settings object; it only masked drift. Two had drifted:

- OPENAI_PROMPT_CACHE_KEY defaults to True on the model but the reader
  fell back to False, and two test stubs relied on that.
- SharePoint's MICROSOFT_AUTHORITY fallback to
  https://login.microsoftonline.com/<tenant> never fired, because the
  attribute always exists (as None), so MSAL got authority=None. The
  connector now derives the tenant authority when the setting is unset,
  as its test always assumed.

Four places read EMBEDDINGS_KEY straight from os.environ, skipping the
"None"/"" normalisation the model applies; they read the setting now.
Test stubs that replaced a module's settings with a SimpleNamespace list
every setting the code under test reads.
2026-09-17 11:14:34 +01:00
Alex 7da46c2bea feat: air-gapped deployment guide, no implicit downloads
- Ship tiktoken's cl100k_base inside the package and build the encoding
  from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
  temp dir, and read tokenizer.json and repo metadata from that cache, so
  a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
  the endpoints return 404, audio files fail to ingest with a clear
  message, /api/config reports tts_available/stt_available, and the UI
  hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
  packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
2026-09-15 17:54:24 +01:00
Alex ad201f8318 fix: refuse oversized TIFF/BMP attachments before converting them
A deflate-compressed TIFF under 1 MB can declare 144 million pixels and
take 1.2 GB to convert to PNG, and Pillow only warns below 179 million.
Read the dimensions from the header and refuse images over 40 million
pixels before any pixel data is decoded. Pillow's DecompressionBombError
is now raised as DocumentParseError, so the upload fails once instead of
being retried.
2026-09-14 17:45:43 +01:00
Alex 3030aba39d fix: handle non text uploads more carefully 2026-09-14 17:27:13 +01:00
Alex b3326b3c33 chore(deps): redis 8, tiktoken 0.14, openapi3-parser 2, daytona 0.211, reportlab 5
Major bumps whose ceilings had to move. redis 8.1.0, tiktoken 0.14.0 and
daytona 0.211.2 needed no code change (the Daytona client, filesystem and
process signatures the sandbox calls are unchanged; redis 8 was checked
against a live server through the app's own sync and async clients).
reportlab 5.0.1 is test-only.

openapi-parser 2.0.0 is a rewrite onto pydantic spec models: `paths` is now
a dict keyed by URL rather than a list of objects carrying their own `url`,
and a path item exposes one field per HTTP method instead of an `operations`
list. `OpenAPI3Parser` reads both accordingly, iterating methods in the
order the spec declares them, and its rendered output is byte-identical to
before. The rewrite also drops prance, openapi-spec-validator and five more
transitive packages.

tokenizers stays at 0.22.2: transformers 5.8.1 caps it at <=0.23.0 and no
such release exists, so it moves with the transformers cap or not at all.
2026-09-12 17:28:58 +01:00
Alex 7e00197678 chore(deps): upgrade backend deps within their declared ranges
`uv lock --upgrade` plus the two code changes the new versions need.

firecrawl-anydoc 0.2.4 raises a dedicated `NeedsOcrError` where 0.2.3 raised
`UnsupportedError("... OCR is required")`, so the anydoc parser no longer
recognised a scanned PDF: the fallback still ran, but a near-empty result was
stored as an empty document instead of failing with the OCR_ENABLED hint.
`_needs_ocr` now accepts both spellings and looks the class up lazily, so an
older anydoc keeps working. 0.2.4 also refuses the CID-font NDA fixture
outright rather than dropping its Chinese column silently, so the PDF
trust-check tests stub that dropped output against the fixture's real bytes
(the check's own inputs) and a new test pins the refusal path.

ruff 0.16 widened its implicit default rule set, turning the dev-group bump
into 7131 findings across the tree. `.ruff.toml` now states the historical
selection (E4, E7, E9, F) explicitly and the CI pin moves to the locked
0.16.7, so lint no longer drifts with the version.
2026-09-12 17:16:59 +01:00
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00
Alex 6860a21541 fix: address review on the slim-image branch
- The frontend image ran the Vite dev server in development mode, so
  .env.development supplied its defaults (notification banner, Google client
  id, local API host). The static build only loads .env.production, so the
  build stage now copies .env.development in as the baseline and the compose
  files pass every VITE_* the app reads through from .env; the runtime script
  skips empty values so a blank passthrough keeps the build-time default.
  .dockerignore kept only the .local variants out.
- VITE_DISABLE_SOURCE_FE disables sources only when it is the string true.
- DoclingParser: find_spec raises when docling itself is absent; the install
  hint now covers that path, with a regression test.
- verify_offline: direct tests for verify(); the PR image check builds and
  verifies the -docling variant as well as slim.
- Workflows this branch adds or rewrites pin actions by commit, pass the
  release tag through env instead of template expansion, and do not persist
  checkout credentials.
- OCR guide no longer claims pre-built images never include docling.
2026-09-06 21:44:34 +01:00
Alex 4707c45b93 fix(parser,vectorstore): stop the first-request downloads in a warmed install
Three things still reached the network from a container whose models were
baked in:

- tiktoken fetched cl100k_base from openaipublic.blob.core.windows.net on
  every fresh container (its cache defaulted to /tmp), and token accounting
  calls it on every chat. prefetch_models now warms it too; the image sets
  TIKTOKEN_CACHE_DIR.
- The chunker loaded its tokenizer with Tokenizer.from_pretrained, which
  revalidates the revision with a HEAD request per process start and stalls
  for the etag timeout (10 s) when huggingface.co is unreachable. It now reads
  tokenizer.json from the hub cache first and only downloads on a miss; the
  repo-metadata read for models outside the registry does the same.
- tldextract fetched the public suffix list on the first web crawl; the
  bundled snapshot is used instead.

application/scripts/verify_offline.py exercises these paths (and docling's
conversion when the extra is installed) so an image can be checked with
docker run --network none.
2026-09-05 15:50:20 +01:00
Alex 3947c66cda Merge branch 'main' into anydoc-support
Conflicts, and how each was taken:

- application/core/settings.py — ours. The renamed OCR_ENABLED /
  OCR_ATTACHMENTS_ENABLED / OCR_MIN_CHARS_PER_PAGE accept main's
  DOCLING_OCR_* spellings as AliasChoices, so nothing is dropped.
- application/Dockerfile — both. Main's install layers plus the
  INSTALL_DOCLING build arg.
- application/parser/file/constants.py — both imports.
- deployment/docker-compose.yaml — both. The INSTALL_DOCLING /
  INSTALL_TESSERACT build args on backend and worker, and main's
  -Q docsgpt,parsing,embeddings, which query embedding needs.
- tests/conftest.py — theirs. Both sides fixed the same pytest-postgresql
  9.0.0 autocommit= breakage; main's spelling is the one already on main.
- application/requirements.txt — the comments claimed different reasons
  torch is in core. Main's is the true one now: it removed
  sentence-transformers, so docling is torch's only remaining consumer.

Two things the merge broke without conflicting:

- onnxruntime. This branch moved it out of core into the docling extra;
  main meanwhile made it the runtime local embeddings execute on
  (fastembed). Git took the deletion, leaving fastembed with no pinned
  runtime in a repo that pins everything. Restored to core, and no longer
  pinned twice from the extra.
- The frontend copy of ATTACHMENT_PARSER_EXTENSIONS. The backend list is
  derived and picked up the anydoc suffixes; the hand-kept frontend mirror
  did not, so the composer would refuse files the API accepts.
  tests/parser/file/test_constants.py is what caught it.

ruff, pytest (9897 passed), frontend build and docs build all pass. The
image build is unverified: no Docker daemon on this machine.
2026-09-04 16:42:48 +01:00
Pavel 4e14a79923 Fixes batch 2 2026-09-04 13:38:29 +04:00
Alex f7cd94668f fix(attachments): close the .txt gap in the parseability gate
Review follow-ups on the attachment gate.

.txt was listed as parser-backed, but it has no parser — it *is* the
plain-text fallthrough. That let it skip the content check, so renaming a
video to notes.txt walked straight back into the bug the gate exists for
(verified: 5132 chars of binary "extracted" and stored). The list is now
exactly the file extractor's keys, .txt included in the content check like
any other unparsed suffix, and the drift test asserts equality rather than
containment. The sniff now recognises a UTF-16/32 BOM as text, so a
Notepad "Unicode" .txt is not caught by the NUL-byte rule.

The picker's accept filter listed parser-backed suffixes only, hiding .txt,
.py and .log — files the gate reads happily — behind "All files". It now
carries text/* as well, so it can never be narrower than what the upload
accepts.

A rejected batch carries one errors entry per file, but the non-200 branch
applied the top-level message to every chip, so two files failing for two
reasons both reported the first one. Reasons are now matched by
upload_index, with the top-level message as fallback.

_get_store_attachment_user_error no longer reads str(exc): the
unsupported-type message is rebuilt from the filename, so no exception
state can reach a response body (CodeQL py/stack-trace-exposure).
2026-09-03 09:13:24 +01:00
Alex 96cad1f8e2 fix(attachments): refuse unparseable chat attachments
A chat attachment with no parser fell through to SimpleDirectoryReader's
plain-text open(), so a phone-uploaded video was "extracted" into megabytes
of binary garbage, truncated, and stored with extraction.status == "ok".

Gate attachments in two tiers instead. A suffix with a dedicated parser is
admitted on its name — a PDF is binary and parses fine. Anything else has to
read as text: the first 8KB are sampled and refused on a NUL byte or too many
other control bytes. That keeps source, config and log files working through
the plain-text fallthrough, and keeps out videos, archives and renamed
binaries alike. The route checks the staged spool before anything is stored
or queued; the worker repeats the check where the local file exists, raising
the non-retryable AttachmentRejectedError.

SUPPORTED_ATTACHMENT_EXTENSIONS gains the parser-backed suffixes it was
missing (.tiff, .tif, .bmp, .webp, .vtt, .xml) and is now exactly the file
extractor's keys plus .txt, with a test asserting the two agree. The composer
applies the same rule client-side, so an unsupported file is named before it
costs an upload, and a test pins the frontend list to the backend one.

Attachment failures now show their reason inline under the chips rather than
only in a hover tooltip, which a touch user can never see, and only after a
send was attempted. Dropped `accept` from the dropzone: it discarded rejected
drops with no feedback and disagreed with the server about text files.
2026-09-03 08:52:12 +01:00
Pavel ed0892b39b Batch fixes 2 2026-09-03 00:30:59 +04:00
Pavel 47eb92fc42 Fixes batch 1 2026-09-02 23:42:35 +04:00
Pavel 96b878217d standalone OCR 2026-09-02 23:12:36 +04:00
Pavel d4e92be0ab ocr update 2026-08-27 23:46:12 +04:00
Alex de22be5a21 fix: chunk-budget blowups, FastEmbed built-ins, and re-embed gaps
Follow-up review pass over the embeddings branch.

- Fold an oversized header back into the body, and drop header duplication
  when it would leave under a quarter of the chunk budget. A header at or
  over max_tokens collapsed the body budget to one token, so a document
  became one chunk per body token, each still over the cap: a 95 KB file
  produced 20k chunks of 2563 tokens against a 1250 cap. Also clamp
  max_tokens to at least 1, as the strategy chunkers already do.
- Emit a header-only document as its own chunk. With no body piece to
  attach it to, splitting returned nothing and the document was dropped
  from the index with no error and no log line.
- Skip add_custom_model for a repository FastEmbed already ships. It
  rejects a name it knows, so configuring any of its ~30 built-ins
  (MiniLM, bge, e5, gte, ...) failed every embed call and every query.
- Decide "the user chose this model" by comparing against the field
  default rather than model_fields_set, which is true for anything read
  from .env. Every setup script has always written EMBEDDINGS_NAME, so an
  upgraded remote-embeddings install inherited mpnet's 384-token window
  and silently clipped ~80% off every chunk.
- Cut tiktoken splits at character offsets instead of decoding each token
  window. A multi-byte character straddling a boundary decoded to U+FFFD
  on both sides, destroying one character at roughly one boundary in five
  on CJK text -- including at the default max_tokens of 2000.
- Let the re-embed script open a FAISS index whose width does not match
  the configured model. That mismatch is the main reason to run it, and
  the error recommending the script was raised by the script itself, so
  the advice failed on every source.
- Re-embed graph_nodes.name_embedding when GraphRAG is enabled. Those
  vectors seed every traversal and share the chunk vectors' width, so a
  same-width model swap left the graph retrieving from the old space with
  nothing to report it.
- Prefetch the models before copying the application source, so editing
  any file no longer re-downloads ~780 MB of artifacts on every build.
- Mirror the setup.sh embedding menu into setup.ps1: granite default,
  legacy mpnet as an explicit option, and both engine flows updated.
  Windows users were otherwise stranded on mpnet with no granite path.
- Drop the unused EmbeddingsWrapper.tokenizer property.
2026-08-27 15:55:21 +01:00
Pavel d7a7d4d084 docling separation 2026-08-27 17:36:38 +04:00
Alex 25e07f5cee fix: embeddings registry edge cases and chunk-size accounting
Follow-up to the embeddings work, from a review pass over the branch.

- Route the OpenAI/Azure key handling through the model registry instead of
  matching the canonical name literally, so the `text-embedding-ada-002`
  alias the registry now accepts also reaches the Azure deployment name
  rather than failing every embed with DeploymentNotFound.
- Fall back to a default width where the embeddings model reports no
  dimension. A model outside the registry returns None rather than no
  attribute, so `getattr` with a default did not catch it and the width
  reached the DDL as `vector(None)` / `list_size=None`.
- Point HF_HUB_CACHE at the prefetch directory. Chunking loads the tokenizer
  through `tokenizers`, which reads the hub cache, so a fresh container
  fetched over the network on first ingest and an offline one silently fell
  back to cl100k.
- Charge a token that collapses a long unbroken run by its character span.
  WordPiece emits one [UNK] for any word over its character limit, which made
  base64 and minified content count as near-zero tokens, so nothing split it
  and oversized chunks reached the embedding server.
- Preserve chunk ids and honour --batch-size when rebuilding a FAISS index.
  Fresh uuids orphaned GraphRAG's graph_node_chunks rows, and the whole index
  went out in a single embed call on remote servers.
- Document that granite runs an int8-quantised graph, and scope the
  SentenceTransformer parity claim to mpnet's fp32 graph, which is where it
  was measured.
- Correct the embeddings docs: a matching dimension is not a matching model,
  so a same-width swap raises nothing and silently degrades retrieval.
2026-08-27 13:59:51 +01:00
Pavel e101c0a7f1 Implement anydoc with docling 2026-08-27 13:51:09 +04:00
Alex dd3876fdcb fix: mini fixes 2026-08-26 16:37:03 +01:00
Alex 8380f9bb47 feat: optimise embeds 2026-08-26 14:46:19 +01:00
Alex e0aff39a1b feat: ingestion optimisations 2026-08-25 12:27:09 +01:00
Alex e6fb46560c fix: minor graph rag improvements 2026-08-20 13:33:06 +01:00
Alex 0b36257202 feat: parser improvements and fixes 2026-08-19 23:49:26 +01:00
Alex 7ad2f38512 feat: faster attachment processing 2026-08-13 16:04:45 +01:00
Alex 0a15ce8fbb feat: attachment provenance 2026-08-13 14:30:28 +01:00
Alex ca53a3ac4b chore: bump docling 2026-08-12 09:05:00 +01:00
Alex ca345e3bf8 feat: improve tool call durability on long calls 2026-08-10 14:32:30 +01:00
Alex a166796f6c feat: prune some deps 2026-08-10 12:10:54 +01:00
Alex 0ecb421954 fix: error type fixes and docling parsing improvements 2026-08-06 12:29:51 +01:00
Alex 15b6b03b47 fix: minor size cleanup 2026-08-04 15:07:54 +01:00
Alex a1ad802e22 feat: limit docling attachment file sizes explicitly 2026-08-04 14:04:44 +01:00
Alex e4c0c26927 fix: more sandbox protections to terminate plus max pages 2026-07-09 19:56:13 +01:00
Alex 66786c2760 fix: remove code exec as default and sec improvements 2026-07-08 18:50:10 +01:00
Alex 94a845aa82 fix: more artefact hardening 2026-07-04 11:42:27 +02:00
Alex 75e07af78f fix: minor issue fixes 2026-06-30 18:36:06 +02:00
Alex 37d93cbd86 Parse documents on a Celery parsing worker via a read_document tool
Replace the sandbox Docling extractor with read_document, backed by the in-process
backend parser (the same one ingestion uses) and offloaded to a dedicated
'parsing' Celery queue so it can run on GPU-capable workers with predictable RAM.
The tool resolves the input ref under the run-scoped gate, enqueues the parse,
and awaits it with a timeout (degrading to an error rather than hanging); the
worker independently re-resolves the artifact through the same gate and never
trusts a raw path. Untrusted files get the upload path's safeguards (extension
whitelist, size cap, sanitized temp file, cleanup). Options: output
(markdown/text/structured/chunks), ocr, pages, engine, max_chars, include_tables,
persist, json_schema. The workflow native-file 'extract' fallback now uses the
same worker path, so document parsing no longer needs the sandbox and works on
every backend.

Also fixes the branch's periodic-task test (the sandbox reaper made it 12) and
points the dev and e2e Celery workers at the parsing queue.
2026-06-25 13:24:12 +01:00
Alex a709b7ccdc feat: semantic chunking + pgvector hybrid (BM25+vector) retriever (#2553) 2026-06-22 17:04:13 +01:00