Files
DocsGPT/application/storage
Alex 3242d68fd2 feat: pin the embedding model to the installation, not the release
EMBEDDINGS_NAME has a code-level default, and moving that default re-points an
existing index at a different vector space without anything noticing: mpnet and
granite are both 768-dimensional, so no width check fires and retrieval simply
gets worse. Which model an index was built with is a property of the
installation, not of the release it happens to be running.

So it is resolved once at boot and stored in app_metadata, which already exists
for exactly this kind of one-off state and needs no migration. An installation
that already has sources is pinned to the legacy model it has been using all
along and told, once, how to move; an empty one is pinned to the current
recommendation. An explicit EMBEDDINGS_NAME in the environment still wins over
both, and an unreachable database falls back to the code default.

This also closes a gap in what "new installs get granite" meant: it held for
setup.sh and for anyone copying .env-template, but a hand-written .env plus
docker compose fell through to the settings default and quietly got mpnet --
English-only, and a 384-token window against a 1250-token chunk default.

Emptiness is counted in sources rather than vector rows so the answer is the
same for every vector store, including the FAISS ones whose vectors are not in
this database at all. Concurrent workers converge through the same
INSERT ... ON CONFLICT DO NOTHING the instance id already uses.
2026-08-28 13:14:35 +01:00
..
2026-08-12 12:36:43 +01:00
2026-08-12 12:36:43 +01:00