mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-03 15:11:30 +00:00
Changing EMBEDDINGS_NAME on a populated index is the one failure the width check cannot catch. Two models of the same width -- mpnet and granite are both 768 -- swap without raising anything, and every query is then embedded by a different model than the stored vectors were. Nothing fails; answers just get worse. Boot now compares what each source was built with against the active model and names the mismatched sources and the command that fixes them. The comparison goes through the registry rather than string equality, so a stored alias is not read as a different model. A source with no recorded model pre-dates the column and is therefore the legacy model, not unknown. That check is only as good as sources.model, which reembed was not maintaining: it rewrote the vectors and left the column naming the old model, so a source would be reported stale immediately after being migrated. It is now stamped after each source succeeds, from its own session -- sources lives in the user-data database while the vectors may not.