mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-04 16:13:23 +00:00
Follow-up to the embeddings work, from a review pass over the branch. - Route the OpenAI/Azure key handling through the model registry instead of matching the canonical name literally, so the `text-embedding-ada-002` alias the registry now accepts also reaches the Azure deployment name rather than failing every embed with DeploymentNotFound. - Fall back to a default width where the embeddings model reports no dimension. A model outside the registry returns None rather than no attribute, so `getattr` with a default did not catch it and the width reached the DDL as `vector(None)` / `list_size=None`. - Point HF_HUB_CACHE at the prefetch directory. Chunking loads the tokenizer through `tokenizers`, which reads the hub cache, so a fresh container fetched over the network on first ingest and an offline one silently fell back to cl100k. - Charge a token that collapses a long unbroken run by its character span. WordPiece emits one [UNK] for any word over its character limit, which made base64 and minified content count as near-zero tokens, so nothing split it and oversized chunks reached the embedding server. - Preserve chunk ids and honour --batch-size when rebuilding a FAISS index. Fresh uuids orphaned GraphRAG's graph_node_chunks rows, and the whole index went out in a single embed call on remote servers. - Document that granite runs an int8-quantised graph, and scope the SentenceTransformer parity claim to mpnet's fp32 graph, which is where it was measured. - Correct the embeddings docs: a matching dimension is not a matching model, so a same-width swap raises nothing and silently degrades retrieval.