Files
DocsGPT/tests/scripts
Alex 25e07f5cee fix: embeddings registry edge cases and chunk-size accounting
Follow-up to the embeddings work, from a review pass over the branch.

- Route the OpenAI/Azure key handling through the model registry instead of
  matching the canonical name literally, so the `text-embedding-ada-002`
  alias the registry now accepts also reaches the Azure deployment name
  rather than failing every embed with DeploymentNotFound.
- Fall back to a default width where the embeddings model reports no
  dimension. A model outside the registry returns None rather than no
  attribute, so `getattr` with a default did not catch it and the width
  reached the DDL as `vector(None)` / `list_size=None`.
- Point HF_HUB_CACHE at the prefetch directory. Chunking loads the tokenizer
  through `tokenizers`, which reads the hub cache, so a fresh container
  fetched over the network on first ingest and an offline one silently fell
  back to cl100k.
- Charge a token that collapses a long unbroken run by its character span.
  WordPiece emits one [UNK] for any word over its character limit, which made
  base64 and minified content count as near-zero tokens, so nothing split it
  and oversized chunks reached the embedding server.
- Preserve chunk ids and honour --batch-size when rebuilding a FAISS index.
  Fresh uuids orphaned GraphRAG's graph_node_chunks rows, and the whole index
  went out in a single embed call on remote servers.
- Document that granite runs an int8-quantised graph, and scope the
  SentenceTransformer parity claim to mpnet's fp32 graph, which is where it
  was measured.
- Correct the embeddings docs: a matching dimension is not a matching model,
  so a same-width swap raises nothing and silently degrades retrieval.
2026-08-27 13:59:51 +01:00
..
2026-06-14 21:36:07 +01:00
2026-06-14 21:36:07 +01:00
2026-08-26 16:37:03 +01:00
2026-08-26 16:37:03 +01:00