Files
DocsGPT/tests
Alex 4707c45b93 fix(parser,vectorstore): stop the first-request downloads in a warmed install
Three things still reached the network from a container whose models were
baked in:

- tiktoken fetched cl100k_base from openaipublic.blob.core.windows.net on
  every fresh container (its cache defaulted to /tmp), and token accounting
  calls it on every chat. prefetch_models now warms it too; the image sets
  TIKTOKEN_CACHE_DIR.
- The chunker loaded its tokenizer with Tokenizer.from_pretrained, which
  revalidates the revision with a HEAD request per process start and stalls
  for the etag timeout (10 s) when huggingface.co is unreachable. It now reads
  tokenizer.json from the hub cache first and only downloads on a miss; the
  repo-metadata read for models outside the registry does the same.
- tldextract fetched the public suffix list on the first web crawl; the
  bundled snapshot is used instead.

application/scripts/verify_offline.py exercises these paths (and docling's
conversion when the extra is installed) so an image can be checked with
docker run --network none.
2026-09-05 15:50:20 +01:00
..
2026-08-22 14:41:17 +01:00
2026-08-20 14:54:35 +01:00
2026-07-20 19:57:44 +01:00
2026-04-18 13:13:57 +01:00
2026-03-30 16:13:08 +01:00
2026-08-29 12:21:46 +01:00
2026-03-12 14:46:26 +00:00
2026-09-03 01:00:07 +04:00
2026-04-18 13:13:57 +01:00
2026-08-27 17:36:38 +04:00
2026-08-12 16:55:33 +01:00
2026-07-08 00:23:32 +01:00
2026-06-14 21:36:07 +01:00
2026-08-20 13:56:37 +01:00
2026-03-30 16:13:08 +01:00
2026-07-07 18:51:31 +01:00
2026-01-22 13:11:24 +02:00
2026-08-10 14:51:33 +01:00
2026-04-18 13:13:57 +01:00
2026-06-16 09:21:21 +01:00
2026-07-07 18:51:31 +01:00
2026-04-18 13:13:57 +01:00
2026-04-18 13:13:57 +01:00
2026-08-13 14:30:28 +01:00
2026-04-21 14:22:32 +01:00