mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-05 14:14:39 +00:00
Three things still reached the network from a container whose models were baked in: - tiktoken fetched cl100k_base from openaipublic.blob.core.windows.net on every fresh container (its cache defaulted to /tmp), and token accounting calls it on every chat. prefetch_models now warms it too; the image sets TIKTOKEN_CACHE_DIR. - The chunker loaded its tokenizer with Tokenizer.from_pretrained, which revalidates the revision with a HEAD request per process start and stalls for the etag timeout (10 s) when huggingface.co is unreachable. It now reads tokenizer.json from the hub cache first and only downloads on a miss; the repo-metadata read for models outside the registry does the same. - tldextract fetched the public suffix list on the first web crawl; the bundled snapshot is used instead. application/scripts/verify_offline.py exercises these paths (and docling's conversion when the extra is installed) so an image can be checked with docker run --network none.