4 Commits
Author SHA1 Message Date
Alex 7da46c2bea feat: air-gapped deployment guide, no implicit downloads
- Ship tiktoken's cl100k_base inside the package and build the encoding
  from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
  temp dir, and read tokenizer.json and repo metadata from that cache, so
  a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
  the endpoints return 404, audio files fail to ingest with a clear
  message, /api/config reports tts_available/stt_available, and the UI
  hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
  packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
2026-09-15 17:54:24 +01:00
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00
Alex 4707c45b93 fix(parser,vectorstore): stop the first-request downloads in a warmed install
Three things still reached the network from a container whose models were
baked in:

- tiktoken fetched cl100k_base from openaipublic.blob.core.windows.net on
  every fresh container (its cache defaulted to /tmp), and token accounting
  calls it on every chat. prefetch_models now warms it too; the image sets
  TIKTOKEN_CACHE_DIR.
- The chunker loaded its tokenizer with Tokenizer.from_pretrained, which
  revalidates the revision with a HEAD request per process start and stalls
  for the etag timeout (10 s) when huggingface.co is unreachable. It now reads
  tokenizer.json from the hub cache first and only downloads on a miss; the
  repo-metadata read for models outside the registry does the same.
- tldextract fetched the public suffix list on the first web crawl; the
  bundled snapshot is used instead.

application/scripts/verify_offline.py exercises these paths (and docling's
conversion when the extra is installed) so an image can be checked with
docker run --network none.
2026-09-05 15:50:20 +01:00
Alex dd3876fdcb fix: mini fixes 2026-08-26 16:37:03 +01:00