mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-03 11:11:58 +00:00
Backend (arc53/docsgpt): 4.5 GB compressed -> 0.9 GB with both embedding models and tiktoken baked in. - torch/transformers gone from the default install (docling extra only). - Ubuntu 24.04 ships python3.12: no deadsnakes PPA, no software-properties- common; every pin is a wheel, so no gcc/g++/rust in the builder. - COPY --chown and a prefetch that runs as the process user replace the trailing chown -R, which duplicated the 600 MB model layer. - .dockerignore keeps __pycache__, .coverage, local indexes and .env out. - EXTRAS build arg (INSTALL_DOCLING kept as an alias); the docling variant also bakes docling's layout/table/RapidOCR models (DOCLING_ARTIFACTS_PATH) and tesseract, and drops only the discovery documents of Google APIs the app never builds. - FLASK_DEBUG env removed (unused); OCI labels added. Frontend (arc53/docsgpt-fe): 302 MB Vite dev server -> 25 MB static build behind nginx. VITE_* variables are injected at container start into /config.js and read through src/env.ts, so the image no longer needs a rebuild per deployment; docker-compose.yaml keeps hot reload via the dev target. Publishing: every release and develop build now pushes a slim tag and a -docling tag (docling engine + models + tesseract). docker-compose-hub.yaml takes DOCSGPT_IMAGE_TAG / DOCSGPT_IMAGE_VARIANT; docker-compose-standalone.yaml runs the stack from pre-built images without a checkout and is attached to each release. setup.sh selects the -docling variant for OCR instead of requiring a local build. A new workflow builds the image on PRs that touch it and runs verify_offline under --network none; lint checks the exported requirements match uv.lock.