Backend (arc53/docsgpt): 4.5 GB compressed -> 0.9 GB with both embedding
models and tiktoken baked in.
- torch/transformers gone from the default install (docling extra only).
- Ubuntu 24.04 ships python3.12: no deadsnakes PPA, no software-properties-
common; every pin is a wheel, so no gcc/g++/rust in the builder.
- COPY --chown and a prefetch that runs as the process user replace the
trailing chown -R, which duplicated the 600 MB model layer.
- .dockerignore keeps __pycache__, .coverage, local indexes and .env out.
- EXTRAS build arg (INSTALL_DOCLING kept as an alias); the docling variant
also bakes docling's layout/table/RapidOCR models (DOCLING_ARTIFACTS_PATH)
and tesseract, and drops only the discovery documents of Google APIs the
app never builds.
- FLASK_DEBUG env removed (unused); OCI labels added.
Frontend (arc53/docsgpt-fe): 302 MB Vite dev server -> 25 MB static build
behind nginx. VITE_* variables are injected at container start into
/config.js and read through src/env.ts, so the image no longer needs a
rebuild per deployment; docker-compose.yaml keeps hot reload via the dev
target.
Publishing: every release and develop build now pushes a slim tag and a
-docling tag (docling engine + models + tesseract). docker-compose-hub.yaml
takes DOCSGPT_IMAGE_TAG / DOCSGPT_IMAGE_VARIANT; docker-compose-standalone.yaml
runs the stack from pre-built images without a checkout and is attached to
each release. setup.sh selects the -docling variant for OCR instead of
requiring a local build. A new workflow builds the image on PRs that touch
it and runs verify_offline under --network none; lint checks the exported
requirements match uv.lock.
Conflicts, and how each was taken:
- application/core/settings.py — ours. The renamed OCR_ENABLED /
OCR_ATTACHMENTS_ENABLED / OCR_MIN_CHARS_PER_PAGE accept main's
DOCLING_OCR_* spellings as AliasChoices, so nothing is dropped.
- application/Dockerfile — both. Main's install layers plus the
INSTALL_DOCLING build arg.
- application/parser/file/constants.py — both imports.
- deployment/docker-compose.yaml — both. The INSTALL_DOCLING /
INSTALL_TESSERACT build args on backend and worker, and main's
-Q docsgpt,parsing,embeddings, which query embedding needs.
- tests/conftest.py — theirs. Both sides fixed the same pytest-postgresql
9.0.0 autocommit= breakage; main's spelling is the one already on main.
- application/requirements.txt — the comments claimed different reasons
torch is in core. Main's is the true one now: it removed
sentence-transformers, so docling is torch's only remaining consumer.
Two things the merge broke without conflicting:
- onnxruntime. This branch moved it out of core into the docling extra;
main meanwhile made it the runtime local embeddings execute on
(fastembed). Git took the deletion, leaving fastembed with no pinned
runtime in a repo that pins everything. Restored to core, and no longer
pinned twice from the extra.
- The frontend copy of ATTACHMENT_PARSER_EXTENSIONS. The backend list is
derived and picked up the anydoc suffixes; the hand-kept frontend mirror
did not, so the composer would refuse files the API accepts.
tests/parser/file/test_constants.py is what caught it.
ruff, pytest (9897 passed), frontend build and docs build all pass. The
image build is unverified: no Docker daemon on this machine.
Both setup scripts start with `compose pull && compose up -d`. In the local
compose file backend and worker are build-only services, so `up -d` builds
only when no image exists yet: a rerun that switches OCR on wrote
INSTALL_TESSERACT=true to .env and then reused the image built without it,
leaving OCR_ENABLED=true with no engine. Build explicitly on the local
compose path; the hub path stays pull-only, its services have no build stage.
The OCR message named OCR_ENGINE=deepseek but not OCR_DEEPSEEK_URL, whose
default (localhost:11434) resolves to the container, not the host.
deployment/sandbox/README.md still described Docling as already present in
application/requirements.txt.
setup.ps1 wrote OCR_ENABLED=true but never INSTALL_TESSERACT=true, so a
Windows user answering yes to the OCR question ended up with OCR on and no
engine in the image; it also claimed tesseract was "shipped in the Docker
image", which this change makes false, and offered OCR for the pre-built
Docker Hub images that cannot include it. Mirror setup.sh: skip the question
for hub images (naming all three settings the DeepSeek path needs), and bake
tesseract in for locally built ones.
The upgrade callout said earlier images always included tesseract. They never
did -- they included docling, and OCR ran on the RapidOCR engine bundled with
it, needing no system package. It also covered only OCR_ENABLED, missing
OCR_ATTACHMENTS_ENABLED, which is a separate switch onto the same native OCR
path.
Follow-up review pass over the embeddings branch.
- Fold an oversized header back into the body, and drop header duplication
when it would leave under a quarter of the chunk budget. A header at or
over max_tokens collapsed the body budget to one token, so a document
became one chunk per body token, each still over the cap: a 95 KB file
produced 20k chunks of 2563 tokens against a 1250 cap. Also clamp
max_tokens to at least 1, as the strategy chunkers already do.
- Emit a header-only document as its own chunk. With no body piece to
attach it to, splitting returned nothing and the document was dropped
from the index with no error and no log line.
- Skip add_custom_model for a repository FastEmbed already ships. It
rejects a name it knows, so configuring any of its ~30 built-ins
(MiniLM, bge, e5, gte, ...) failed every embed call and every query.
- Decide "the user chose this model" by comparing against the field
default rather than model_fields_set, which is true for anything read
from .env. Every setup script has always written EMBEDDINGS_NAME, so an
upgraded remote-embeddings install inherited mpnet's 384-token window
and silently clipped ~80% off every chunk.
- Cut tiktoken splits at character offsets instead of decoding each token
window. A multi-byte character straddling a boundary decoded to U+FFFD
on both sides, destroying one character at roughly one boundary in five
on CJK text -- including at the default max_tokens of 2000.
- Let the re-embed script open a FAISS index whose width does not match
the configured model. That mismatch is the main reason to run it, and
the error recommending the script was raised by the script itself, so
the advice failed on every source.
- Re-embed graph_nodes.name_embedding when GraphRAG is enabled. Those
vectors seed every traversal and share the chunk vectors' width, so a
same-width model swap left the graph retrieving from the old space with
nothing to report it.
- Prefetch the models before copying the application source, so editing
any file no longer re-downloads ~780 MB of artifacts on every build.
- Mirror the setup.sh embedding menu into setup.ps1: granite default,
legacy mpnet as an explicit option, and both engine flows updated.
Windows users were otherwise stranded on mpnet with no granite path.
- Drop the unused EmbeddingsWrapper.tokenizer property.
- Replay prior-turn assistant text as input_text, not output_text: a
Responses easy-input message only accepts input_* content parts, so the
old shape would 400 on the second turn of every conversation (default
store=false resends prior turns inline).
- Always request include=["reasoning.encrypted_content"] so in-turn
reasoning carryover works whether or not the response is also stored
server-side (OPENAI_RESPONSES_STORE).
- Add detail="auto" to input_image parts.
- Remove the now-dead azure_openai option from setup.sh / setup.ps1 and
the docs, completing the AzureOpenAILLM removal.
- Tests: correct the input_text / image-detail / include assertions and
add parallel tool-call and stream-error coverage.
* fixes setup scripts
fixes to env handling in setup script plus other minor fixes
* Remove var declarations
Declarations such as `LLM_PROVIDER=$LLM_PROVIDER` override .env variables in compose
Similar issue is present in the frontend - need to choose either to switch to separate frontend env or keep as is.
* Manage apikeys in settings
1. More pydantic management of api keys.
2. Clean up of variable declarations from docker compose files, used to block .env imports. Now should be managed ether by settings.py defaults or .env