docsgpt/core/settings.py had grown to 258 fields in one 600-line class,
touched by about two commits a week, with related settings scattered
(GitHub ingest caps inside the embeddings block, API keys in four places,
the OpenAI Responses knobs 100 lines from the other OpenAI fields).
It is now a package: one module per domain (auth, llm, embeddings,
retrieval, vectorstores, database, workers, ingestion, ocr, storage,
connectors, server, events, agents, guardrails, scheduler, sandbox,
speech), each a SettingsGroup owning its fields and validators, composed
by multiple inheritance into the same flat Settings class. Every
attribute name, type, default, alias and constraint is unchanged, so
settings.NAME reads, .env files and test monkeypatches all keep working;
the import path docsgpt.core.settings is the package. Settings.normalize_api_key
is kept as a classmethod for callers that reuse it.
The comment above or beside each field became its Field(description=...),
so the definitions are visible to tooling; the next commit generates the
docs reference from them.
Pitfall recorded for future groups: pydantic collects validators by
method name across the MRO, so two groups naming a validator the same
would silently keep only one. Each group's validator has a unique name.
The connector OAuth popup posted the session token to window.opener with
a '*' target origin, so any page that opened the popup received it. An
attacker with an account on a multi-user deployment could start a flow for
their own pending session, get a victim to finish the provider consent, and
receive a token backed by the victim's Drive/SharePoint/Confluence tokens.
- Post popup results only to allowed frontend origins: the callback origin,
OIDC_FRONTEND_URL, the new CONNECTOR_ALLOWED_ORIGINS, and localhost:5173
when the callback runs on a loopback host.
- Render the success page from the callback itself so the token never
appears in a URL; callback-status ignores session_token/user_email params.
- ConnectorAuth accepts messages only from the popup it opened, on the
callback origin reported by /api/connectors/auth.
- /api/connectors/disconnect requires auth and only deletes the caller's
session.
- /api/connectors/sync and /api/remote reject session tokens the caller
does not own.
Fixes#2766
A deflate-compressed TIFF under 1 MB can declare 144 million pixels and
take 1.2 GB to convert to PNG, and Pillow only warns below 179 million.
Read the dimensions from the header and refuse images over 40 million
pixels before any pixel data is decoded. Pillow's DecompressionBombError
is now raised as DocumentParseError, so the upload fails once instead of
being retried.
Move GET /api/events, the remote-device command stream and artifact
downloads from Flask to Starlette routes mounted ahead of the Flask
catch-all. On Flask each held an a2wsgi threadpool slot for as long as
its response stayed open, and because uvicorn drops writes after a
client disconnects, a closed tab never released it.
- asgi_auth: one JWT/OIDC gate for Starlette routes; the chat reconnect
reader uses it too
- ClosingStreamingResponse closes the body iterator and releases the
SSE slot or file handle even when the client leaves before the first
frame
- AsyncTopic liveness probe replaces the sync client's socket_timeout
guard against half-open pub/sub sockets
- ASYNC_REDIS_MAX_CONNECTIONS sizes the async Redis pool; every open
stream holds a connection and redis-py defaults to 100
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.
Kept for one release:
- A top-level application package whose meta-path finder resolves
application.x.y to the already-imported docsgpt.x.y object, so old imports
and entry points (celery -A application.app.celery,
uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
docsgpt.* task on start-up, so messages queued by the previous release still
run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
the previous release wrote are left unread instead of firing twice.
The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
- The frontend image ran the Vite dev server in development mode, so
.env.development supplied its defaults (notification banner, Google client
id, local API host). The static build only loads .env.production, so the
build stage now copies .env.development in as the baseline and the compose
files pass every VITE_* the app reads through from .env; the runtime script
skips empty values so a blank passthrough keeps the build-time default.
.dockerignore kept only the .local variants out.
- VITE_DISABLE_SOURCE_FE disables sources only when it is the string true.
- DoclingParser: find_spec raises when docling itself is absent; the install
hint now covers that path, with a regression test.
- verify_offline: direct tests for verify(); the PR image check builds and
verifies the -docling variant as well as slim.
- Workflows this branch adds or rewrites pin actions by commit, pass the
release tag through env instead of template expansion, and do not persist
checkout credentials.
- OCR guide no longer claims pre-built images never include docling.
Backend (arc53/docsgpt): 4.5 GB compressed -> 0.9 GB with both embedding
models and tiktoken baked in.
- torch/transformers gone from the default install (docling extra only).
- Ubuntu 24.04 ships python3.12: no deadsnakes PPA, no software-properties-
common; every pin is a wheel, so no gcc/g++/rust in the builder.
- COPY --chown and a prefetch that runs as the process user replace the
trailing chown -R, which duplicated the 600 MB model layer.
- .dockerignore keeps __pycache__, .coverage, local indexes and .env out.
- EXTRAS build arg (INSTALL_DOCLING kept as an alias); the docling variant
also bakes docling's layout/table/RapidOCR models (DOCLING_ARTIFACTS_PATH)
and tesseract, and drops only the discovery documents of Google APIs the
app never builds.
- FLASK_DEBUG env removed (unused); OCI labels added.
Frontend (arc53/docsgpt-fe): 302 MB Vite dev server -> 25 MB static build
behind nginx. VITE_* variables are injected at container start into
/config.js and read through src/env.ts, so the image no longer needs a
rebuild per deployment; docker-compose.yaml keeps hot reload via the dev
target.
Publishing: every release and develop build now pushes a slim tag and a
-docling tag (docling engine + models + tesseract). docker-compose-hub.yaml
takes DOCSGPT_IMAGE_TAG / DOCSGPT_IMAGE_VARIANT; docker-compose-standalone.yaml
runs the stack from pre-built images without a checkout and is attached to
each release. setup.sh selects the -docling variant for OCR instead of
requiring a local build. A new workflow builds the image on PRs that touch
it and runs verify_offline under --network none; lint checks the exported
requirements match uv.lock.
setup.ps1 wrote OCR_ENABLED=true but never INSTALL_TESSERACT=true, so a
Windows user answering yes to the OCR question ended up with OCR on and no
engine in the image; it also claimed tesseract was "shipped in the Docker
image", which this change makes false, and offered OCR for the pre-built
Docker Hub images that cannot include it. Mirror setup.sh: skip the question
for hub images (naming all three settings the DeepSeek path needs), and bake
tesseract in for locally built ones.
The upgrade callout said earlier images always included tesseract. They never
did -- they included docling, and OCR ran on the RapidOCR engine bundled with
it, needing no system package. It also covered only OCR_ENABLED, missing
OCR_ATTACHMENTS_ENABLED, which is a separate switch onto the same native OCR
path.