`docsgpt up --native` installs services meant to outlive the shell. Development
wants the opposite, and until now it meant three terminals from the guide:
uvicorn, celery, and vite.
`docsgpt dev` runs this checkout's API and worker as children of one terminal,
both restarting when a file is saved, their output interleaved and labelled, and
Ctrl-C stopping them together. `--ui` adds the Vite dev server, `--mock-llm`
runs the bundled mock model so no API key is needed, and `--no-worker` leaves
the worker to your editor's debugger. Celery has no reloader of its own, so the
worker is wrapped in watchfiles when it is installed, and runs plain when it is
not.
Alongside it, the commands a dev loop keeps reaching for:
- `docsgpt doctor` checks what usually breaks a new setup: PostgreSQL answering
and its schema matching this version, Redis answering, a model provider being
configured, and the port being free.
- `docsgpt restart [api|worker]` bounces services without rewriting settings or
rerunning migrations, which `down` plus `up` did.
- `docsgpt logs -f` follows a native install instead of telling you to run
`tail -f` yourself.
- `docsgpt env set` applies itself to a running native install rather than
asking you to run `docsgpt up` again to change one value.
Two bugs found on the way, both older than this change:
- `docsgpt api --reload` watched the working directory, which in a checkout is
178,425 files: .venv, node_modules, and the indexes/ and inputs/ the app
writes to while ingesting, so the server restarted itself mid-request. It
watches the package now — 1,217 files.
- The VS Code "Flask Debugger" ran `flask run`, which serves only the WSGI app:
/mcp, the SSE streams and artifact downloads 404 under it. The guide warned
about this in prose while the debug config did it anyway. It runs uvicorn on
the ASGI app now, like production.
- Ship tiktoken's cl100k_base inside the package and build the encoding
from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
temp dir, and read tokenizer.json and repo metadata from that cache, so
a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
the endpoints return 404, audio files fail to ingest with a clear
message, /api/config reports tts_available/stt_available, and the UI
hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
Move GET /api/events, the remote-device command stream and artifact
downloads from Flask to Starlette routes mounted ahead of the Flask
catch-all. On Flask each held an a2wsgi threadpool slot for as long as
its response stayed open, and because uvicorn drops writes after a
client disconnects, a closed tab never released it.
- asgi_auth: one JWT/OIDC gate for Starlette routes; the chat reconnect
reader uses it too
- ClosingStreamingResponse closes the body iterator and releases the
SSE slot or file handle even when the client leaves before the first
frame
- AsyncTopic liveness probe replaces the sync client's socket_timeout
guard against half-open pub/sub sockets
- ASYNC_REDIS_MAX_CONNECTIONS sizes the async Redis pool; every open
stream holds a connection and redis-py defaults to 100
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.
Kept for one release:
- A top-level application package whose meta-path finder resolves
application.x.y to the already-imported docsgpt.x.y object, so old imports
and entry points (celery -A application.app.celery,
uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
docsgpt.* task on start-up, so messages queued by the previous release still
run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
the previous release wrote are left unread instead of firing twice.
The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
A named volume mounted on /app/inputs, /app/indexes or /app/vectors inherits
the ownership of the image directory, so the image creates them as appuser;
before this the volume came up root-owned and every upload failed with a
permission error unless the container ran as root. docker-compose-standalone.yaml
still runs backend and worker as root, like docker-compose-hub.yaml, so it
also works with image tags that predate these directories.
requirements-docling.txt adds the PyTorch CPU index, which uv resolves only
with UV_INDEX_STRATEGY=unsafe-best-match (pip is unaffected); the file header
and the docs say so and point uv users at uv sync --extra docling.
requirements.txt pinned torch and transformers in core although only docling
needs them, and on Linux torch pulls the CUDA 13 stack: 2.7 GB of the 3.0 GB
wheel download. Direct dependencies now live in pyproject.toml, uv.lock pins
everything, and application/requirements*.txt are exported from the lock by
scripts/export_requirements.sh (each file is the core set plus one extra).
The docling extra pins torch/torchvision/transformers itself and, on Linux,
resolves torch from the CPU-only PyTorch index (no nvidia packages). milvus
(pymilvus + milvus-lite, which pulls pyarrow) is the second extra.
application/core/optional_deps.py is the one place install hints come from;
the milvus store and the docling call sites use it so a missing extra fails
with the exact command to run.
Conflicts, and how each was taken:
- application/core/settings.py — ours. The renamed OCR_ENABLED /
OCR_ATTACHMENTS_ENABLED / OCR_MIN_CHARS_PER_PAGE accept main's
DOCLING_OCR_* spellings as AliasChoices, so nothing is dropped.
- application/Dockerfile — both. Main's install layers plus the
INSTALL_DOCLING build arg.
- application/parser/file/constants.py — both imports.
- deployment/docker-compose.yaml — both. The INSTALL_DOCLING /
INSTALL_TESSERACT build args on backend and worker, and main's
-Q docsgpt,parsing,embeddings, which query embedding needs.
- tests/conftest.py — theirs. Both sides fixed the same pytest-postgresql
9.0.0 autocommit= breakage; main's spelling is the one already on main.
- application/requirements.txt — the comments claimed different reasons
torch is in core. Main's is the true one now: it removed
sentence-transformers, so docling is torch's only remaining consumer.
Two things the merge broke without conflicting:
- onnxruntime. This branch moved it out of core into the docling extra;
main meanwhile made it the runtime local embeddings execute on
(fastembed). Git took the deletion, leaving fastembed with no pinned
runtime in a repo that pins everything. Restored to core, and no longer
pinned twice from the extra.
- The frontend copy of ATTACHMENT_PARSER_EXTENSIONS. The backend list is
derived and picked up the anydoc suffixes; the hand-kept frontend mirror
did not, so the composer would refuse files the API accepts.
tests/parser/file/test_constants.py is what caught it.
ruff, pytest (9897 passed), frontend build and docs build all pass. The
image build is unverified: no Docker daemon on this machine.
The API embeds every query it serves, so it held its own copy of the model:
~890 MB it never needed. EMBEDDINGS_DELEGATE_TO_WORKER (on by default) sends
the text to the Celery worker instead and gets the vector back, taking an API
process from 1176 MB to 285 MB with no ONNX Runtime imported at all. The client
embeds locally when it finds itself inside a worker task, so the worker never
dispatches to itself -- the same self-deadlock DOCUMENT_PARSE_QUEUE avoids on
the parsing side. EMBEDDINGS_BASE_URL still wins over it, and remains the right
answer for production.
ensure_vector_schema was constructing the embeddings instance purely to read
.dimension off it, loading several hundred MB of ONNX into every API and worker
process at import. For a model the registry describes that is a lookup; only an
unregistered name now falls back to loading.
EMBEDDINGS_BATCH_SIZE was sizing two unrelated things: chunks per store
transaction (and per remote embed request) and documents per ONNX forward pass.
Each pass pads every input up to its longest, and that waste grows with the
square of chunk length, so at the 1250-token default a batch of 32 peaked at
6.6 GB and took 326s where a batch of 1 peaked at 2.9 GB and took 90s. The
forward pass is now sized by EMBEDDINGS_MODEL_BATCH_SIZE, defaulting to 1;
storage and remote batching are unchanged at 32.
reembed embeds in-process: a batch job that walks the whole index should not
round-trip every chunk through a broker, and loading the model there reports a
real failure instead of timing out against an empty queue.
Also drops the mpnet zip download from the docs and the devcontainer, which
pointed at a SentenceTransformers export with no ONNX graph and had been inert
since the FastEmbed swap; corrects the claim that any sentence-transformers
model works; and settles the Configuring/Settings pages on what the registry
and the repository metadata actually decide.