16 Commits
Author SHA1 Message Date
Alex aecb596e99 build(docker): slim backend image, static frontend image, -docling variant
Backend (arc53/docsgpt): 4.5 GB compressed -> 0.9 GB with both embedding
models and tiktoken baked in.
- torch/transformers gone from the default install (docling extra only).
- Ubuntu 24.04 ships python3.12: no deadsnakes PPA, no software-properties-
  common; every pin is a wheel, so no gcc/g++/rust in the builder.
- COPY --chown and a prefetch that runs as the process user replace the
  trailing chown -R, which duplicated the 600 MB model layer.
- .dockerignore keeps __pycache__, .coverage, local indexes and .env out.
- EXTRAS build arg (INSTALL_DOCLING kept as an alias); the docling variant
  also bakes docling's layout/table/RapidOCR models (DOCLING_ARTIFACTS_PATH)
  and tesseract, and drops only the discovery documents of Google APIs the
  app never builds.
- FLASK_DEBUG env removed (unused); OCI labels added.

Frontend (arc53/docsgpt-fe): 302 MB Vite dev server -> 25 MB static build
behind nginx. VITE_* variables are injected at container start into
/config.js and read through src/env.ts, so the image no longer needs a
rebuild per deployment; docker-compose.yaml keeps hot reload via the dev
target.

Publishing: every release and develop build now pushes a slim tag and a
-docling tag (docling engine + models + tesseract). docker-compose-hub.yaml
takes DOCSGPT_IMAGE_TAG / DOCSGPT_IMAGE_VARIANT; docker-compose-standalone.yaml
runs the stack from pre-built images without a checkout and is attached to
each release. setup.sh selects the -docling variant for OCR instead of
requiring a local build. A new workflow builds the image on PRs that touch
it and runs verify_offline under --network none; lint checks the exported
requirements match uv.lock.
2026-09-05 15:50:21 +01:00
Alex 3947c66cda Merge branch 'main' into anydoc-support
Conflicts, and how each was taken:

- application/core/settings.py — ours. The renamed OCR_ENABLED /
  OCR_ATTACHMENTS_ENABLED / OCR_MIN_CHARS_PER_PAGE accept main's
  DOCLING_OCR_* spellings as AliasChoices, so nothing is dropped.
- application/Dockerfile — both. Main's install layers plus the
  INSTALL_DOCLING build arg.
- application/parser/file/constants.py — both imports.
- deployment/docker-compose.yaml — both. The INSTALL_DOCLING /
  INSTALL_TESSERACT build args on backend and worker, and main's
  -Q docsgpt,parsing,embeddings, which query embedding needs.
- tests/conftest.py — theirs. Both sides fixed the same pytest-postgresql
  9.0.0 autocommit= breakage; main's spelling is the one already on main.
- application/requirements.txt — the comments claimed different reasons
  torch is in core. Main's is the true one now: it removed
  sentence-transformers, so docling is torch's only remaining consumer.

Two things the merge broke without conflicting:

- onnxruntime. This branch moved it out of core into the docling extra;
  main meanwhile made it the runtime local embeddings execute on
  (fastembed). Git took the deletion, leaving fastembed with no pinned
  runtime in a repo that pins everything. Restored to core, and no longer
  pinned twice from the extra.
- The frontend copy of ATTACHMENT_PARSER_EXTENSIONS. The backend list is
  derived and picked up the anydoc suffixes; the hand-kept frontend mirror
  did not, so the composer would refuse files the API accepts.
  tests/parser/file/test_constants.py is what caught it.

ruff, pytest (9897 passed), frontend build and docs build all pass. The
image build is unverified: no Docker daemon on this machine.
2026-09-04 16:42:48 +01:00
Alex b473498baf fix(setup,docs): rebuild locally built images, name OCR_DEEPSEEK_URL
Both setup scripts start with `compose pull && compose up -d`. In the local
compose file backend and worker are build-only services, so `up -d` builds
only when no image exists yet: a rerun that switches OCR on wrote
INSTALL_TESSERACT=true to .env and then reused the image built without it,
leaving OCR_ENABLED=true with no engine. Build explicitly on the local
compose path; the hub path stays pull-only, its services have no build stage.

The OCR message named OCR_ENGINE=deepseek but not OCR_DEEPSEEK_URL, whose
default (localhost:11434) resolves to the container, not the host.

deployment/sandbox/README.md still described Docling as already present in
application/requirements.txt.
2026-09-04 15:28:29 +01:00
Alex 2a9a02427f fix(setup,docs): write INSTALL_TESSERACT from setup.ps1, correct the OCR upgrade note
setup.ps1 wrote OCR_ENABLED=true but never INSTALL_TESSERACT=true, so a
Windows user answering yes to the OCR question ended up with OCR on and no
engine in the image; it also claimed tesseract was "shipped in the Docker
image", which this change makes false, and offered OCR for the pre-built
Docker Hub images that cannot include it. Mirror setup.sh: skip the question
for hub images (naming all three settings the DeepSeek path needs), and bake
tesseract in for locally built ones.

The upgrade callout said earlier images always included tesseract. They never
did -- they included docling, and OCR ran on the RapidOCR engine bundled with
it, needing no system package. It also covered only OCR_ENABLED, missing
OCR_ATTACHMENTS_ENABLED, which is a separate switch onto the same native OCR
path.
2026-09-04 15:13:23 +01:00
Pavel ed0892b39b Batch fixes 2 2026-09-03 00:30:59 +04:00
Alex de22be5a21 fix: chunk-budget blowups, FastEmbed built-ins, and re-embed gaps
Follow-up review pass over the embeddings branch.

- Fold an oversized header back into the body, and drop header duplication
  when it would leave under a quarter of the chunk budget. A header at or
  over max_tokens collapsed the body budget to one token, so a document
  became one chunk per body token, each still over the cap: a 95 KB file
  produced 20k chunks of 2563 tokens against a 1250 cap. Also clamp
  max_tokens to at least 1, as the strategy chunkers already do.
- Emit a header-only document as its own chunk. With no body piece to
  attach it to, splitting returned nothing and the document was dropped
  from the index with no error and no log line.
- Skip add_custom_model for a repository FastEmbed already ships. It
  rejects a name it knows, so configuring any of its ~30 built-ins
  (MiniLM, bge, e5, gte, ...) failed every embed call and every query.
- Decide "the user chose this model" by comparing against the field
  default rather than model_fields_set, which is true for anything read
  from .env. Every setup script has always written EMBEDDINGS_NAME, so an
  upgraded remote-embeddings install inherited mpnet's 384-token window
  and silently clipped ~80% off every chunk.
- Cut tiktoken splits at character offsets instead of decoding each token
  window. A multi-byte character straddling a boundary decoded to U+FFFD
  on both sides, destroying one character at roughly one boundary in five
  on CJK text -- including at the default max_tokens of 2000.
- Let the re-embed script open a FAISS index whose width does not match
  the configured model. That mismatch is the main reason to run it, and
  the error recommending the script was raised by the script itself, so
  the advice failed on every source.
- Re-embed graph_nodes.name_embedding when GraphRAG is enabled. Those
  vectors seed every traversal and share the chunk vectors' width, so a
  same-width model swap left the graph retrieving from the old space with
  nothing to report it.
- Prefetch the models before copying the application source, so editing
  any file no longer re-downloads ~780 MB of artifacts on every build.
- Mirror the setup.sh embedding menu into setup.ps1: granite default,
  legacy mpnet as an explicit option, and both engine flows updated.
  Windows users were otherwise stranded on mpnet with no granite path.
- Drop the unused EmbeddingsWrapper.tokenizer property.
2026-08-27 15:55:21 +01:00
Kushal 14cfe8618b Fix RandomNumberGenerator.Fill incompatibility with Windows PowerShell 5.1 2026-08-10 18:29:51 +05:30
Alex 27f9ddc8e0 fix: address PR review — Responses input encoding + finish Azure removal
- Replay prior-turn assistant text as input_text, not output_text: a
  Responses easy-input message only accepts input_* content parts, so the
  old shape would 400 on the second turn of every conversation (default
  store=false resends prior turns inline).
- Always request include=["reasoning.encrypted_content"] so in-turn
  reasoning carryover works whether or not the response is also stored
  server-side (OPENAI_RESPONSES_STORE).
- Add detail="auto" to input_image parts.
- Remove the now-dead azure_openai option from setup.sh / setup.ps1 and
  the docs, completing the AzureOpenAILLM removal.
- Tests: correct the input_text / image-detail / include assertions and
  add parallel tool-call and stream-error coverage.
2026-06-03 14:56:28 +01:00
Alex d9a92a7208 feat: improve setup scripts 2026-04-03 17:15:21 +01:00
Alex-wuhu eaf39bb15b feat: add Novita AI as LLM provider
Add Novita AI (https://novita.ai) as a new LLM provider option.
Novita offers OpenAI-compatible API endpoints with competitive pricing.
2026-03-23 10:52:26 +08:00
Pavel 8aa44c415b Advanced settings (#2281)
Add additional settings to setup scripts
2026-02-17 11:54:59 +00:00
Pavel e7d2af2405 Setup plus env fixes (#2265)
* fixes setup scripts

fixes to env handling in setup script plus other minor fixes

* Remove var declarations

Declarations such as `LLM_PROVIDER=$LLM_PROVIDER` override .env variables in compose

Similar issue is present in the frontend - need to choose either to switch to separate frontend env or keep as is.

* Manage apikeys in settings

1. More pydantic management of api keys.
2. Clean up of variable declarations from docker compose files, used to block .env imports. Now should be managed ether by settings.py defaults or .env
2026-01-22 12:21:01 +02:00
Marco Ponce a8d2024791 Windows deployment powershell and renamed LLM_PROVIDER and runtime (#2050)
* Windows deployment powershell and  renamed LLM_PROVIDER and runtime

* added LLM_NAME back

* revert changes on docker-compose-hub.yaml
2025-10-12 15:25:42 +03:00
Siddhant Rai 9da4215d1f feat: implement Docker Hub integration for building and pushing images in CI/CD workflow 2025-08-28 12:01:04 +05:30
Ankit Matth b3af4ee50b speed up scripts by using docker hub 2025-08-24 08:59:19 +05:30
Pavel 1223fd2149 win-setup 2025-03-20 21:48:59 +03:00