Files
DocsGPT/deployment/docker-compose.yaml
T
Alex 00be2c05ad fix: make worker-delegated embedding survive the shipped deployments
Query embedding moved to the Celery worker, but nothing that ships was
updated to consume the queue it dispatches to.

- Add `embeddings` to every worker `-Q` list (compose x3, k8s, devcontainer,
  sandbox README). Without it a search blocked for EMBEDDINGS_DELEGATE_TIMEOUT
  and then answered with no retrieved context, because classic_rag swallows the
  dispatch error and skips the source -- bad answers, not an error.

- Skip the task_postrun heap reclaim for the embed task. The full gc.collect()
  was written for docling/torch parses; on a worker holding the ONNX model it
  measured ~86ms against ~8ms for the embed itself, a 9x slowdown of the round
  trip for a task that allocates a few kilobytes.

- Resolve the installation pin in the re-embed script. It never imports
  application.app, so an install pinned in app_metadata with no EMBEDDINGS_NAME
  set -- every stock k8s deployment, whose manifests carry no embedding config
  -- would rewrite its whole index with the legacy default and stamp
  sources.model to match, then be told by the boot warning to run it again.

- Fail fast for 30s after a failed dispatch. fanout.embed_questions falls back
  to letting each store embed its own query, so one dead-worker retrieval paid
  the timeout once in the fan-out and again per source.

- Forget the task result. Nothing reads it back: the key is per-dispatch UUID,
  not content-addressed, so a repeated query mints another. Left alone every
  search leaked ~17KB for result_expires (7 days) into the Redis the broker
  shares -- on the bundled k8s manifest (1Gi, no maxmemory policy) that is an
  OOMKill that takes the broker with it.

- Release the model ensure_vector_schema loads to read the width of an
  unregistered model, in a process that delegates and would never call it.
  The width still comes from the model, not the table, so the mismatch check
  the hook exists for keeps working.

- Correct the docs that said otherwise: embeddings.md claimed the standard
  deployment worked unchanged, upgrading.mdx said no action was needed, and
  the settings table listed none of the three delegation settings.
2026-08-28 14:31:19 +01:00

91 lines
2.6 KiB
YAML

name: docsgpt-oss
services:
frontend:
build: ../frontend
volumes:
- ../frontend/src:/app/src
environment:
- VITE_API_HOST=http://localhost:7091
- VITE_API_STREAMING=$VITE_API_STREAMING
- VITE_GOOGLE_CLIENT_ID=$VITE_GOOGLE_CLIENT_ID
ports:
- "5173:5173"
depends_on:
- backend
backend:
user: root
build: ../application
env_file:
- ../.env
environment:
# Override URLs to use docker service names
- CELERY_BROKER_URL=redis://redis:6379/0
- CELERY_RESULT_BACKEND=redis://redis:6379/1
- CACHE_REDIS_URL=redis://redis:6379/2
- POSTGRES_URI=postgresql://docsgpt:docsgpt@postgres:5432/docsgpt
ports:
- "7091:7091"
volumes:
- ../application/indexes:/app/indexes
- ../application/inputs:/app/inputs
- ../application/vectors:/app/vectors
depends_on:
redis:
condition: service_started
postgres:
condition: service_healthy
worker:
user: root
build: ../application
# Consumes the default queue AND the dedicated `parsing` (read_document /
# parse_document) and `embeddings` (query embedding) queues. Without `parsing`
# the read_document await never resolves; without `embeddings` every search
# fails after EMBEDDINGS_DELEGATE_TIMEOUT, because EMBEDDINGS_DELEGATE_TO_WORKER
# is on by default. For heavy/OCR parsing run a separate worker with `-Q parsing`;
# to keep query latency off the ingest pool, another with `-Q embeddings`.
command: celery -A application.app.celery worker -l INFO -B -Q docsgpt,parsing,embeddings
env_file:
- ../.env
environment:
# Override URLs to use docker service names
- CELERY_BROKER_URL=redis://redis:6379/0
- CELERY_RESULT_BACKEND=redis://redis:6379/1
- API_URL=http://backend:7091
- CACHE_REDIS_URL=redis://redis:6379/2
- POSTGRES_URI=postgresql://docsgpt:docsgpt@postgres:5432/docsgpt
volumes:
- ../application/indexes:/app/indexes
- ../application/inputs:/app/inputs
- ../application/vectors:/app/vectors
depends_on:
redis:
condition: service_started
postgres:
condition: service_healthy
redis:
image: redis:6-alpine
ports:
- 6379:6379
postgres:
image: postgres:16-alpine
environment:
- POSTGRES_USER=docsgpt
- POSTGRES_PASSWORD=docsgpt
- POSTGRES_DB=docsgpt
ports:
- "5432:5432"
volumes:
- postgres_data:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U docsgpt -d docsgpt"]
interval: 5s
timeout: 5s
retries: 10
volumes:
postgres_data: