Files
DocsGPT/application/celeryconfig.py
T
Alex 47e53a71c4 feat: embed on the worker, and stop batching the ONNX pass
The API embeds every query it serves, so it held its own copy of the model:
~890 MB it never needed. EMBEDDINGS_DELEGATE_TO_WORKER (on by default) sends
the text to the Celery worker instead and gets the vector back, taking an API
process from 1176 MB to 285 MB with no ONNX Runtime imported at all. The client
embeds locally when it finds itself inside a worker task, so the worker never
dispatches to itself -- the same self-deadlock DOCUMENT_PARSE_QUEUE avoids on
the parsing side. EMBEDDINGS_BASE_URL still wins over it, and remains the right
answer for production.

ensure_vector_schema was constructing the embeddings instance purely to read
.dimension off it, loading several hundred MB of ONNX into every API and worker
process at import. For a model the registry describes that is a lookup; only an
unregistered name now falls back to loading.

EMBEDDINGS_BATCH_SIZE was sizing two unrelated things: chunks per store
transaction (and per remote embed request) and documents per ONNX forward pass.
Each pass pads every input up to its longest, and that waste grows with the
square of chunk length, so at the 1250-token default a batch of 32 peaked at
6.6 GB and took 326s where a batch of 1 peaked at 2.9 GB and took 90s. The
forward pass is now sized by EMBEDDINGS_MODEL_BATCH_SIZE, defaulting to 1;
storage and remote batching are unchanged at 32.

reembed embeds in-process: a batch job that walks the whole index should not
round-trip every chunk through a broker, and loading the model there reports a
real failure instead of timing out against an empty queue.

Also drops the mpnet zip download from the docs and the devcontainer, which
pointed at a SentenceTransformers export with no ONNX graph and had been inert
since the FastEmbed swap; corrects the claim that any sentence-transformers
model works; and settles the Configuring/Settings pages on what the registry
and the repository metadata actually decide.
2026-08-28 12:23:18 +01:00

70 lines
2.9 KiB
Python

from kombu import Queue
from application.core.settings import settings
# Pydantic loads .env into ``settings`` but does not inject values into
# ``os.environ`` — read directly from settings so beat startup (which
# imports this module before any explicit env load) sees a real URL.
broker_url = settings.CELERY_BROKER_URL
result_backend = settings.CELERY_RESULT_BACKEND
task_serializer = 'json'
result_serializer = 'json'
accept_content = ['json']
# Autodiscover tasks
imports = (
'application.api.user.tasks',
'application.vectorstore.embeddings_tasks',
)
# Project-scoped queue so a stray sibling worker on the same broker
# (other repo, same default ``celery`` queue) can't grab DocsGPT tasks.
task_default_queue = "docsgpt"
task_default_exchange = "docsgpt"
task_default_routing_key = "docsgpt"
# Route document parsing to a dedicated queue so a parse enqueued from inside a
# Celery worker (headless/scheduled agent) is served by a separate parsing worker
# and never self-deadlocks the awaiting worker. The tool also passes the queue at
# apply_async time, so this routing is the default for any other enqueuer.
# Query embedding gets its own queue for the same reason parsing does: a query
# waiting behind a multi-minute ingest is a query that has timed out. A bare
# worker still consumes it, but its concurrency is shared -- run a separate
# ``-Q embeddings`` worker to actually isolate query latency from ingest.
task_routes = {
"application.api.user.tasks.parse_document": {"queue": settings.DOCUMENT_PARSE_QUEUE},
"application.vectorstore.embeddings_tasks.embed_texts": {"queue": settings.EMBEDDINGS_QUEUE},
}
# Declare every queue so a bare ``celery worker`` (no -Q) consumes ALL of them —
# the default worker does the whole job, parsing included. Operators who want
# heavy OCR isolated run one worker with ``-Q docsgpt`` and another with
# ``-Q parsing``. (dict.fromkeys dedupes if DOCUMENT_PARSE_QUEUE == "docsgpt".)
task_queues = tuple(
Queue(name)
for name in dict.fromkeys(
["docsgpt", settings.DOCUMENT_PARSE_QUEUE, settings.EMBEDDINGS_QUEUE]
)
)
beat_scheduler = "redbeat.RedBeatScheduler"
redbeat_redis_url = broker_url
redbeat_key_prefix = "redbeat:docsgpt:"
redbeat_lock_timeout = 90
# Survive worker SIGKILL/OOM without silently dropping in-flight tasks.
task_acks_late = True
task_reject_on_worker_lost = True
worker_prefetch_multiplier = settings.CELERY_WORKER_PREFETCH_MULTIPLIER
broker_transport_options = {"visibility_timeout": settings.CELERY_VISIBILITY_TIMEOUT}
result_expires = 86400 * 7
task_track_started = True
# Recycle the prefork worker child to bound native-heap growth from
# docling/torch parsing. Left unset (Celery's unlimited default) when 0.
if settings.CELERY_WORKER_MAX_MEMORY_PER_CHILD > 0:
worker_max_memory_per_child = settings.CELERY_WORKER_MAX_MEMORY_PER_CHILD
if settings.CELERY_WORKER_MAX_TASKS_PER_CHILD > 0:
worker_max_tasks_per_child = settings.CELERY_WORKER_MAX_TASKS_PER_CHILD