Files
DocsGPT/application
Alex 37d93cbd86 Parse documents on a Celery parsing worker via a read_document tool
Replace the sandbox Docling extractor with read_document, backed by the in-process
backend parser (the same one ingestion uses) and offloaded to a dedicated
'parsing' Celery queue so it can run on GPU-capable workers with predictable RAM.
The tool resolves the input ref under the run-scoped gate, enqueues the parse,
and awaits it with a timeout (degrading to an error rather than hanging); the
worker independently re-resolves the artifact through the same gate and never
trusts a raw path. Untrusted files get the upload path's safeguards (extension
whitelist, size cap, sanitized temp file, cleanup). Options: output
(markdown/text/structured/chunks), ocr, pages, engine, max_chars, include_tables,
persist, json_schema. The workflow native-file 'extract' fallback now uses the
same worker path, so document parsing no longer needs the sandbox and works on
every backend.

Also fixes the branch's periodic-task test (the sandbox reaper made it 12) and
points the dev and e2e Celery workers at the parsing queue.
2026-06-25 13:24:12 +01:00
..
2026-05-15 12:23:31 +01:00
2026-04-18 13:13:57 +01:00
2026-06-18 16:21:42 +01:00
2026-03-17 14:27:48 +00:00
2026-04-25 14:57:37 +01:00
2026-04-12 00:07:24 +01:00
2026-06-23 17:11:49 +01:00
2026-06-08 22:02:21 +01:00
2026-05-15 12:23:31 +01:00
2026-06-08 15:07:31 +01:00
2026-06-08 15:07:31 +01:00
2026-04-28 02:27:02 +01:00
2026-05-31 18:00:51 +02:00
2026-06-16 18:08:04 +01:00
2026-05-23 02:39:20 +01:00