Commit Graph
8 Commits
Author SHA1 Message Date
Alex 0a15ce8fbb feat: attachment provenance 2026-08-13 14:30:28 +01:00
Alex ca53a3ac4b chore: bump docling 2026-08-12 09:05:00 +01:00
Alex 15b6b03b47 fix: minor size cleanup 2026-08-04 15:07:54 +01:00
Alex e4c0c26927 fix: more sandbox protections to terminate plus max pages 2026-07-09 19:56:13 +01:00
Alex 66786c2760 fix: remove code exec as default and sec improvements 2026-07-08 18:50:10 +01:00
Alex 94a845aa82 fix: more artefact hardening 2026-07-04 11:42:27 +02:00
Alex 75e07af78f fix: minor issue fixes 2026-06-30 18:36:06 +02:00
Alex 37d93cbd86 Parse documents on a Celery parsing worker via a read_document tool
Replace the sandbox Docling extractor with read_document, backed by the in-process
backend parser (the same one ingestion uses) and offloaded to a dedicated
'parsing' Celery queue so it can run on GPU-capable workers with predictable RAM.
The tool resolves the input ref under the run-scoped gate, enqueues the parse,
and awaits it with a timeout (degrading to an error rather than hanging); the
worker independently re-resolves the artifact through the same gate and never
trusts a raw path. Untrusted files get the upload path's safeguards (extension
whitelist, size cap, sanitized temp file, cleanup). Options: output
(markdown/text/structured/chunks), ocr, pages, engine, max_chars, include_tables,
persist, json_schema. The workflow native-file 'extract' fallback now uses the
same worker path, so document parsing no longer needs the sandbox and works on
every backend.

Also fixes the branch's periodic-task test (the sandbox reaper made it 12) and
points the dev and e2e Celery workers at the parsing queue.
2026-06-25 13:24:12 +01:00