mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-04 12:13:05 +00:00
Replace the sandbox Docling extractor with read_document, backed by the in-process backend parser (the same one ingestion uses) and offloaded to a dedicated 'parsing' Celery queue so it can run on GPU-capable workers with predictable RAM. The tool resolves the input ref under the run-scoped gate, enqueues the parse, and awaits it with a timeout (degrading to an error rather than hanging); the worker independently re-resolves the artifact through the same gate and never trusts a raw path. Untrusted files get the upload path's safeguards (extension whitelist, size cap, sanitized temp file, cleanup). Options: output (markdown/text/structured/chunks), ocr, pages, engine, max_chars, include_tables, persist, json_schema. The workflow native-file 'extract' fallback now uses the same worker path, so document parsing no longer needs the sandbox and works on every backend. Also fixes the branch's periodic-task test (the sandbox reaper made it 12) and points the dev and e2e Celery workers at the parsing queue.
61 lines
1.9 KiB
Markdown
61 lines
1.9 KiB
Markdown
# Welcome to DocsGPT Devcontainer
|
|
|
|
Welcome to the DocsGPT development environment! This guide will help you get started quickly.
|
|
|
|
## Starting Services
|
|
|
|
To run DocsGPT, you need to start three main services: Flask (backend), Celery (task queue), and Vite (frontend). Here are the commands to start each service within the devcontainer:
|
|
|
|
### Vite (Frontend)
|
|
|
|
```bash
|
|
cd frontend
|
|
npm run dev -- --host
|
|
```
|
|
|
|
### Backend (ASGI)
|
|
|
|
Run the full app under uvicorn (serves `/mcp` and the async SSE reconnect
|
|
routes, and matches production):
|
|
|
|
```bash
|
|
uvicorn application.asgi:asgi_app --host 0.0.0.0 --port 7091 --reload
|
|
```
|
|
|
|
`flask --app application/app.py run --host=0.0.0.0 --port=7091` is faster but
|
|
serves only the WSGI Flask app — it omits `/mcp` and the reconnect reader
|
|
`GET /api/messages/<id>/events`, so a dropped stream won't auto-resume.
|
|
|
|
### Celery (Task Queue)
|
|
|
|
```bash
|
|
celery -A application.app.celery worker -l INFO -Q docsgpt,parsing
|
|
```
|
|
|
|
The `parsing` queue serves document parsing (the `read_document` tool / workflow
|
|
native-file parse); without it those calls hang `DOCUMENT_PARSE_TIMEOUT` then
|
|
error. A dedicated `-Q parsing` worker can be GPU-enabled for heavier parsers.
|
|
|
|
## Github Codespaces Instructions
|
|
|
|
### 1. Make Ports Public:
|
|
|
|
Go to the "Ports" panel in Codespaces (usually located at the bottom of the VS Code window).
|
|
|
|
For both port 5173 and 7091, right-click on the port and select "Make Public".
|
|
|
|

|
|
|
|
|
|
### 2. Update VITE_API_HOST:
|
|
|
|
After making port 7091 public, copy the public URL provided by Codespaces for port 7091.
|
|
|
|
Open the file frontend/.env.development.
|
|
|
|
Find the line VITE_API_HOST=http://localhost:7091.
|
|
|
|
Replace http://localhost:7091 with the public URL you copied from Codespaces.
|
|
|
|

|