11 KiB
AGENTS.md
- Read
CONTRIBUTING.mdbefore making non-trivial changes. - For day-to-day development and feature work, follow the development-environment workflow rather than defaulting to
setup.sh/setup.ps1. - Avoid using the setup scripts during normal feature work unless the user explicitly asks for them. Users configure
.envusually. - Try to follow red/green TDD
Check existing dev prerequisites first
For feature work, do not assume the environment needs to be recreated.
- Check whether the user already has a Python virtual environment such as
venv/or.venv/. - Check whether Postgres is already running and reachable via
POSTGRES_URI(the canonical user-data store). - Check whether Redis is already running.
- Reuse what is already working. Do not stop or recreate Postgres, Redis, or the Python environment unless the task is environment setup or troubleshooting.
MongoDB is not required for the default install. It is only needed if the user opts into the Mongo vector-store backend (
VECTOR_STORE=mongodb) or is running the one-shotscripts/db/backfill.pyto migrate existing user data from the legacy Mongo-based install. In those cases,pymongois available as an optional extra, not a core dependency.
Normal local development commands
Use these commands once the dev prerequisites above are satisfied.
Backend
source .venv/bin/activate # macOS/Linux
uv pip install -r docsgpt/requirements.txt # or: pip install -r docsgpt/requirements.txt
# Optional extras (not installed by default; each file = core + the extra):
# uv pip install -r docsgpt/requirements-docling.txt # docling parser engine (OCR backend, structured output)
# uv pip install -r docsgpt/requirements-milvus.txt # VECTOR_STORE=milvus
# With uv alone: `uv sync --extra docling` (pyproject.toml + uv.lock are the source of truth).
# `uv pip install -r docsgpt/requirements-docling.txt` needs UV_INDEX_STRATEGY=unsafe-best-match
# (the file adds the PyTorch CPU index; prefer `uv sync --extra docling`).
The backend is also an installable package (pyproject.toml, hatchling).
uv sync installs it editable and puts a docsgpt command on PATH:
docsgpt api --reload, docsgpt worker, docsgpt migrate,
docsgpt prefetch-models, docsgpt verify-offline. Runtime data (.env,
inputs/, indexes/) lives in the checkout by default; DOCSGPT_HOME moves
that data home, and DOCSGPT_ENV_FILE selects only the .env file (see
docsgpt/core/paths.py). bash scripts/build_frontend.sh builds the web UI
into docsgpt/static (gitignored); the API serves it when present
(docsgpt/ui.py, switch SERVE_UI), and the package workflows run the script
before uv build so the wheel ships it.
Dependencies are declared in pyproject.toml and locked in uv.lock; the
docsgpt/requirements*.txt files are exported from the lock. To add or
bump a package: edit pyproject.toml, run uv lock, then
bash scripts/export_requirements.sh (CI fails if the exports are stale).
Never edit the requirements files by hand.
Run the API. For local dev, prefer the ASGI entrypoint under uvicorn — it serves the whole app, matches production, and hot-reloads:
uvicorn docsgpt.asgi:asgi_app --host 0.0.0.0 --port 7091 --reload
flask --app docsgpt/app.py run --host=0.0.0.0 --port=7091 is a faster
inner loop (quick startup, the Werkzeug interactive debugger), but it serves
only the WSGI Flask app and omits the routes mounted on the ASGI shell
in docsgpt/asgi.py:
- the
/mcpFastMCP endpoint, - the chat reconnect reader
GET /api/messages/<id>/events, - the notification stream
GET /api/events, - the remote-device command stream
GET /api/devices/sessions/<id>/events, and - artifact downloads
GET /api/artifacts/<id>/download.
These are native-async Starlette routes because they hold a response open
for a long time; on Flask each would pin a WSGI threadpool slot. Under
flask run those paths 404: chat still works (POST /stream is a Flask
route), but live notifications, stream auto-resume, paired devices and
artifact downloads don't. Use flask run only when you don't need them.
Production uses gunicorn -k uvicorn_worker.UvicornWorker against the same
docsgpt.asgi:asgi_app target; see docsgpt/Dockerfile for the
full flag set.
Run the Celery worker in a separate terminal:
celery -A docsgpt.app.celery worker -l INFO
The worker is required for retrieval, not optional. EMBEDDINGS_DELEGATE_TO_WORKER
defaults on, so the API embeds each query by dispatching to the worker rather than
loading a model of its own — which keeps the API process around 285 MB instead of
1.2 GB. Without a worker consuming EMBEDDINGS_QUEUE, every search fails after
EMBEDDINGS_DELEGATE_TIMEOUT. To run the API on its own, either set
EMBEDDINGS_DELEGATE_TO_WORKER=false (loads the model in-process) or point
EMBEDDINGS_BASE_URL at an embeddings service.
On macOS, prefer the solo pool for Celery:
python -m celery -A docsgpt.app.celery worker -l INFO --pool=solo
Note that --pool=solo costs roughly 350 ms per query embed against ~55 ms on the
default prefork pool — nearly all of it the solo worker picking the message up, not
the embedding itself. That only affects local dev; production runs prefork.
A bare worker (no -Q) consumes every configured queue, so one worker does the
whole job — app tasks, query embedding, and document parsing (the read_document
tool / workflow native-file parse) alike. Use -Q only to split load: run the main
worker with -Q docsgpt, a dedicated (e.g. GPU-enabled) parser worker with
-Q parsing for heavy OCR, and -Q embeddings to keep query latency off the ingest
pool. Note the main ingest task parses in-process on docsgpt; only
read_document is routed to parsing.
Frontend
Install dependencies only when needed, then run the dev server:
cd frontend
npm install --include=dev
npm run dev
Docs site
cd docs
npm install
Python / backend changes validation
ruff check .
python -m pytest
On macOS, run the suite with KMP_DUPLICATE_LIB_OK=TRUE:
KMP_DUPLICATE_LIB_OK=TRUE python -m pytest
faiss-cpu and torch each ship their own LLVM OpenMP runtime, and loading
both into one process makes libomp.dylib abort the interpreter
(OMP: Error #15). It is a macOS-only packaging clash, not a code fault: Linux
resolves both to libgomp, which tolerates duplicates, so CI (ubuntu-latest)
and the Docker images are unaffected. Without the variable, whether the run
aborts depends on which tests happen to load faiss and torch in the same
process, so a green run on one selection and an abort on another is expected.
Frontend changes
cd frontend && npm run lint
cd frontend && npm run build
Documentation changes
cd docs && npm run build
If Vale is installed locally and you edited prose, also run:
vale .
Repository map
docsgpt/: Flask backend, API routes, agent logic, retrieval, parsing, security, storage, Celery worker, and WSGI entrypoints.tests/: backend unit/integration tests and test-only Python dependencies.frontend/: Vite + React + TypeScript application.frontend/src/: main UI code, includingcomponents,conversation,hooks,locale,settings,upload, and Redux store wiring instore.ts.docs/: separate documentation site built with Next.js/Nextra.extensions/: integrations and widgets — currently the Chatwoot webhook bridge and the React widget (published to npm asdocsgpt). The Discord bot, Slack bot, and Chrome extension have been moved to their own repos underarc53/.deployment/: Docker Compose variants and Kubernetes manifests.
Coding rules
Backend
- Follow PEP 8 and keep Python line length at or under 120 characters.
- Use type hints for function arguments and return values.
- Add Google-style docstrings to new or substantially changed functions and classes.
- Add or update tests under
tests/for backend behavior changes. - Keep changes narrow in
api,auth,security,parser,retriever, andstorageareas.
Backend Abstractions
- LLM providers implement a common interface in
docsgpt/llm/(add new providers by extending the base class). - Vector stores are abstracted in
docsgpt/vectorstore/. - Parsers live in
docsgpt/parser/and handle different document formats in the ingestion stage. - Agents and tools are in
docsgpt/agents/anddocsgpt/agents/tools/. - Celery setup/config lives in
docsgpt/celery_init.pyanddocsgpt/celeryconfig.py. - Settings and env vars are managed via Pydantic in
docsgpt/core/settings/(one module per domain, composed intoSettings). Every field needs adescription; regenerate the docs reference withpython -m docsgpt.core.settings.reference --write.
Frontend
- Follow the existing ESLint + Prettier setup.
- Prefer small, reusable functional components and hooks.
- If shared state must be added, use Redux rather than introducing a new global state library.
- Avoid broad UI refactors unless the task explicitly asks for them.
- Do not re-create components if we already have some in the app.
- Follow
frontend/DESIGN.md: composecomponents/ui/parts and pick their look with props, use theme tokens, and keep typography, spacing, radius and motion on its roles. - Every user-visible string, attributes included (
aria-label,label,placeholder,title,alt), is at()key in all seven locales underfrontend/src/locale/(de en es jp ru zh zh-TW). Admin pages stay English.
Icons
DocsGPT historically mixed three icon sources: lucide-react, inline SVG components, and
.svg assets loaded via <img src=…>. For new code:
- Prefer
lucide-reactfor standard UI affordances (close, chevron, search, trash, plus, etc.). It tokenizes viacurrentColor, ships tree-shaken icons, and the codebase already imports it in 30+ places.<X className="size-4" />,<ChevronDown />, etc. - Use
assets/<name>.svg?reactwhen you need a brand-specific or domain illustration that doesn't exist in lucide (the app logo, robot fallback, send arrow, etc.). Always setfill="currentColor"/stroke="currentColor"in the SVG file so consumers can theme via Tailwind text classes. - Avoid
<img src={Asset}>for new icons. It blockscurrentColortheming and forces dark-variant duplicates (the audit removed several orphan dark/purple/white variants in this branch). The pattern is acceptable for existing call sites — don't bulk-migrate without a reason.
Three pre-existing dark-variant pairs (documentation, no-files, science-spark) are
hand-tuned multi-color illustrations, not pure inverts; they keep their -dark companion
files until a per-illustration refactor.
PR readiness
Before opening a PR:
- run the relevant validation commands above
- confirm backend changes still work end-to-end after ingesting sample data when applicable
- clearly summarize user-visible behavior changes
- mention any config, dependency, or deployment implications
- Ask your user to attach a screenshot or a video to it