Files
Alex f9f52e99f0 fix: make a native install take effect on systemd and stay out of Docker's way
From review of #2800:

- systemd `enable --now` starts nothing when the unit is already active, so a
  second `up --native` kept the old ExecStart and left the API on its previous
  port. start enables and then restarts, as the launchd path already did by
  booting the job out first.
- An explicit `home` now wins over XDG_CONFIG_HOME, which is what callers pass
  it for.
- The ExecStart program must be a real executable: when `docsgpt` is not on
  PATH, sys.argv[0] is accepted only if it can be run, and otherwise the
  failure is raised before any unit is written.
- `up --native` over a directory holding a Docker install now refuses and says
  how to proceed, instead of starting native services beside containers that
  down, status and uninstall would no longer see.
- Docs: without a terminal only --postgres-uri is required, and the Windows
  fallback names `docsgpt beat`, which the worker cannot embed there.

SystemdServices was the least covered part of the module and cannot be run on
this machine, so it now has tests for install, start, stop, remove, is_running
and a failing systemctl.
2026-09-16 22:35:14 +01:00

139 lines
7.4 KiB
Plaintext

---
title: Install with pip
description: Run the DocsGPT backend from the PyPI package, inside your own Python environment.
---
import { Callout } from 'nextra/components'
# Install with pip
DocsGPT is on PyPI as [`docsgpt`](https://pypi.org/project/docsgpt/): the API server, the web UI, the Celery worker and the maintenance scripts in one package, behind a single `docsgpt` command. Use it when you want DocsGPT inside your own Python environment or process manager rather than the [Docker images](/Deploying/Docker-Deploying).
<Callout type="info">
`docsgpt api` serves the web UI on the same port as the API. Set `SERVE_UI=false` to run the API alone, for example behind the frontend Docker image or a UI you host yourself.
</Callout>
<Callout type="info">
With Docker available, the same package can run the whole stack for you, including Postgres and Redis: `docsgpt up`. See [Run it with `docsgpt up`](/Deploying/Docker-Deploying#run-it-with-docsgpt-up).
</Callout>
## Requirements
- Python 3.12 or newer
- [PostgreSQL](/Deploying/Postgres-Migration) for user data, with the `vector` extension if you set `VECTOR_STORE=pgvector`
- Redis for the task queue and the cache
- An LLM: an API key for a hosted provider, or a local model server
## Install
```bash
python -m venv .venv && source .venv/bin/activate
pip install docsgpt
```
Extras add the optional engines:
```bash
pip install "docsgpt[docling]" # DOC_PARSER_ENGINE=docling: OCR backend, structured output
pip install "docsgpt[milvus]" # VECTOR_STORE=milvus
```
The `docling` extra pulls in PyTorch, and on Linux the PyPI torch wheels bring the CUDA stack with them. On a CPU-only machine install the CPU build first, then the extra; pip keeps the torch it already has:
```bash
pip install --index-url https://download.pytorch.org/whl/cpu torch torchvision
pip install "docsgpt[docling]"
```
`uv pip install` accepts the same two commands. With `pipx`, install `docsgpt[docling]` first, then replace the CUDA build inside its environment (`--no-deps` keeps pip from touching torch's dependencies, which the PyTorch index carries in older copies):
```bash
pipx runpip docsgpt install --force-reinstall --no-deps --index-url https://download.pytorch.org/whl/cpu torch torchvision
```
## Configure
DocsGPT keeps its runtime files in a **data home**: the `.env` file it reads settings from, uploaded files under `inputs/`, vector indexes under `indexes/` and downloaded embedding models under `models/`. The data home is `~/.docsgpt/server` (`/opt/docsgpt` when you run as root on Linux), whatever directory you run the commands from. `DOCSGPT_HOME` moves it, and `DOCSGPT_ENV_FILE` points at a `.env` kept somewhere else. Both variables must be set in the process environment, not in `.env`: they decide where `.env` is read from. In a source checkout the data home is the checkout.
<Callout type="warning">
Up to 0.20 the data home of an installed package was the directory you ran the command from. If you kept `.env` and your data there, move them to `~/.docsgpt/server` or set `DOCSGPT_HOME` to that directory. `docsgpt api` and `docsgpt worker` point out a `.env` in the working directory that they no longer read.
</Callout>
Create a `.env` in the data home. The minimum for a hosted LLM:
```ini
# LLM provider and model: see App Configuration for the options
LLM_PROVIDER=openai
LLM_NAME=<model name>
API_KEY=<provider API key>
# User data
POSTGRES_URI=postgresql://docsgpt:<password>@localhost:5432/docsgpt
# Worker-to-API authentication (required for uploads)
INTERNAL_KEY=<a long random string>
```
Redis defaults to `localhost:6379` (`CELERY_BROKER_URL`, `CELERY_RESULT_BACKEND`, `CACHE_REDIS_URL`). Every setting is listed in [App Configuration](/Deploying/DocsGPT-Settings).
## Run
```bash
docsgpt migrate # create the database if it is missing and apply the migrations
docsgpt api # the API and the web UI on http://127.0.0.1:7091
docsgpt worker # in a second terminal: the Celery worker, with the scheduler
```
`docsgpt api` listens on localhost only. Pass `--host 0.0.0.0` to accept connections from other machines or containers, and put a reverse proxy with TLS in front of it for anything public.
The API applies pending migrations when it starts (`AUTO_MIGRATE`), so `docsgpt migrate` is the explicit step for deployments that want the schema in place before the first request or that run the API with a restricted database role.
Both commands print the data home they resolved on start-up. They share it as long as `DOCSGPT_HOME` is the same for both (or unset), so the worker finds the files the API stores and the API finds the indexes the worker builds.
The worker is not optional: query embedding runs on it, so search fails without one. `docsgpt worker --help` lists the queue, concurrency and pool options; `--no-beat` starts a worker without the scheduler when another worker already runs it. On Windows the scheduler cannot be embedded, so run `docsgpt beat` in a third terminal.
Other commands:
- `docsgpt api --reload`: a development server with auto-reload.
- `docsgpt prefetch-models`: download the embedding models and their tokenizers ahead of time, for machines that go offline (see [Air-Gapped Deployment](/Deploying/Air-Gapped)).
- `docsgpt verify-offline`: check that a prepared install starts with networking off.
- `docsgpt reembed`: re-embed every index after changing `EMBEDDINGS_NAME` (see [Upgrading](/upgrading)).
## Run it as services, without Docker
`docsgpt up --native` runs the API and the worker as services on the machine itself: launchd agents on macOS, systemd user units on Linux. It does not start PostgreSQL or Redis; point it at ones you already run.
```bash
docsgpt up --native \
--postgres-uri postgresql://docsgpt:<password>@localhost:5432/docsgpt \
--redis-url redis://localhost:6379
```
Without a terminal only `--postgres-uri` is required, since Redis defaults to `redis://localhost:6379`; with one, it asks for both and for the model provider. It writes the same `.env` a Docker install uses (minus the image settings), generates `INTERNAL_KEY` and `JWT_SECRET_KEY` on the first run, applies the migrations, then starts `docsgpt-api` and `docsgpt-worker` and waits for the API to answer.
One Redis URL covers all three uses: the Celery broker, its result backend and the cache go on databases 0, 1 and 2 of it. Name a database in the URL and the three start there instead, so `redis://localhost:6379/5` puts them on 5, 6 and 7 — that is how you share a Redis that already holds something else.
The same commands manage it:
| Command | In native mode |
| --- | --- |
| `docsgpt status` | Which services run, the address, and whether the API answers |
| `docsgpt logs [api\|worker]` | The service log files under `<stack>/logs` |
| `docsgpt down` | Stops both services; settings stay |
| `docsgpt uninstall [--purge]` | Removes the services; `--purge` also deletes the stack directory |
`uninstall` never touches the database or Redis: they were yours to begin with.
<Callout type="info">
Windows has neither launchd nor systemd, so native mode is macOS and Linux only. On Windows, run DocsGPT on Docker with `docsgpt up`, or start `docsgpt api`, `docsgpt worker` and `docsgpt beat` yourself — the worker cannot run the scheduler in-process there, so `docsgpt beat` has to run alongside it for scheduled tasks to fire.
</Callout>
## Upgrade
```bash
pip install -U docsgpt
docsgpt migrate
```
Read the [upgrade notes](/upgrading) first when moving between minor versions.