--- title: Install with pip description: Run the DocsGPT backend from the PyPI package, inside your own Python environment. --- import { Callout } from 'nextra/components' # Install with pip DocsGPT is on PyPI as [`docsgpt`](https://pypi.org/project/docsgpt/): the API server, the web UI, the Celery worker and the maintenance scripts in one package, behind a single `docsgpt` command. Use it when you want DocsGPT inside your own Python environment or process manager rather than the [Docker images](/Deploying/Docker-Deploying). `docsgpt api` serves the web UI on the same port as the API. Set `SERVE_UI=false` to run the API alone, for example behind the frontend Docker image or a UI you host yourself. With Docker available, the same package can run the whole stack for you, including Postgres and Redis: `docsgpt up`. See [Run it with `docsgpt up`](/Deploying/Docker-Deploying#run-it-with-docsgpt-up). ## Requirements - Python 3.12 or newer - [PostgreSQL](/Deploying/Postgres-Migration) for user data, with the `vector` extension if you set `VECTOR_STORE=pgvector` - Redis for the task queue and the cache - An LLM: an API key for a hosted provider, or a local model server ## Install ```bash python -m venv .venv && source .venv/bin/activate pip install docsgpt ``` Extras add the optional engines: ```bash pip install "docsgpt[docling]" # DOC_PARSER_ENGINE=docling: OCR backend, structured output pip install "docsgpt[milvus]" # VECTOR_STORE=milvus ``` The `docling` extra pulls in PyTorch, and on Linux the PyPI torch wheels bring the CUDA stack with them. On a CPU-only machine install the CPU build first, then the extra; pip keeps the torch it already has: ```bash pip install --index-url https://download.pytorch.org/whl/cpu torch torchvision pip install "docsgpt[docling]" ``` `uv pip install` accepts the same two commands. With `pipx`, install `docsgpt[docling]` first, then replace the CUDA build inside its environment (`--no-deps` keeps pip from touching torch's dependencies, which the PyTorch index carries in older copies): ```bash pipx runpip docsgpt install --force-reinstall --no-deps --index-url https://download.pytorch.org/whl/cpu torch torchvision ``` ## Configure DocsGPT keeps its runtime files in a **data home**: the `.env` file it reads settings from, uploaded files under `inputs/`, vector indexes under `indexes/` and downloaded embedding models under `models/`. The data home is `~/.docsgpt/server` (`/opt/docsgpt` when you run as root on Linux), whatever directory you run the commands from. `DOCSGPT_HOME` moves it, and `DOCSGPT_ENV_FILE` points at a `.env` kept somewhere else. Both variables must be set in the process environment, not in `.env`: they decide where `.env` is read from. In a source checkout the data home is the checkout. Up to 0.20 the data home of an installed package was the directory you ran the command from. If you kept `.env` and your data there, move them to `~/.docsgpt/server` or set `DOCSGPT_HOME` to that directory. `docsgpt api` and `docsgpt worker` point out a `.env` in the working directory that they no longer read. Create a `.env` in the data home. The minimum for a hosted LLM: ```ini # LLM provider and model: see App Configuration for the options LLM_PROVIDER=openai LLM_NAME= API_KEY= # User data POSTGRES_URI=postgresql://docsgpt:@localhost:5432/docsgpt # Worker-to-API authentication (required for uploads) INTERNAL_KEY= ``` Redis defaults to `localhost:6379` (`CELERY_BROKER_URL`, `CELERY_RESULT_BACKEND`, `CACHE_REDIS_URL`). Every setting is listed in [App Configuration](/Deploying/DocsGPT-Settings). ## Run ```bash docsgpt migrate # create the database if it is missing and apply the migrations docsgpt api # the API and the web UI on http://127.0.0.1:7091 docsgpt worker # in a second terminal: the Celery worker, with the scheduler ``` `docsgpt api` listens on localhost only. Pass `--host 0.0.0.0` to accept connections from other machines or containers, and put a reverse proxy with TLS in front of it for anything public. The API applies pending migrations when it starts (`AUTO_MIGRATE`), so `docsgpt migrate` is the explicit step for deployments that want the schema in place before the first request or that run the API with a restricted database role. Both commands print the data home they resolved on start-up. They share it as long as `DOCSGPT_HOME` is the same for both (or unset), so the worker finds the files the API stores and the API finds the indexes the worker builds. The worker is not optional: query embedding runs on it, so search fails without one. `docsgpt worker --help` lists the queue, concurrency and pool options; `--no-beat` starts a worker without the scheduler when another worker already runs it. On Windows the scheduler cannot be embedded, so run `docsgpt beat` in a third terminal. Other commands: - `docsgpt api --reload`: a development server with auto-reload. - `docsgpt prefetch-models`: download the embedding models and their tokenizers ahead of time, for machines that go offline (see [Air-Gapped Deployment](/Deploying/Air-Gapped)). - `docsgpt verify-offline`: check that a prepared install starts with networking off. - `docsgpt reembed`: re-embed every index after changing `EMBEDDINGS_NAME` (see [Upgrading](/upgrading)). ## Upgrade ```bash pip install -U docsgpt docsgpt migrate ``` Read the [upgrade notes](/upgrading) first when moving between minor versions.