Docs revamp

This commit is contained in:
Pavel committed 2026-06-26 11:56:20 +04:00
1 parent 1acb836332
commit 5c19d972d2
27 files changed
+1101 -39

No files matched your search

+20 -1
View File
@@ -25,13 +25,17 @@ DocsGPT offers direct, streamlined support for the following cloud LLM providers
| :--------------------------- | :------------- | :-------------------------- |
| DocsGPT Public API | `docsgpt` | `None` |
| OpenAI | `openai` | `gpt-5.1` |
| OpenAI-compatible (BYOM) | `openai_compatible` | (any; with per-model `base_url`/`api_key`) |
| Google (Vertex AI, Gemini) | `google` | `gemini-3.5-flash` |
| Anthropic (Claude) | `anthropic` | `claude-3-5-sonnet-20241022`|
| Groq | `groq` | `llama-3.3-70b-versatile` |
| OpenRouter | `openrouter` | (See OpenRouter docs) |
| Novita AI | `novita` | (See Novita docs) |
| HuggingFace Inference API | `huggingface` | `meta-llama/Llama-3.1-8B-Instruct` |
| Prem AI | `premai` | (See Prem AI docs) |
| AWS SageMaker | `sagemaker` | (See SageMaker docs) |
| Novita AI | `novita` | (See Novita docs) |
DocsGPT also ships a **model catalog** (`application/core/models/*.yaml`) that the in-app model picker reads, so common models from these providers — including DeepSeek — appear ready to select once the matching API key is set.
## Connecting to OpenAI-Compatible Cloud APIs
@@ -52,6 +56,21 @@ OPENAI_BASE_URL=https://api.deepseek.com/v1 # DeepSeek's OpenAI API URL
Remember to consult the documentation of your chosen OpenAI-compatible cloud provider for their specific API endpoint, required model names, and authentication methods.
### Dedicated `openai_compatible` provider (bring-your-own-model)
Beyond the global `OPENAI_BASE_URL`, DocsGPT has a first-class `openai_compatible` provider. It lets a model carry its **own** `base_url` and `api_key`, which is how per-user "bring your own model" (BYOM) endpoints work — each model can point at a different OpenAI-compatible server without changing instance-wide settings. Outbound requests use an SSRF-pinned HTTP client for safety.
This is the mechanism behind catalog entries like DeepSeek, which declare their own `base_url` and API key environment variable rather than relying on `OPENAI_BASE_URL`.
## OpenAI Responses API and reasoning
For OpenAI models that support it, DocsGPT can call the newer **Responses API** (`/v1/responses`) instead of Chat Completions. This is selected per model in the catalog via an `api_flavor: responses` capability and enables features like server-side reasoning. Related settings:
- `reasoning_effort` — per-model reasoning effort hint (for example `medium`) declared in the model catalog.
- `OPENAI_RESPONSES_STORE` (default `false`) — when `true`, lets OpenAI persist Responses API state server-side.
See [App Configuration](/Deploying/DocsGPT-Settings) for the full settings reference.
## Adding Support for Other Cloud Providers
If you wish to connect to a cloud provider that is not explicitly listed above or doesn't offer OpenAI API compatibility, you can extend DocsGPT to support it. Within the DocsGPT repository, navigate to the `application/llm` directory. Here, you will find Python files defining the existing LLM integrations. You can use these files as examples to create a new module for your desired cloud provider. After creating your new LLM module, you will need to register it within the `llm_creator.py` file. This process involves some coding, but it allows for virtually unlimited extensibility to connect to any cloud-based LLM service with an accessible API.
+35 -3
View File
@@ -22,10 +22,12 @@ In essence, embedding models are the bridge that allows DocsGPT to understand th
## Out-of-the-Box Embedding Model Support in DocsGPT
DocsGPT is designed to be flexible and supports a wide range of embedding models right out of the box. Currently, DocsGPT provides native support for models from two major sources:
DocsGPT is designed to be flexible and supports a wide range of embedding models right out of the box:
* **Sentence Transformers:** DocsGPT supports all models available through the [Sentence Transformers library](https://www.sbert.net/). This library offers a vast selection of pre-trained embedding models, known for their quality and efficiency in various semantic tasks.
* **OpenAI Embeddings:** DocsGPT also supports using embedding models from OpenAI, specifically the `text-embedding-ada-002` model, which is a powerful and widely used embedding model from OpenAI's API.
* **Sentence Transformers:** DocsGPT supports all models available through the [Sentence Transformers library](https://www.sbert.net/). This library offers a vast selection of pre-trained embedding models, known for their quality and efficiency in various semantic tasks. This is the default (`EMBEDDINGS_NAME=huggingface_sentence-transformers/all-mpnet-base-v2`).
* **OpenAI Embeddings:** DocsGPT supports OpenAI embedding models (for example `text-embedding-ada-002`, `text-embedding-3-small`, `text-embedding-3-large`) via the OpenAI API.
* **Azure OpenAI Embeddings:** Set `AZURE_EMBEDDINGS_DEPLOYMENT_NAME` alongside your Azure OpenAI configuration.
* **Remote OpenAI-compatible Embeddings:** Any server that exposes an OpenAI-compatible `/v1/embeddings` endpoint (for example llama.cpp, vLLM, TEI, or a hosted provider) by setting `EMBEDDINGS_BASE_URL`. See [Remote Embeddings](#remote-openai-compatible-embeddings) below.
## Configuring Sentence Transformer Models
@@ -65,6 +67,36 @@ API_KEY=YOUR_OPENAI_API_KEY # Your OpenAI API Key
EMBEDDINGS_NAME=openai_text-embedding-ada-002
```
## Remote (OpenAI-compatible) Embeddings
If you run your own embedding server, or use a provider that exposes an OpenAI-style embeddings API, point DocsGPT at it with `EMBEDDINGS_BASE_URL`. When this is set, all embedding calls (ingestion and querying) are sent to `{EMBEDDINGS_BASE_URL}/v1/embeddings` in OpenAI format instead of running a local model.
```env
EMBEDDINGS_BASE_URL=http://localhost:8080 # your OpenAI-compatible embeddings server
EMBEDDINGS_NAME=your-model-name # sent as the "model" field in the request
EMBEDDINGS_KEY=YOUR_API_KEY # optional; sent as a Bearer token
```
- `EMBEDDINGS_BASE_URL` — base URL of the remote server. Setting it switches DocsGPT into remote-embeddings mode.
- `EMBEDDINGS_NAME` — forwarded as the `model` field in each request.
- `EMBEDDINGS_KEY` — optional bearer token. If you are using OpenAI directly you can copy `API_KEY` here.
### Guarding against oversized inputs
Some remote servers (notably llama.cpp) reject any single input larger than their physical batch size with a `500` error. Set `EMBEDDINGS_MAX_INPUT_TOKENS` to clip each input to a fixed number of tokens before it is sent:
```env
EMBEDDINGS_MAX_INPUT_TOKENS=512
```
When set, each input string is truncated to that many tokens and the overflow is dropped (lossy by design). Token counts use DocsGPT's shared tiktoken encoding, which differs from your server's tokenizer, so choose a limit with some headroom below the server's true limit to absorb tokenizer skew. Leave the setting unset (or `0`) to disable truncation.
## Important: Embedding Dimensions Must Stay Consistent
Each embedding model produces vectors of a fixed dimension, and your vector store is created with that dimension. **Changing `EMBEDDINGS_NAME` to a model with a different dimension is not compatible with an existing index** — FAISS and LanceDB will raise a dimension-mismatch error, and pgvector/Qdrant tables are sized to the original dimension.
If you need to switch embedding models, you must re-ingest your sources so the index is rebuilt with the new dimension. This also applies to the [GraphRAG](/Sources/GraphRAG) graph tables, which are sized to the embedding dimension at creation time.
## Adding Support for Other Embedding Models
If you wish to use an embedding model that is not supported out-of-the-box, a good starting point for adding custom embedding model support is to examine the `base.py` file located in the `application/vectorstore` directory.