The API embeds every query it serves, so it held its own copy of the model: ~890 MB it never needed. EMBEDDINGS_DELEGATE_TO_WORKER (on by default) sends the text to the Celery worker instead and gets the vector back, taking an API process from 1176 MB to 285 MB with no ONNX Runtime imported at all. The client embeds locally when it finds itself inside a worker task, so the worker never dispatches to itself -- the same self-deadlock DOCUMENT_PARSE_QUEUE avoids on the parsing side. EMBEDDINGS_BASE_URL still wins over it, and remains the right answer for production. ensure_vector_schema was constructing the embeddings instance purely to read .dimension off it, loading several hundred MB of ONNX into every API and worker process at import. For a model the registry describes that is a lookup; only an unregistered name now falls back to loading. EMBEDDINGS_BATCH_SIZE was sizing two unrelated things: chunks per store transaction (and per remote embed request) and documents per ONNX forward pass. Each pass pads every input up to its longest, and that waste grows with the square of chunk length, so at the 1250-token default a batch of 32 peaked at 6.6 GB and took 326s where a batch of 1 peaked at 2.9 GB and took 90s. The forward pass is now sized by EMBEDDINGS_MODEL_BATCH_SIZE, defaulting to 1; storage and remote batching are unchanged at 32. reembed embeds in-process: a batch job that walks the whole index should not round-trip every chunk through a broker, and loading the model there reports a real failure instead of timing out against an empty queue. Also drops the mpnet zip download from the docs and the devcontainer, which pointed at a SentenceTransformers export with no ONNX graph and had been inert since the FastEmbed swap; corrects the claim that any sentence-transformers model works; and settles the Configuring/Settings pages on what the registry and the repository metadata actually decide.
nextra-docsgpt
Setting Up Docs Folder of DocsGPT Locally
1. Clone the DocsGPT repository:
git clone https://github.com/arc53/DocsGPT.git
2. Navigate to the docs folder:
cd DocsGPT/docs
The docs folder contains the markdown files that make up the documentation. The majority of the files are in the pages directory. Some notable files in this folder include:
index.mdx: The main documentation file.
_app.js: This file is used to customize the default Next.js application shell.
theme.config.jsx: This file is for configuring the Nextra theme for the documentation.
3. Verify that you have Node.js and npm installed in your system. You can check by running:
node --version
npm --version
4. If not installed, download Node.js and npm from the respective official websites.
5. Once you have Node.js and npm running, proceed to install yarn - another package manager that helps to manage project dependencies:
npm install --global yarn
6. Install the project dependencies using yarn:
yarn install
7. After the successful installation of the project dependencies, start the local server:
yarn dev
-
Now, you should be able to view the docs on your local environment by visiting
http://localhost:3000. You can explore the different markdown files and make changes as you see fit. -
Footnotes: This guide assumes you have Node.js and npm installed. The guide involves running a local server using yarn, and viewing the documentation offline. If you encounter any issues, it may be worth verifying your Node.js and npm installations and whether you have installed yarn correctly.