Graph retrieval tied plain vector search at best and never beat it. Measured
across five corpora, the bottleneck was seeding, not the graph: the walk
started from nodes whose embeddings were computed from bare entity names, and
a whole question shares almost nothing with a name like "Quill".
Extraction now embeds each node from "name (type): description" and each
relationship as the fact it asserts ("Alder streams_to Quill: ..."), stored on
a new nullable graph_edges.fact_embedding column that ensure_vector_schema adds
in place. Entity names are canonicalised (case, punctuation, word breaks and a
cautious plural) so "VECTOR_STORE" and "vector stores" land on one node. Extraction calls run
concurrently (GRAPHRAG_EXTRACTION_WORKERS, default 8) while embedding and graph
writes stay serial on the task thread, so ordering and idempotency are
unchanged; that measured 8.4x faster with identical output.
Retrieval gains per-source options, stored under retrieval.graph and read live
at query time:
- seed_strategy: start from matching entities (default) or matching
relationships, which can reach an entity the question never names;
- passage_nodes (on): walk the source's passages alongside entities, with
PageRank damping 0.5 instead of 0.85;
- blend_vector (on): fuse the graph ranking with the source's vector ranking
by reciprocal rank.
The defaults are the measured-best configuration. Through GraphRAGRetriever,
the new seeding moved recall@4 from 0.41 to 0.68 on a multi-hop corpus and
from 0.50 to 1.00 on the docs corpus, and regressed none of the corpora
measured. Existing graphs keep name-only embeddings until rebuilt.
nextra-docsgpt
Setting Up Docs Folder of DocsGPT Locally
1. Clone the DocsGPT repository:
git clone https://github.com/arc53/DocsGPT.git
2. Navigate to the docs folder:
cd DocsGPT/docs
The docs folder contains the markdown files that make up the documentation. The majority of the files are in the pages directory. Some notable files in this folder include:
index.mdx: The main documentation file.
_app.js: This file is used to customize the default Next.js application shell.
theme.config.jsx: This file is for configuring the Nextra theme for the documentation.
3. Verify that you have Node.js and npm installed in your system. You can check by running:
node --version
npm --version
4. If not installed, download Node.js and npm from the respective official websites.
5. Once you have Node.js and npm running, proceed to install yarn - another package manager that helps to manage project dependencies:
npm install --global yarn
6. Install the project dependencies using yarn:
yarn install
7. After the successful installation of the project dependencies, start the local server:
yarn dev
-
Now, you should be able to view the docs on your local environment by visiting
http://localhost:3000. You can explore the different markdown files and make changes as you see fit. -
Footnotes: This guide assumes you have Node.js and npm installed. The guide involves running a local server using yarn, and viewing the documentation offline. If you encounter any issues, it may be worth verifying your Node.js and npm installations and whether you have installed yarn correctly.