Alex a83e1dc0af feat(graphrag): seed the walk from what entities are, and rank with passages and vector hits
Graph retrieval tied plain vector search at best and never beat it. Measured
across five corpora, the bottleneck was seeding, not the graph: the walk
started from nodes whose embeddings were computed from bare entity names, and
a whole question shares almost nothing with a name like "Quill".

Extraction now embeds each node from "name (type): description" and each
relationship as the fact it asserts ("Alder streams_to Quill: ..."), stored on
a new nullable graph_edges.fact_embedding column that ensure_vector_schema adds
in place. Entity names are canonicalised (case, punctuation, word breaks and a
cautious plural) so "VECTOR_STORE" and "vector stores" land on one node. Extraction calls run
concurrently (GRAPHRAG_EXTRACTION_WORKERS, default 8) while embedding and graph
writes stay serial on the task thread, so ordering and idempotency are
unchanged; that measured 8.4x faster with identical output.

Retrieval gains per-source options, stored under retrieval.graph and read live
at query time:

- seed_strategy: start from matching entities (default) or matching
  relationships, which can reach an entity the question never names;
- passage_nodes (on): walk the source's passages alongside entities, with
  PageRank damping 0.5 instead of 0.85;
- blend_vector (on): fuse the graph ranking with the source's vector ranking
  by reciprocal rank.

The defaults are the measured-best configuration. Through GraphRAGRetriever,
the new seeding moved recall@4 from 0.41 to 0.68 on a multi-hop corpus and
from 0.50 to 1.00 on the docs corpus, and regressed none of the corpora
measured. Existing graphs keep name-only embeddings until rebuilt.
2026-09-19 14:07:41 +01:00
2026-09-16 21:51:43 +01:00
2025-06-18 22:17:23 +01:00
2026-04-12 00:29:23 +01:00
2023-10-29 14:03:05 +05:30
2023-02-02 11:03:24 +00:00
2025-07-25 14:48:49 +05:30

DocsGPT 🦖

Private AI for agents, assistants and enterprise search

DocsGPT is an open-source AI platform for building intelligent agents and assistants. Features Agent Builder, deep research tools, document analysis (PDF, Office, web content, and audio), Multi-model support (choose your provider or run locally), and rich API connectivity for agents with actionable tools and integrations. Deploy anywhere with complete privacy control.


video-example-of-docs-gpt

Key Features:

  • 🗂️ Wide Format Support: Reads PDF, DOCX, CSV, XLSX, EPUB, MD, RST, HTML, MDX, JSON, PPTX, images, and audio files such as MP3, WAV, M4A, OGG, and WebM.
  • 🎙️ Speech Workflows: Record voice input into chat, transcribe audio on the backend, and ingest meeting recordings or voice notes as searchable knowledge.
  • 🌐 Web & Data Integration: Ingests from URLs, sitemaps, Reddit, GitHub and web crawlers.
  • ✅ Reliable Answers: Get accurate, hallucination-free responses with source citations viewable in a clean UI.
  • 🔑 Streamlined API Keys: Generate keys linked to your settings, documents, and models, simplifying chatbot and integration setup.
  • 🔗 Actionable Tooling: Connect to APIs, tools, and other services to enable LLM actions.
  • 🧩 Pre-built Integrations: Use readily available HTML/React chat widgets, search tools, Discord/Telegram bots, and more.
  • 🔌 Flexible Deployment: Works with major LLMs (OpenAI, Google, Anthropic) and local models (Ollama, llama_cpp).
  • 🏢 Secure & Scalable: Run privately and securely with Kubernetes support, designed for enterprise-grade reliability.

Roadmap

  • Agent Workflow Builder with conditional nodes ( February 2026 )
  • Research mode ( March 2026 )
  • SharePoint & Confluence connectors ( March – April 2026 )
  • Postgres migration for user data ( April 2026 )
  • OpenTelemetry observability ( April 2026 )
  • Bring Your Own Model (BYOM) ( April 2026 )
  • Agent scheduling (RedBeat-backed) ( April 2026 )
  • Notifications & conversation search ( May 2026 )
  • Analytics & logs revamp with per-agent attribution ( June 2026 )
  • OIDC / SSO login with SCIM provisioning & groups ( June 2026 )
  • Admin dashboard & role-based access control (RBAC) ( June 2026 )
  • Agent import / export ( June 2026 )
  • Teams with team-scoped sharing & roles ( June 2026 )

You can find our full roadmap here. Please don't hesitate to contribute or create issues, it helps us improve DocsGPT!

Production Support / Help for Companies:

We're eager to provide personalized assistance when deploying your DocsGPT to a live environment.

Get a Demo 👋⁠

Send Email 📧

Join the Lighthouse Program 🌟

Calling all developers and GenAI innovators! The DocsGPT Lighthouse Program connects technical leaders actively deploying or extending DocsGPT in real-world scenarios. Collaborate directly with our team to shape the roadmap, access priority support, and build enterprise-ready solutions with exclusive community insights.

Learn More & Apply →

QuickStart

Note

DocsGPT runs on Docker. The installer checks for it first.

macOS and Linux:

curl -fsSL https://docs.ac/install | bash

Windows (PowerShell):

irm https://docs.ac/install.ps1 | iex

The installer gets uv, installs the docsgpt Python package with it, and runs docsgpt up. That asks who should reach DocsGPT (only this computer, your network, or a domain with HTTPS) and which model provider to use, then starts it, at http://localhost:7091 for a local install. Afterwards, docsgpt status, docsgpt logs, docsgpt upgrade, docsgpt down and docsgpt uninstall manage it.

To read the script before running it:

curl -fsSL https://docs.ac/install -o install.sh
less install.sh
bash install.sh

A more detailed Quickstart is available in our documentation.

From a clone, with the setup script

  1. Clone the repository:

    git clone https://github.com/arc53/DocsGPT.git
    cd DocsGPT
    

For macOS and Linux:

  1. Run the setup script:

    ./setup.sh
    

For Windows:

  1. Run the PowerShell setup script:

    PowerShell -ExecutionPolicy Bypass -File .\setup.ps1
    

Either script will guide you through setting up DocsGPT. Five options are available: using the public API, running locally, connecting to a local inference engine, using a cloud API provider, or building the docker image locally. The scripts will automatically configure your .env file and handle necessary downloads and installations based on your chosen option.

Navigate to http://localhost:5173/

To stop DocsGPT, open a terminal in the DocsGPT directory and run:

docker compose -f deployment/docker-compose.yaml down

(or use the specific docker compose down command shown after running the setup script).

Note

For development environment setup instructions, please refer to the Development Environment Guide.

Contributing

Please refer to the CONTRIBUTING.md file for information about how to get involved. We welcome issues, questions, and pull requests.

Architecture

Architecture chart

Project Structure

  • docsgpt - Backend Flask application (the docsgpt Python package).

  • Extensions - Integrations and widgets (e.g., Chatwoot, React widget).

  • Frontend - Web UI built with Vite and React.

  • Scripts - Miscellaneous utility scripts.

Code Of Conduct

We as members, contributors, and leaders, pledge to make participation in our community a harassment-free experience for everyone, regardless of age, body size, visible or invisible disability, ethnicity, sex characteristics, gender identity and expression, level of experience, education, socio-economic status, nationality, personal appearance, race, religion, or sexual identity and orientation. Please refer to the CODE_OF_CONDUCT.md file for more information about contributing.

Many Thanks To Our Contributors⚡

Contributors

License

The source code license is MIT, as described in the LICENSE file.

This project is supported by:

color

S
Description
Private AI platform for agents, assistants and enterprise search. Built-in Agent Builder, Deep research, Document analysis, Multi-model support, and API connectivity for agents.
Readme MIT
107 MiB
0 Stars 1 Watchers 0 Forks
Languages
Python 73.2%
TypeScript 25.8%
Shell 0.4%
PowerShell 0.3%
JavaScript 0.1%