Files
DocsGPT/docs/content/Guides/compression.md
T
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00

2.8 KiB

Context Compression

DocsGPT implements a smart context compression system to manage long conversations effectively. This feature prevents conversations from hitting the LLM's context window limit while preserving critical information and continuity.

How It Works

The compression system operates on a "summarize and truncate" principle:

  1. Threshold Check: Before each request, the system calculates the total token count of the conversation history.
  2. Trigger: If the token count exceeds a configured threshold (default: 80% of the model's context limit), compression is triggered.
  3. Summarization: An LLM (potentially a different, cheaper/faster one) processes the older part of the conversation—including previous summaries, user messages, agent responses, and tool outputs.
  4. Context Replacement: The system generates a comprehensive summary of the older history. For subsequent requests, the LLM receives this Summary + Recent Messages instead of the full raw history.

Key Features

  • Recursive Summarization: New summaries incorporate previous summaries, ensuring that information from the very beginning of a long chat is not lost.
  • Tool Call Support: The compression logic explicitly handles tool calls and their outputs (e.g., file readings, search results), summarizing their results so the agent retains knowledge of what it has already done.
  • "Needle in a Haystack" Preservation: The prompts are designed to identify and preserve specific, critical details (like passwords, keys, or specific user instructions) even when compressing large amounts of text.

Configuration

You can configure the compression behavior in your .env file or docsgpt/core/settings.py:

Setting Default Description
ENABLE_CONVERSATION_COMPRESSION True Master switch to enable/disable the feature.
COMPRESSION_THRESHOLD_PERCENTAGE 0.8 The fraction of the context window (0.0 to 1.0) that triggers compression.
COMPRESSION_MODEL_OVERRIDE None (Optional) Specify a different model ID to use specifically for the summarization task (e.g., using gpt-3.5-turbo to compress for gpt-4).
COMPRESSION_MAX_HISTORY_POINTS 3 The number of past compression points to keep in the database (older ones are discarded as they are incorporated into newer summaries).

Architecture

The system is modularized into several components:

  • CompressionThresholdChecker: Calculates token usage and decides when to compress.
  • CompressionService: Orchestrates the compression process, manages DB updates, and reconstructs the context (Summary + Recent Messages) for the LLM.
  • CompressionPromptBuilder: Constructs the specific prompts used to instruct the LLM to summarize the conversation effectively.