mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-03 11:11:58 +00:00
docsgpt/core/settings.py had grown to 258 fields in one 600-line class, touched by about two commits a week, with related settings scattered (GitHub ingest caps inside the embeddings block, API keys in four places, the OpenAI Responses knobs 100 lines from the other OpenAI fields). It is now a package: one module per domain (auth, llm, embeddings, retrieval, vectorstores, database, workers, ingestion, ocr, storage, connectors, server, events, agents, guardrails, scheduler, sandbox, speech), each a SettingsGroup owning its fields and validators, composed by multiple inheritance into the same flat Settings class. Every attribute name, type, default, alias and constraint is unchanged, so settings.NAME reads, .env files and test monkeypatches all keep working; the import path docsgpt.core.settings is the package. Settings.normalize_api_key is kept as a classmethod for callers that reuse it. The comment above or beside each field became its Field(description=...), so the definitions are visible to tooling; the next commit generates the docs reference from them. Pitfall recorded for future groups: pydantic collects validators by method name across the MRO, so two groups naming a validator the same would silently keep only one. Each group's validator has a unique name.
2.8 KiB
2.8 KiB
Context Compression
DocsGPT implements a smart context compression system to manage long conversations effectively. This feature prevents conversations from hitting the LLM's context window limit while preserving critical information and continuity.
How It Works
The compression system operates on a "summarize and truncate" principle:
- Threshold Check: Before each request, the system calculates the total token count of the conversation history.
- Trigger: If the token count exceeds a configured threshold (default: 80% of the model's context limit), compression is triggered.
- Summarization: An LLM (potentially a different, cheaper/faster one) processes the older part of the conversation—including previous summaries, user messages, agent responses, and tool outputs.
- Context Replacement: The system generates a comprehensive summary of the older history. For subsequent requests, the LLM receives this Summary + Recent Messages instead of the full raw history.
Key Features
- Recursive Summarization: New summaries incorporate previous summaries, ensuring that information from the very beginning of a long chat is not lost.
- Tool Call Support: The compression logic explicitly handles tool calls and their outputs (e.g., file readings, search results), summarizing their results so the agent retains knowledge of what it has already done.
- "Needle in a Haystack" Preservation: The prompts are designed to identify and preserve specific, critical details (like passwords, keys, or specific user instructions) even when compressing large amounts of text.
Configuration
You can configure the compression behavior in your .env file or docsgpt/core/settings/agents.py:
| Setting | Default | Description |
|---|---|---|
ENABLE_CONVERSATION_COMPRESSION |
True |
Master switch to enable/disable the feature. |
COMPRESSION_THRESHOLD_PERCENTAGE |
0.8 |
The fraction of the context window (0.0 to 1.0) that triggers compression. |
COMPRESSION_MODEL_OVERRIDE |
None |
(Optional) Specify a different model ID to use specifically for the summarization task (e.g., using gpt-3.5-turbo to compress for gpt-4). |
COMPRESSION_MAX_HISTORY_POINTS |
3 |
The number of past compression points to keep in the database (older ones are discarded as they are incorporated into newer summaries). |
Architecture
The system is modularized into several components:
CompressionThresholdChecker: Calculates token usage and decides when to compress.CompressionService: Orchestrates the compression process, manages DB updates, and reconstructs the context (Summary + Recent Messages) for the LLM.CompressionPromptBuilder: Constructs the specific prompts used to instruct the LLM to summarize the conversation effectively.