Files
DocsGPT/docsgpt/core/settings/retrieval.py
T
arc53-machine 5578039c19 refactor(settings): treat unset spellings of every optional string as None
Review follow-up. The per-group secret validators normalised a hand-picked
list of API keys, which left other optional credentials and overrides
(OPEN_ROUTER_API_KEY, S3 and Daytona keys, ELASTIC_PASSWORD, the OIDC
trio, connector client ids, MICROSOFT_AUTHORITY, MCP_OAUTH_REDIRECT_URI)
holding the literal "None" or "" a .env file spells "unset" with, so
truthiness checks and fallbacks downstream saw a value. One rule on the
group base replaces those lists: every Optional[str] field maps "", "None"
and whitespace to None and strips real values. Plain str fields are left
alone. The OIDC required-settings check therefore also rejects those
spellings.

EMBEDDINGS_POOLING is Literal["cls", "mean"] with case-insensitive
parsing; its consumer silently ignored anything else.

Bounds added where the consumer rejects or misbehaves on the value:
SCHEDULE_RUN_OUTPUT_RETENTION_DAYS and MESSAGE_EVENTS_RETENTION_DAYS (the
cleanup repositories raise on <= 0), EMBEDDINGS_DELEGATE_TIMEOUT, the
remote-device idle/pairing/invocation TTLs and CELERY_VISIBILITY_TIMEOUT
(> 0), REMOTE_DEVICE_CMD_QUEUE_TTL_SECONDS (> 605, the documented drain
deadline), GRAPHRAG_MAX_CHUNKS_FOR_EXTRACTION (>= 0; negative would slice
the pending list from the end).

The generated reference now renders generic type arguments
(dict[str, int] rather than dict).
2026-09-17 11:37:57 +01:00

39 lines
1.5 KiB
Python

"""Retrieval strategy and GraphRAG."""
from __future__ import annotations
from typing import Literal, Optional
from pydantic import Field, field_validator
from docsgpt.core.settings._shared import SettingsGroup, normalize_choice
class RetrievalSettings(SettingsGroup):
"""Which vector store answers searches and how retrieval fans out across sources."""
VECTOR_STORE: Literal["faiss", "elasticsearch", "mongodb", "qdrant", "milvus", "pgvector"] = Field(
default="faiss", description="Vector store backend."
)
RETRIEVAL_MAX_PARALLEL_SOURCES: int = Field(
default=4,
ge=1,
description="Concurrent per-source searches in one retrieval; the query is embedded once and shared.",
)
PER_SOURCE_RETRIEVAL_ENABLED: bool = Field(
default=True,
description="Kill-switch for per-source retrieval dispatch; False collapses to a single retriever.",
)
GRAPHRAG_ENABLED: bool = Field(default=False, description="Gates graph-aware ingestion and retrieval.")
GRAPHRAG_EXTRACTION_MODEL: Optional[str] = Field(
default=None, description="Model for ingest-time graph extraction; unset reuses LLM_PROVIDER/LLM_NAME."
)
GRAPHRAG_MAX_CHUNKS_FOR_EXTRACTION: int = Field(
default=2000, ge=0, description="Hard cap on chunks extracted per source (cost control); 0 extracts nothing."
)
@field_validator("VECTOR_STORE", mode="before")
@classmethod
def _normalize_vector_store(cls, v):
return normalize_choice(v)