Files
DocsGPT/application/parser/chunking_creator.py
T
Alex f6400cd736 feat: per-source RAG configuration (retrieval strategies, chunking, exposure, prescreen)
Introduces a per-source config contract that makes RAG behavior strategy-dispatched instead of a single hardcoded path. Every source gains a validated JSONB config; an empty/absent config reproduces current behavior byte-for-byte, and the whole path is gated by PER_SOURCE_RETRIEVAL_ENABLED.

Foundation: sources.config JSONB column + migration 0022_source_config; SourceConfig/ChunkingConfig/RetrievalConfig pydantic models (strict on write, lenient on read); ChunkerCreator and RetrieverCreator.register registries; config threaded through the upload routes, ingest/remote/connector workers, and reingest.

Retrieval: a Dispatcher groups sources by retriever key (all-classic collapses to today's single ClassicRAG under one shared token budget; non-classic retrievers get their own instance), removing the previous single-global-retriever collapse in stream_processor. Per-source chunks, score_threshold (honored for pgvector/mongodb, safely ignored elsewhere), and rephrase_query toggle. New PATCH /api/sources/<id>/config with team-aware (effective_write_owner) authz and a requires_reingest signal.

Chunking strategies: recursive, markdown, parent_child (selectable per source; re-ingest to apply). Search exposure: per-source prefetch vs agentic_tool for agentic/research agents. Map-reduce prescreen: optional LLM relevance pre-filter implemented as a composable post-retrieval stage that wraps any retriever.

Backend and frontend (shared Retrieval options panel + edit modal) with tests; backend suite and frontend vitest green. Excludes the wiki and GraphRAG flagships.
2026-06-20 21:54:23 +01:00

58 lines
2.1 KiB
Python

"""String-keyed registry for chunking strategies.
Mirrors ``RetrieverCreator``: features register new strategies (``recursive``,
``markdown``, ``parent_child``, ...) without touching the dispatch site. The
classic strategy is registered under ``classic_chunk`` by ``chunking.py``.
"""
from __future__ import annotations
from typing import Type
class ChunkerCreator:
chunkers: dict[str, Type] = {}
_strategies_loaded: bool = False
@classmethod
def _ensure_builtin(cls) -> None:
"""Register built-in chunkers if they are not registered yet.
Self-bootstraps so ``create_chunker`` works regardless of import order:
``application.parser.chunking`` registers ``classic_chunk`` and
``application.parser.chunking_strategies`` registers ``recursive`` /
``markdown`` / ``parent_child``.
"""
if not cls.chunkers:
import application.parser.chunking # noqa: F401 (registers classic_chunk)
if not cls._strategies_loaded:
cls._strategies_loaded = True
import application.parser.chunking_strategies # noqa: F401
@classmethod
def create_chunker(cls, strategy: str, *args, **kwargs):
"""Instantiate the chunker registered under ``strategy``.
Args:
strategy: Registry key (e.g. ``classic_chunk``).
*args: Positional args forwarded to the chunker constructor.
**kwargs: Keyword args forwarded to the chunker constructor.
Returns:
A chunker instance exposing ``chunk(documents) -> List[Document]``.
Raises:
ValueError: If no chunker is registered for ``strategy``.
"""
cls._ensure_builtin()
key = (strategy or "classic_chunk").lower()
chunker_class = cls.chunkers.get(key)
if not chunker_class:
raise ValueError(f"No chunker class found for strategy {strategy}")
return chunker_class(*args, **kwargs)
@classmethod
def register(cls, key: str, chunker_class: Type) -> None:
"""Register ``chunker_class`` under ``key`` (idempotent)."""
cls.chunkers[key] = chunker_class