Files
DocsGPT/docsgpt/parser/chunking_creator.py
T
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00

58 lines
2.1 KiB
Python

"""String-keyed registry for chunking strategies.
Mirrors ``RetrieverCreator``: features register new strategies (``recursive``,
``markdown``, ``parent_child``, ...) without touching the dispatch site. The
classic strategy is registered under ``classic_chunk`` by ``chunking.py``.
"""
from __future__ import annotations
from typing import Type
class ChunkerCreator:
chunkers: dict[str, Type] = {}
_strategies_loaded: bool = False
@classmethod
def _ensure_builtin(cls) -> None:
"""Register built-in chunkers if they are not registered yet.
Self-bootstraps so ``create_chunker`` works regardless of import order:
``docsgpt.parser.chunking`` registers ``classic_chunk`` and
``docsgpt.parser.chunking_strategies`` registers ``recursive`` /
``markdown`` / ``parent_child``.
"""
if not cls.chunkers:
import docsgpt.parser.chunking # noqa: F401 (registers classic_chunk)
if not cls._strategies_loaded:
cls._strategies_loaded = True
import docsgpt.parser.chunking_strategies # noqa: F401
@classmethod
def create_chunker(cls, strategy: str, *args, **kwargs):
"""Instantiate the chunker registered under ``strategy``.
Args:
strategy: Registry key (e.g. ``classic_chunk``).
*args: Positional args forwarded to the chunker constructor.
**kwargs: Keyword args forwarded to the chunker constructor.
Returns:
A chunker instance exposing ``chunk(documents) -> List[Document]``.
Raises:
ValueError: If no chunker is registered for ``strategy``.
"""
cls._ensure_builtin()
key = (strategy or "classic_chunk").lower()
chunker_class = cls.chunkers.get(key)
if not chunker_class:
raise ValueError(f"No chunker class found for strategy {strategy}")
return chunker_class(*args, **kwargs)
@classmethod
def register(cls, key: str, chunker_class: Type) -> None:
"""Register ``chunker_class`` under ``key`` (idempotent)."""
cls.chunkers[key] = chunker_class