mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-03 22:13:01 +00:00
Introduces a per-source config contract that makes RAG behavior strategy-dispatched instead of a single hardcoded path. Every source gains a validated JSONB config; an empty/absent config reproduces current behavior byte-for-byte, and the whole path is gated by PER_SOURCE_RETRIEVAL_ENABLED. Foundation: sources.config JSONB column + migration 0022_source_config; SourceConfig/ChunkingConfig/RetrievalConfig pydantic models (strict on write, lenient on read); ChunkerCreator and RetrieverCreator.register registries; config threaded through the upload routes, ingest/remote/connector workers, and reingest. Retrieval: a Dispatcher groups sources by retriever key (all-classic collapses to today's single ClassicRAG under one shared token budget; non-classic retrievers get their own instance), removing the previous single-global-retriever collapse in stream_processor. Per-source chunks, score_threshold (honored for pgvector/mongodb, safely ignored elsewhere), and rephrase_query toggle. New PATCH /api/sources/<id>/config with team-aware (effective_write_owner) authz and a requires_reingest signal. Chunking strategies: recursive, markdown, parent_child (selectable per source; re-ingest to apply). Search exposure: per-source prefetch vs agentic_tool for agentic/research agents. Map-reduce prescreen: optional LLM relevance pre-filter implemented as a composable post-retrieval stage that wraps any retriever. Backend and frontend (shared Retrieval options panel + edit modal) with tests; backend suite and frontend vitest green. Excludes the wiki and GraphRAG flagships.
42 lines
1.5 KiB
Python
42 lines
1.5 KiB
Python
from typing import List, Optional
|
|
from uuid import uuid4
|
|
|
|
|
|
from application.core.settings import settings
|
|
from application.vectorstore.base import BaseVectorStore
|
|
|
|
|
|
class MilvusStore(BaseVectorStore):
|
|
def __init__(self, source_id: str = "", embeddings_key: str = "embeddings"):
|
|
super().__init__()
|
|
from langchain_milvus import Milvus
|
|
|
|
connection_args = {
|
|
"uri": settings.MILVUS_URI,
|
|
"token": settings.MILVUS_TOKEN,
|
|
}
|
|
self._docsearch = Milvus(
|
|
embedding_function=self._get_embeddings(settings.EMBEDDINGS_NAME, embeddings_key),
|
|
collection_name=settings.MILVUS_COLLECTION_NAME,
|
|
connection_args=connection_args,
|
|
)
|
|
self._source_id = source_id
|
|
|
|
def search(self, question, k=2, *args, **kwargs):
|
|
# Drop the per-source score_threshold (unsupported here) so it is safely
|
|
# ignored instead of being forwarded into the langchain call.
|
|
kwargs.pop("score_threshold", None)
|
|
expr = f"source_id == '{self._source_id}'"
|
|
return self._docsearch.similarity_search(query=question, k=k, expr=expr, *args, **kwargs)
|
|
|
|
def add_texts(self, texts: List[str], metadatas: Optional[List[dict]], *args, **kwargs):
|
|
ids = [str(uuid4()) for _ in range(len(texts))]
|
|
|
|
return self._docsearch.add_texts(texts=texts, metadatas=metadatas, ids=ids, *args, **kwargs)
|
|
|
|
def save_local(self, *args, **kwargs):
|
|
pass
|
|
|
|
def delete_index(self, *args, **kwargs):
|
|
pass
|