3 Commits
Author SHA1 Message Date
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00
Alex 636d4d35d1 feat(prompts): revamp preset prompts, tool naming, and prompt templating
Rewrite the default/creative/strict presets (classic + agentic) into
structured sections: grounding and cite-by-title guidance, insufficient-
context behavior, current date, respond-in-user-language, scoped mermaid
usage, an untrusted-content guardrail, and a conditional XML-tagged
document context block. A memory directory listing is injected at render
time via the template prefetch mechanism so the model starts oriented
without burning a tool call.

Fixes along the way:
- Agentic preset swap was dead code: _get_prompt_content cached the
  classic preset before create_agent's swap check ran, so agentic and
  research agents always got the classic preset. The swap now happens
  inside _get_prompt_content.
- Jinja autoescape corrupted document content in custom prompts
  (< -> &lt;); prompts are not HTML, autoescape is now off.
- Literal {summaries} leaked into the prompt when no docs were
  retrieved; the placeholder is now stripped.
- Agentic/research prompts referenced tool names from a dropped naming
  scheme (search_internal, reason_think); they now reference the real
  names (search, reason).
- The strict preset told the model to "be very creative and use your
  imagination" right after "never make up information".
- extract_tool_usages recorded intermediate attribute chains as
  bare-tool usages, which meant "run all actions" at prefetch; only
  maximal chains are recorded now.
- Headless runs retrieved docs but never rendered them into the
  prompt; the prompt is now rendered like the streaming path.
- Default tools were unreachable by name in prompt templates
  (prefetch results were keyed by synthetic id only); defaults now
  claim the name key unless an explicit row shadows it.

Tool layer: memory/notes/todo actions are namespaced (memory_view,
note_overwrite, todo_create, ...) with legacy unprefixed names still
accepted via prefix stripping; duplicate action names across tools are
disambiguated with the owning tool's name instead of numeric suffixes;
thin tool descriptions rewritten (brave, duckduckgo, telegram, ntfy,
cryptoprice, read_webpage, internal_search, think).

Docs are now wrapped per chunk in <document index>/<source>/<content>
tags for citation-by-title support.
2026-06-11 11:18:46 +01:00
Alex 00a621f33a tests: retriever and seeder 2026-03-29 09:46:07 +01:00