Files
DocsGPT/tests/parser/conftest.py
T
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00

24 lines
701 B
Python

import pytest
@pytest.fixture(scope="session")
def _mpnet_tokenizer():
"""Real WordPiece tokenizer; skipped when the hub is unreachable."""
pytest.importorskip("tokenizers")
from tokenizers import Tokenizer
try:
tokenizer = Tokenizer.from_pretrained("sentence-transformers/all-mpnet-base-v2")
except Exception as exc: # offline CI
pytest.skip(f"tokenizer unavailable: {exc}")
tokenizer.no_padding()
tokenizer.no_truncation()
return tokenizer
@pytest.fixture
def hf_counter(_mpnet_tokenizer):
from docsgpt.parser.tokenization import HuggingFaceCounter
return HuggingFaceCounter(_mpnet_tokenizer, "sentence-transformers/all-mpnet-base-v2")