Files
DocsGPT/tests/parser/file/test_markdown_parser.py
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00

63 lines
1.8 KiB
Python

from pathlib import Path
from unittest.mock import mock_open, patch
import pytest
from docsgpt.parser.file.markdown_parser import MarkdownParser
from docsgpt import utils
class _Enc:
def encode(self, s: str):
return list(s)
def encode_ordinary(self, s: str):
return list(s)
@pytest.fixture(autouse=True)
def _patch_tokenizer(monkeypatch):
monkeypatch.setattr(utils, "get_encoding", lambda: _Enc())
def test_markdown_init_parser():
parser = MarkdownParser()
assert isinstance(parser._init_parser(), dict)
assert not parser.parser_config_set
parser.init_parser()
assert parser.parser_config_set
def test_markdown_parse_file_basic_structure():
content = "# Title\npara1\npara2\n## Sub\ntext\n"
parser = MarkdownParser()
with patch("builtins.open", mock_open(read_data=content)):
result = parser.parse_file(Path("doc.md"))
assert isinstance(result, list) and len(result) >= 2
assert "Title" in result[0]
assert "para1" in result[0] and "para2" in result[0]
assert "Sub" in result[1]
assert "text" in result[1]
def test_markdown_removes_links_and_images_in_parse():
content = "# T\nSee [link](http://x) and ![[img.png]] here.\n"
parser = MarkdownParser()
with patch("builtins.open", mock_open(read_data=content)):
result = parser.parse_file(Path("doc.md"))
joined = "\n".join(result)
assert "(http://x)" not in joined
assert "![[img.png]]" not in joined
assert "link" in joined
def test_markdown_token_chunking_via_max_tokens():
raw = "abcdefghij" # 10 chars
parser = MarkdownParser(max_tokens=4)
with patch("builtins.open", mock_open(read_data=raw)):
tups = parser.parse_tups(Path("doc.md"))
assert len(tups) > 1
for _hdr, chunk in tups:
assert len(chunk) <= 4