With GITHUB_ACCESS_TOKEN set in the environment, load_data checked the
repository's visibility against the real GitHub API. The tests now clear
the instance token unless they set one.
An issue deleted after it was listed, or comments the token cannot read,
now lose that detail with a warning instead of failing the whole sync; an
unreadable document is skipped.
The Linear connector now syncs as well as giving agents its tools, from
one connection. Linear's MCP server is its own OAuth issuer, so its
tokens are read through the same MCP tools the agents use (list_issues,
get_issue, list_comments, list_documents) rather than Linear's GraphQL
API, and no OAuth app has to be registered.
A source picks teams and projects, with comments (on by default) and the
projects' documents. Each issue becomes one document with its state,
assignee, priority, labels, description and comments, filed under its
team and citing its Linear URL. Each sync reads up to 500 issues and 100
documents again. /api/connections/<id>/linear lists the teams and
projects to pick from. Sources sync on their schedule with the owner's
connection, and pause when the sign-in needs reconnecting.
One GitHub connection feeds both a Knowledge source and an agent tool.
Users connect with a personal access token, checked against GitHub and
named after the account, or, when an admin registers a GitHub App
(GITHUB_CLIENT_ID, GITHUB_CLIENT_SECRET, GITHUB_APP_SLUG), with Sign in
with GitHub. App tokens expire after eight hours and are refreshed before
each sync or tool call; tokens with no expiry are never treated as expired.
The catalog gains an optional second sign-in method (oauth_settings,
exposed as sign_in_methods) without changing any other connector.
Sync lists the repositories the connection can read (the token's own, or
the App installations') and ingests one with the connection's token. The
tool is GitHub's read-only MCP server (api.githubcopilot.com/mcp/readonly):
setup discovers its actions and creates it bound to the connection, and
the executor sends the connection's token only to that server.
GITHUB_ACCESS_TOKEN belongs to the server, but any user could ingest any
repository it can see, private ones included. It is now used only for
public repositories; the loader checks the repository's visibility first
and asks for a GitHub connection otherwise.
The loader also takes a token from the connection a source syncs from.
Merging a connection's keys into a plain repository URL no longer fails
on json.loads, manual Sync now passes the source's connection, and a
token GitHub rejects pauses the connection's sources for reconnect.
- An MCP connection's tokens only go to its own server, and a
connection_id can no longer come from a client: the MCP test/save
routes drop it and the tool executor uses only the resolved one.
- API-key connections are reused only when they hold the same
credentials. Two keys that share a hint or a label get separate
connections ("…abcd (2)"), at runtime and in the 0038 backfill.
- Reconnecting a connection whose secrets cannot be decrypted no longer
flags it from a second transaction while holding its row lock.
- Connection setup validates the sync request before claiming its
Idempotency-Key, so a corrected retry is queued.
- A request that names no connection resolves to none without opening
a database connection; fixes two tests that reached a real database.
Migration 0038 moves every stored secret (OAuth tokens, MCP OAuth
tokens and client registrations, API keys) into the connection's
encrypted envelope, links API-key tools to one connection per distinct
credential, allows several accounts per provider, and adds
credential_mode to sources and tools. OAuth MCP tools keep resolving
each member's own token, as they did before.
docsgpt.connectors.service is now the only reader of OAuth tokens:
get_valid_token_info refreshes under a row lock and persists rotated
refresh tokens, and a revoked grant flags the connection, pauses its
sources and notifies the owner. Loaders build from a connection
(BaseConnectorLoader.from_connection), so scheduled sync covers Drive,
SharePoint and Confluence sources with no browser. S3 and Reddit keys
stay on the connection instead of in remote_data.
New endpoints: POST /api/connections, /setup, /reconnect,
/picker-token, /claim, DELETE /api/connections/<id>, per-action
permissions and MCP refresh-tools. Upload, file listing, sync and
validate-session take a connection_id; session tokens keep working for
this release. The tool executor reads credentials from the resolved
connection (owner or member mode) and pauses on a Connect card when a
connection needs signing in. docsgpt connectors reencrypt rewrites
stored credentials after a key rotation.
About 85 call sites read a setting as getattr(settings, "NAME", fallback),
each carrying its own copy of the default. Every one of those names is a
field with a default on the model, so the fallback could never apply to
the real settings object; it only masked drift. Two had drifted:
- OPENAI_PROMPT_CACHE_KEY defaults to True on the model but the reader
fell back to False, and two test stubs relied on that.
- SharePoint's MICROSOFT_AUTHORITY fallback to
https://login.microsoftonline.com/<tenant> never fired, because the
attribute always exists (as None), so MSAL got authority=None. The
connector now derives the tenant authority when the setting is unset,
as its test always assumed.
Four places read EMBEDDINGS_KEY straight from os.environ, skipping the
"None"/"" normalisation the model applies; they read the setting now.
Test stubs that replaced a module's settings with a SimpleNamespace list
every setting the code under test reads.
- Ship tiktoken's cl100k_base inside the package and build the encoding
from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
temp dir, and read tokenizer.json and repo metadata from that cache, so
a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
the endpoints return 404, audio files fail to ingest with a clear
message, /api/config reports tts_available/stt_available, and the UI
hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
A deflate-compressed TIFF under 1 MB can declare 144 million pixels and
take 1.2 GB to convert to PNG, and Pillow only warns below 179 million.
Read the dimensions from the header and refuse images over 40 million
pixels before any pixel data is decoded. Pillow's DecompressionBombError
is now raised as DocumentParseError, so the upload fails once instead of
being retried.
Major bumps whose ceilings had to move. redis 8.1.0, tiktoken 0.14.0 and
daytona 0.211.2 needed no code change (the Daytona client, filesystem and
process signatures the sandbox calls are unchanged; redis 8 was checked
against a live server through the app's own sync and async clients).
reportlab 5.0.1 is test-only.
openapi-parser 2.0.0 is a rewrite onto pydantic spec models: `paths` is now
a dict keyed by URL rather than a list of objects carrying their own `url`,
and a path item exposes one field per HTTP method instead of an `operations`
list. `OpenAPI3Parser` reads both accordingly, iterating methods in the
order the spec declares them, and its rendered output is byte-identical to
before. The rewrite also drops prance, openapi-spec-validator and five more
transitive packages.
tokenizers stays at 0.22.2: transformers 5.8.1 caps it at <=0.23.0 and no
such release exists, so it moves with the transformers cap or not at all.
`uv lock --upgrade` plus the two code changes the new versions need.
firecrawl-anydoc 0.2.4 raises a dedicated `NeedsOcrError` where 0.2.3 raised
`UnsupportedError("... OCR is required")`, so the anydoc parser no longer
recognised a scanned PDF: the fallback still ran, but a near-empty result was
stored as an empty document instead of failing with the OCR_ENABLED hint.
`_needs_ocr` now accepts both spellings and looks the class up lazily, so an
older anydoc keeps working. 0.2.4 also refuses the CID-font NDA fixture
outright rather than dropping its Chinese column silently, so the PDF
trust-check tests stub that dropped output against the fixture's real bytes
(the check's own inputs) and a new test pins the refusal path.
ruff 0.16 widened its implicit default rule set, turning the dev-group bump
into 7131 findings across the tree. `.ruff.toml` now states the historical
selection (E4, E7, E9, F) explicitly and the CI pin moves to the locked
0.16.7, so lint no longer drifts with the version.
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.
Kept for one release:
- A top-level application package whose meta-path finder resolves
application.x.y to the already-imported docsgpt.x.y object, so old imports
and entry points (celery -A application.app.celery,
uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
docsgpt.* task on start-up, so messages queued by the previous release still
run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
the previous release wrote are left unread instead of firing twice.
The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
- The frontend image ran the Vite dev server in development mode, so
.env.development supplied its defaults (notification banner, Google client
id, local API host). The static build only loads .env.production, so the
build stage now copies .env.development in as the baseline and the compose
files pass every VITE_* the app reads through from .env; the runtime script
skips empty values so a blank passthrough keeps the build-time default.
.dockerignore kept only the .local variants out.
- VITE_DISABLE_SOURCE_FE disables sources only when it is the string true.
- DoclingParser: find_spec raises when docling itself is absent; the install
hint now covers that path, with a regression test.
- verify_offline: direct tests for verify(); the PR image check builds and
verifies the -docling variant as well as slim.
- Workflows this branch adds or rewrites pin actions by commit, pass the
release tag through env instead of template expansion, and do not persist
checkout credentials.
- OCR guide no longer claims pre-built images never include docling.
Three things still reached the network from a container whose models were
baked in:
- tiktoken fetched cl100k_base from openaipublic.blob.core.windows.net on
every fresh container (its cache defaulted to /tmp), and token accounting
calls it on every chat. prefetch_models now warms it too; the image sets
TIKTOKEN_CACHE_DIR.
- The chunker loaded its tokenizer with Tokenizer.from_pretrained, which
revalidates the revision with a HEAD request per process start and stalls
for the etag timeout (10 s) when huggingface.co is unreachable. It now reads
tokenizer.json from the hub cache first and only downloads on a miss; the
repo-metadata read for models outside the registry does the same.
- tldextract fetched the public suffix list on the first web crawl; the
bundled snapshot is used instead.
application/scripts/verify_offline.py exercises these paths (and docling's
conversion when the extra is installed) so an image can be checked with
docker run --network none.
Conflicts, and how each was taken:
- application/core/settings.py — ours. The renamed OCR_ENABLED /
OCR_ATTACHMENTS_ENABLED / OCR_MIN_CHARS_PER_PAGE accept main's
DOCLING_OCR_* spellings as AliasChoices, so nothing is dropped.
- application/Dockerfile — both. Main's install layers plus the
INSTALL_DOCLING build arg.
- application/parser/file/constants.py — both imports.
- deployment/docker-compose.yaml — both. The INSTALL_DOCLING /
INSTALL_TESSERACT build args on backend and worker, and main's
-Q docsgpt,parsing,embeddings, which query embedding needs.
- tests/conftest.py — theirs. Both sides fixed the same pytest-postgresql
9.0.0 autocommit= breakage; main's spelling is the one already on main.
- application/requirements.txt — the comments claimed different reasons
torch is in core. Main's is the true one now: it removed
sentence-transformers, so docling is torch's only remaining consumer.
Two things the merge broke without conflicting:
- onnxruntime. This branch moved it out of core into the docling extra;
main meanwhile made it the runtime local embeddings execute on
(fastembed). Git took the deletion, leaving fastembed with no pinned
runtime in a repo that pins everything. Restored to core, and no longer
pinned twice from the extra.
- The frontend copy of ATTACHMENT_PARSER_EXTENSIONS. The backend list is
derived and picked up the anydoc suffixes; the hand-kept frontend mirror
did not, so the composer would refuse files the API accepts.
tests/parser/file/test_constants.py is what caught it.
ruff, pytest (9897 passed), frontend build and docs build all pass. The
image build is unverified: no Docker daemon on this machine.
Review follow-ups on the attachment gate.
.txt was listed as parser-backed, but it has no parser — it *is* the
plain-text fallthrough. That let it skip the content check, so renaming a
video to notes.txt walked straight back into the bug the gate exists for
(verified: 5132 chars of binary "extracted" and stored). The list is now
exactly the file extractor's keys, .txt included in the content check like
any other unparsed suffix, and the drift test asserts equality rather than
containment. The sniff now recognises a UTF-16/32 BOM as text, so a
Notepad "Unicode" .txt is not caught by the NUL-byte rule.
The picker's accept filter listed parser-backed suffixes only, hiding .txt,
.py and .log — files the gate reads happily — behind "All files". It now
carries text/* as well, so it can never be narrower than what the upload
accepts.
A rejected batch carries one errors entry per file, but the non-200 branch
applied the top-level message to every chip, so two files failing for two
reasons both reported the first one. Reasons are now matched by
upload_index, with the top-level message as fallback.
_get_store_attachment_user_error no longer reads str(exc): the
unsupported-type message is rebuilt from the filename, so no exception
state can reach a response body (CodeQL py/stack-trace-exposure).
A chat attachment with no parser fell through to SimpleDirectoryReader's
plain-text open(), so a phone-uploaded video was "extracted" into megabytes
of binary garbage, truncated, and stored with extraction.status == "ok".
Gate attachments in two tiers instead. A suffix with a dedicated parser is
admitted on its name — a PDF is binary and parses fine. Anything else has to
read as text: the first 8KB are sampled and refused on a NUL byte or too many
other control bytes. That keeps source, config and log files working through
the plain-text fallthrough, and keeps out videos, archives and renamed
binaries alike. The route checks the staged spool before anything is stored
or queued; the worker repeats the check where the local file exists, raising
the non-retryable AttachmentRejectedError.
SUPPORTED_ATTACHMENT_EXTENSIONS gains the parser-backed suffixes it was
missing (.tiff, .tif, .bmp, .webp, .vtt, .xml) and is now exactly the file
extractor's keys plus .txt, with a test asserting the two agree. The composer
applies the same rule client-side, so an unsupported file is named before it
costs an upload, and a test pins the frontend list to the backend one.
Attachment failures now show their reason inline under the chips rather than
only in a hover tooltip, which a touch user can never see, and only after a
send was attempted. Dropped `accept` from the dropzone: it discarded rejected
drops with no feedback and disagreed with the server about text files.
Follow-up review pass over the embeddings branch.
- Fold an oversized header back into the body, and drop header duplication
when it would leave under a quarter of the chunk budget. A header at or
over max_tokens collapsed the body budget to one token, so a document
became one chunk per body token, each still over the cap: a 95 KB file
produced 20k chunks of 2563 tokens against a 1250 cap. Also clamp
max_tokens to at least 1, as the strategy chunkers already do.
- Emit a header-only document as its own chunk. With no body piece to
attach it to, splitting returned nothing and the document was dropped
from the index with no error and no log line.
- Skip add_custom_model for a repository FastEmbed already ships. It
rejects a name it knows, so configuring any of its ~30 built-ins
(MiniLM, bge, e5, gte, ...) failed every embed call and every query.
- Decide "the user chose this model" by comparing against the field
default rather than model_fields_set, which is true for anything read
from .env. Every setup script has always written EMBEDDINGS_NAME, so an
upgraded remote-embeddings install inherited mpnet's 384-token window
and silently clipped ~80% off every chunk.
- Cut tiktoken splits at character offsets instead of decoding each token
window. A multi-byte character straddling a boundary decoded to U+FFFD
on both sides, destroying one character at roughly one boundary in five
on CJK text -- including at the default max_tokens of 2000.
- Let the re-embed script open a FAISS index whose width does not match
the configured model. That mismatch is the main reason to run it, and
the error recommending the script was raised by the script itself, so
the advice failed on every source.
- Re-embed graph_nodes.name_embedding when GraphRAG is enabled. Those
vectors seed every traversal and share the chunk vectors' width, so a
same-width model swap left the graph retrieving from the old space with
nothing to report it.
- Prefetch the models before copying the application source, so editing
any file no longer re-downloads ~780 MB of artifacts on every build.
- Mirror the setup.sh embedding menu into setup.ps1: granite default,
legacy mpnet as an explicit option, and both engine flows updated.
Windows users were otherwise stranded on mpnet with no granite path.
- Drop the unused EmbeddingsWrapper.tokenizer property.
Follow-up to the embeddings work, from a review pass over the branch.
- Route the OpenAI/Azure key handling through the model registry instead of
matching the canonical name literally, so the `text-embedding-ada-002`
alias the registry now accepts also reaches the Azure deployment name
rather than failing every embed with DeploymentNotFound.
- Fall back to a default width where the embeddings model reports no
dimension. A model outside the registry returns None rather than no
attribute, so `getattr` with a default did not catch it and the width
reached the DDL as `vector(None)` / `list_size=None`.
- Point HF_HUB_CACHE at the prefetch directory. Chunking loads the tokenizer
through `tokenizers`, which reads the hub cache, so a fresh container
fetched over the network on first ingest and an offline one silently fell
back to cl100k.
- Charge a token that collapses a long unbroken run by its character span.
WordPiece emits one [UNK] for any word over its character limit, which made
base64 and minified content count as near-zero tokens, so nothing split it
and oversized chunks reached the embedding server.
- Preserve chunk ids and honour --batch-size when rebuilding a FAISS index.
Fresh uuids orphaned GraphRAG's graph_node_chunks rows, and the whole index
went out in a single embed call on remote servers.
- Document that granite runs an int8-quantised graph, and scope the
SentenceTransformer parity claim to mpnet's fp32 graph, which is where it
was measured.
- Correct the embeddings docs: a matching dimension is not a matching model,
so a same-width swap raises nothing and silently degrades retrieval.
Replace the sandbox Docling extractor with read_document, backed by the in-process
backend parser (the same one ingestion uses) and offloaded to a dedicated
'parsing' Celery queue so it can run on GPU-capable workers with predictable RAM.
The tool resolves the input ref under the run-scoped gate, enqueues the parse,
and awaits it with a timeout (degrading to an error rather than hanging); the
worker independently re-resolves the artifact through the same gate and never
trusts a raw path. Untrusted files get the upload path's safeguards (extension
whitelist, size cap, sanitized temp file, cleanup). Options: output
(markdown/text/structured/chunks), ocr, pages, engine, max_chars, include_tables,
persist, json_schema. The workflow native-file 'extract' fallback now uses the
same worker path, so document parsing no longer needs the sandbox and works on
every backend.
Also fixes the branch's periodic-task test (the sandbox reaper made it 12) and
points the dev and e2e Celery workers at the parsing queue.