15 Commits
Author SHA1 Message Date
arc53-machine f2349edb4c Keep a Linear sync going when one issue or document cannot be read
An issue deleted after it was listed, or comments the token cannot read,
now lose that detail with a warning instead of failing the whole sync; an
unreadable document is skipped.
2026-09-29 12:14:10 +01:00
arc53-machine 03d1b10897 Sync Linear issues and documents into Knowledge with the Linear sign-in
The Linear connector now syncs as well as giving agents its tools, from
one connection. Linear's MCP server is its own OAuth issuer, so its
tokens are read through the same MCP tools the agents use (list_issues,
get_issue, list_comments, list_documents) rather than Linear's GraphQL
API, and no OAuth app has to be registered.

A source picks teams and projects, with comments (on by default) and the
projects' documents. Each issue becomes one document with its state,
assignee, priority, labels, description and comments, filed under its
team and citing its Linear URL. Each sync reads up to 500 issues and 100
documents again. /api/connections/<id>/linear lists the teams and
projects to pick from. Sources sync on their schedule with the owner's
connection, and pause when the sign-in needs reconnecting.
2026-09-29 11:32:40 +01:00
arc53-machine 7b516ade82 Add a built-in GitHub connector for repository sync and read-only tools
One GitHub connection feeds both a Knowledge source and an agent tool.
Users connect with a personal access token, checked against GitHub and
named after the account, or, when an admin registers a GitHub App
(GITHUB_CLIENT_ID, GITHUB_CLIENT_SECRET, GITHUB_APP_SLUG), with Sign in
with GitHub. App tokens expire after eight hours and are refreshed before
each sync or tool call; tokens with no expiry are never treated as expired.

The catalog gains an optional second sign-in method (oauth_settings,
exposed as sign_in_methods) without changing any other connector.

Sync lists the repositories the connection can read (the token's own, or
the App installations') and ingests one with the connection's token. The
tool is GitHub's read-only MCP server (api.githubcopilot.com/mcp/readonly):
setup discovers its actions and creates it bound to the connection, and
the executor sends the connection's token only to that server.
2026-09-29 10:41:03 +01:00
arc53-machine 3572d968ec Read private repositories only with the user's own GitHub token
GITHUB_ACCESS_TOKEN belongs to the server, but any user could ingest any
repository it can see, private ones included. It is now used only for
public repositories; the loader checks the repository's visibility first
and asks for a GitHub connection otherwise.

The loader also takes a token from the connection a source syncs from.
Merging a connection's keys into a plain repository URL no longer fails
on json.loads, manual Sync now passes the source's connection, and a
token GitHub rejects pauses the connection's sources for reconnect.
2026-09-29 10:41:03 +01:00
arc53-machine 3e11442a83 Merge remote-tracking branch 'origin/main' into connectors
# Conflicts:
#	frontend/DESIGN.md
#	frontend/src/locale/de.json
#	frontend/src/locale/en.json
#	frontend/src/locale/es.json
#	frontend/src/locale/jp.json
#	frontend/src/locale/ru.json
#	frontend/src/locale/zh-TW.json
#	frontend/src/locale/zh.json
#	frontend/src/upload/Upload.tsx
2026-09-28 18:07:12 +01:00
arc53-machine 3511004f50 Connections own their credentials
Migration 0038 moves every stored secret (OAuth tokens, MCP OAuth
tokens and client registrations, API keys) into the connection's
encrypted envelope, links API-key tools to one connection per distinct
credential, allows several accounts per provider, and adds
credential_mode to sources and tools. OAuth MCP tools keep resolving
each member's own token, as they did before.

docsgpt.connectors.service is now the only reader of OAuth tokens:
get_valid_token_info refreshes under a row lock and persists rotated
refresh tokens, and a revoked grant flags the connection, pauses its
sources and notifies the owner. Loaders build from a connection
(BaseConnectorLoader.from_connection), so scheduled sync covers Drive,
SharePoint and Confluence sources with no browser. S3 and Reddit keys
stay on the connection instead of in remote_data.

New endpoints: POST /api/connections, /setup, /reconnect,
/picker-token, /claim, DELETE /api/connections/<id>, per-action
permissions and MCP refresh-tools. Upload, file listing, sync and
validate-session take a connection_id; session tokens keep working for
this release. The tool executor reads credentials from the resolved
connection (owner or member mode) and pauses on a Connect card when a
connection needs signing in. docsgpt connectors reencrypt rewrites
stored credentials after a key rotation.
2026-09-28 17:03:09 +01:00
Pavel f4331cd3a3 rabbit fixes 2026-09-28 18:40:49 +04:00
Pavel 706a0cb2b2 Big source revamp 2026-09-28 17:35:16 +04:00
arc53-machine f882ef49a7 refactor: read settings directly instead of getattr with a second default
About 85 call sites read a setting as getattr(settings, "NAME", fallback),
each carrying its own copy of the default. Every one of those names is a
field with a default on the model, so the fallback could never apply to
the real settings object; it only masked drift. Two had drifted:

- OPENAI_PROMPT_CACHE_KEY defaults to True on the model but the reader
  fell back to False, and two test stubs relied on that.
- SharePoint's MICROSOFT_AUTHORITY fallback to
  https://login.microsoftonline.com/<tenant> never fired, because the
  attribute always exists (as None), so MSAL got authority=None. The
  connector now derives the tenant authority when the setting is unset,
  as its test always assumed.

Four places read EMBEDDINGS_KEY straight from os.environ, skipping the
"None"/"" normalisation the model applies; they read the setting now.
Test stubs that replaced a module's settings with a SimpleNamespace list
every setting the code under test reads.
2026-09-17 11:14:34 +01:00
Alex 7da46c2bea feat: air-gapped deployment guide, no implicit downloads
- Ship tiktoken's cl100k_base inside the package and build the encoding
  from it, so token counting never downloads anything.
- Default EMBEDDINGS_CACHE_DIR to <data home>/models instead of FastEmbed's
  temp dir, and read tokenizer.json and repo metadata from that cache, so
  a model downloads once and survives reboots.
- TTS_PROVIDER=none and STT_PROVIDER=none switch the speech features off:
  the endpoints return 404, audio files fail to ingest with a clear
  message, /api/config reports tts_available/stt_available, and the UI
  hides the Speak and microphone buttons.
- Drop the Google Fonts Roboto import from the web UI.
- prefetch-models fills the cache the app reads; verify-offline checks the
  packaged encoding.
- Docs: new Air-Gapped Deployment guide, settings and cache notes.
2026-09-15 17:54:24 +01:00
Alex ad201f8318 fix: refuse oversized TIFF/BMP attachments before converting them
A deflate-compressed TIFF under 1 MB can declare 144 million pixels and
take 1.2 GB to convert to PNG, and Pillow only warns below 179 million.
Read the dimensions from the header and refuse images over 40 million
pixels before any pixel data is decoded. Pillow's DecompressionBombError
is now raised as DocumentParseError, so the upload fails once instead of
being retried.
2026-09-14 17:45:43 +01:00
Alex 3030aba39d fix: handle non text uploads more carefully 2026-09-14 17:27:13 +01:00
Alex b3326b3c33 chore(deps): redis 8, tiktoken 0.14, openapi3-parser 2, daytona 0.211, reportlab 5
Major bumps whose ceilings had to move. redis 8.1.0, tiktoken 0.14.0 and
daytona 0.211.2 needed no code change (the Daytona client, filesystem and
process signatures the sandbox calls are unchanged; redis 8 was checked
against a live server through the app's own sync and async clients).
reportlab 5.0.1 is test-only.

openapi-parser 2.0.0 is a rewrite onto pydantic spec models: `paths` is now
a dict keyed by URL rather than a list of objects carrying their own `url`,
and a path item exposes one field per HTTP method instead of an `operations`
list. `OpenAPI3Parser` reads both accordingly, iterating methods in the
order the spec declares them, and its rendered output is byte-identical to
before. The rewrite also drops prance, openapi-spec-validator and five more
transitive packages.

tokenizers stays at 0.22.2: transformers 5.8.1 caps it at <=0.23.0 and no
such release exists, so it moves with the transformers cap or not at all.
2026-09-12 17:28:58 +01:00
Alex 7e00197678 chore(deps): upgrade backend deps within their declared ranges
`uv lock --upgrade` plus the two code changes the new versions need.

firecrawl-anydoc 0.2.4 raises a dedicated `NeedsOcrError` where 0.2.3 raised
`UnsupportedError("... OCR is required")`, so the anydoc parser no longer
recognised a scanned PDF: the fallback still ran, but a near-empty result was
stored as an empty document instead of failing with the OCR_ENABLED hint.
`_needs_ocr` now accepts both spellings and looks the class up lazily, so an
older anydoc keeps working. 0.2.4 also refuses the CID-font NDA fixture
outright rather than dropping its Chinese column silently, so the PDF
trust-check tests stub that dropped output against the fixture's real bytes
(the check's own inputs) and a new test pins the refusal path.

ruff 0.16 widened its implicit default rule set, turning the dev-group bump
into 7131 findings across the tree. `.ruff.toml` now states the historical
selection (E4, E7, E9, F) explicitly and the CI pin moves to the locked
0.16.7, so lint no longer drifts with the version.
2026-09-12 17:16:59 +01:00
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00