27 Commits
Author SHA1 Message Date
arc53-machine 26bd4e2a3f Keep the GitHub loader tests off the network
With GITHUB_ACCESS_TOKEN set in the environment, load_data checked the
repository's visibility against the real GitHub API. The tests now clear
the instance token unless they set one.
2026-09-29 15:09:50 +01:00
arc53-machine f2349edb4c Keep a Linear sync going when one issue or document cannot be read
An issue deleted after it was listed, or comments the token cannot read,
now lose that detail with a warning instead of failing the whole sync; an
unreadable document is skipped.
2026-09-29 12:14:10 +01:00
arc53-machine 03d1b10897 Sync Linear issues and documents into Knowledge with the Linear sign-in
The Linear connector now syncs as well as giving agents its tools, from
one connection. Linear's MCP server is its own OAuth issuer, so its
tokens are read through the same MCP tools the agents use (list_issues,
get_issue, list_comments, list_documents) rather than Linear's GraphQL
API, and no OAuth app has to be registered.

A source picks teams and projects, with comments (on by default) and the
projects' documents. Each issue becomes one document with its state,
assignee, priority, labels, description and comments, filed under its
team and citing its Linear URL. Each sync reads up to 500 issues and 100
documents again. /api/connections/<id>/linear lists the teams and
projects to pick from. Sources sync on their schedule with the owner's
connection, and pause when the sign-in needs reconnecting.
2026-09-29 11:32:40 +01:00
arc53-machine 3572d968ec Read private repositories only with the user's own GitHub token
GITHUB_ACCESS_TOKEN belongs to the server, but any user could ingest any
repository it can see, private ones included. It is now used only for
public repositories; the loader checks the repository's visibility first
and asks for a GitHub connection otherwise.

The loader also takes a token from the connection a source syncs from.
Merging a connection's keys into a plain repository URL no longer fails
on json.loads, manual Sync now passes the source's connection, and a
token GitHub rejects pauses the connection's sources for reconnect.
2026-09-29 10:41:03 +01:00
arc53-machine 3e11442a83 Merge remote-tracking branch 'origin/main' into connectors
# Conflicts:
#	frontend/DESIGN.md
#	frontend/src/locale/de.json
#	frontend/src/locale/en.json
#	frontend/src/locale/es.json
#	frontend/src/locale/jp.json
#	frontend/src/locale/ru.json
#	frontend/src/locale/zh-TW.json
#	frontend/src/locale/zh.json
#	frontend/src/upload/Upload.tsx
2026-09-28 18:07:12 +01:00
arc53-machine 3511004f50 Connections own their credentials
Migration 0038 moves every stored secret (OAuth tokens, MCP OAuth
tokens and client registrations, API keys) into the connection's
encrypted envelope, links API-key tools to one connection per distinct
credential, allows several accounts per provider, and adds
credential_mode to sources and tools. OAuth MCP tools keep resolving
each member's own token, as they did before.

docsgpt.connectors.service is now the only reader of OAuth tokens:
get_valid_token_info refreshes under a row lock and persists rotated
refresh tokens, and a revoked grant flags the connection, pauses its
sources and notifies the owner. Loaders build from a connection
(BaseConnectorLoader.from_connection), so scheduled sync covers Drive,
SharePoint and Confluence sources with no browser. S3 and Reddit keys
stay on the connection instead of in remote_data.

New endpoints: POST /api/connections, /setup, /reconnect,
/picker-token, /claim, DELETE /api/connections/<id>, per-action
permissions and MCP refresh-tools. Upload, file listing, sync and
validate-session take a connection_id; session tokens keep working for
this release. The tool executor reads credentials from the resolved
connection (owner or member mode) and pauses on a Connect card when a
connection needs signing in. docsgpt connectors reencrypt rewrites
stored credentials after a key rotation.
2026-09-28 17:03:09 +01:00
Pavel f4331cd3a3 rabbit fixes 2026-09-28 18:40:49 +04:00
Pavel 706a0cb2b2 Big source revamp 2026-09-28 17:35:16 +04:00
Alex 574f96341e refactor: rename the application package to docsgpt
The backend import package is now docsgpt, the name it will carry on PyPI;
application was far too generic to install into anyone's site-packages.
git mv plus a mechanical rewrite of every import, dotted string and path
reference: 734 Python files, the compose files, Dockerfile, workflows, docs,
setup scripts, devcontainer, k8s manifests, vscode config, pytest and coverage
config, .gitignore. Behaviour is unchanged.

Kept for one release:
- A top-level application package whose meta-path finder resolves
  application.x.y to the already-imported docsgpt.x.y object, so old imports
  and entry points (celery -A application.app.celery,
  uvicorn application.asgi:asgi_app) keep working with a FutureWarning.
- Celery registers every application.* task name as an alias of its
  docsgpt.* task on start-up, so messages queued by the previous release still
  run. The redbeat key prefix moves to redbeat:docsgpt:v2: so schedule entries
  the previous release wrote are left unread instead of firing twice.

The backend image builds from the repository root (docker build -f
docsgpt/Dockerfile .) so it can ship the alias package; a root .dockerignore
allow-lists docsgpt/ and application/ and keeps caches, local data, .env
files, the sample index files and the Dockerfile out. Compose and the image
workflows point at the new context.
2026-09-07 10:20:43 +01:00
Alex 4707c45b93 fix(parser,vectorstore): stop the first-request downloads in a warmed install
Three things still reached the network from a container whose models were
baked in:

- tiktoken fetched cl100k_base from openaipublic.blob.core.windows.net on
  every fresh container (its cache defaulted to /tmp), and token accounting
  calls it on every chat. prefetch_models now warms it too; the image sets
  TIKTOKEN_CACHE_DIR.
- The chunker loaded its tokenizer with Tokenizer.from_pretrained, which
  revalidates the revision with a HEAD request per process start and stalls
  for the etag timeout (10 s) when huggingface.co is unreachable. It now reads
  tokenizer.json from the hub cache first and only downloads on a miss; the
  repo-metadata read for models outside the registry does the same.
- tldextract fetched the public suffix list on the first web crawl; the
  bundled snapshot is used instead.

application/scripts/verify_offline.py exercises these paths (and docling's
conversion when the extra is installed) so an image can be checked with
docker run --network none.
2026-09-05 15:50:20 +01:00
Pavel ed0892b39b Batch fixes 2 2026-09-03 00:30:59 +04:00
Alex e0aff39a1b feat: ingestion optimisations 2026-08-25 12:27:09 +01:00
Alex a166796f6c feat: prune some deps 2026-08-10 12:10:54 +01:00
PavelandAlex aaad51f951 Harden protection with pinned requests and path-param encoding (#2486)
* Harden protection with pinned requests and path-param encoding

* fix: domain pinning

* fix: tests

* fix: html test 2

---------

Co-authored-by: Alex <a@tushynski.me>
2026-05-23 02:33:29 +01:00
Alex e167cf8247 fix: broken syncs (#2480)
* fix: broken syncs

* fix: mini fixes
2026-05-17 23:58:28 +01:00
81b6ee5daa Pg 4 (#2390)
* feat: postgres tests

* feat: mongo cutoff

* feat: mongo cutoff

* feat: adjust docs and compose files

* fix: mini code mongo removals

* fix: tests and k8s mongo stuff

* feat: test fixes

* fix: ruff

* fix: vale

* Potential fix for pull request finding 'CodeQL / Clear-text logging of sensitive information'

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>

* fix: mini suggestions

* vale lint fix 2

* fix: codeql columns thing

* fix: test mongo

* fix: tests coverage

* feat: better tests 4

* feat: more tests

* feat: decent coverage

* fix: ruff fixes

* fix: remove mongo mock

* feat: enhance workflow engine and API routes; add document retrieval and source handling

* feat: e2e tests

* fix: mcp, mongo and more

* fix: mini codeql warning

* fix: agent chunk view

* fix: mini issues

* fix: more pg fixes

* feat: postgres prep on start

* feat: qa tests

* fix: mini improvements

* fix: tests

---------

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
Co-authored-by: Siddhant Rai <siddhant.rai.5686@gmail.com>
2026-04-18 13:13:57 +01:00
Alex b5b6538762 fix: tests 2026-04-12 12:35:23 +01:00
Alex a3b08a5b44 More tests 2026-03-31 00:07:19 +01:00
Alex d5c0322e2a chore: more tests 2026-03-30 16:13:08 +01:00
918bbf0369 Sharepoint (#2283)
* feat: add Microsoft Entra ID integration

- Updated .env-template and settings.py for Microsoft Entra ID configuration.
- Enhanced ConnectorsCallback to support SharePoint authentication.
- Introduced SharePointAuth and SharePointLoader classes.
- Added required dependencies in requirements.txt.

* feat: agent templates and seeding premade agents (#1910)

* feat: agent templates and seeding premade agents

* fix: ensure ObjectId is used for source reference in agent configuration

* fix: improve source handling in DatabaseSeeder and update tool config processing

* feat: add prompt handling in DatabaseSeeder for agent configuration

* Docs premade agents

* link to prescraped docs

* feat: add template agent retrieval and adopt agent functionality

* feat: simplify agent descriptions in premade_agents.yaml  added docs

---------

Co-authored-by: Pavel <pabin@yandex.ru>
Co-authored-by: Alex <a@tushynski.me>

* feat: add GitHub access token support and fix file content fetching logic (#2032)

* feat: add init for Share Point connector module

* chore(deps): bump mermaid from 11.6.0 to 11.12.0 in /frontend

Bumps [mermaid](https://github.com/mermaid-js/mermaid) from 11.6.0 to 11.12.0.
- [Release notes](https://github.com/mermaid-js/mermaid/releases)
- [Commits](https://github.com/mermaid-js/mermaid/compare/mermaid@11.6.0...mermaid@11.12.0)

---
updated-dependencies:
- dependency-name: mermaid
  dependency-version: 11.12.0
  dependency-type: direct:production
  update-type: version-update:semver-minor
...

Signed-off-by: dependabot[bot] <support@github.com>

* Feat: Notification section (#2033)

* Feature/Notification-section

* fix notification ui and add local storage variable to save the state

* add notification component to app.tsx

* refactor: remove MICROSOFT_REDIRECT_URI and update SharePointAuth to use CONNECTOR_REDIRECT_BASE_URI

* feat: Add button to cancel LLM response (#1978)

* feat: Add button to cancel LLM response
- Replace text area with cancel button when loading.
- Add useEffect to change elipsis in cancel button text.
- Add new SVG icon for cancel response.
- Button colors match Figma designs.

* fix: Cancel button UI matches new design
- Delete cancel-response svg.
- Change previous cancel button to match the new Figma design.
- Remove console log in handleCancel function.

* fix: Adjust cancel button rounding

* feat: Update UI for send button
- Add SendArrowIcon component, enables dynamic svg color changes
- Replace original icon
- Update colors and hover effects

* (fix:send-button) minor blink in transition

---------

Co-authored-by: Manish Madan <manishmadan321@gmail.com>

* feat: add SharePoint integration with session validation and UI components

* (feat:oneDrive) file loading for ingestion

* feat(oneDrive): shared user files

* (feat:oneDrive) rm shared file support, as sharedWithMe is degraded

* (feat:sharepoint) shared files for work msa

* (feat:sharepoint) retry on auth failure, decorator

* (fix) tests/ruff

* test: fix sharepoint loader expecting client id

---------

Signed-off-by: dependabot[bot] <support@github.com>
Co-authored-by: Abhishek Malviya <abfeb8@gmail.com>
Co-authored-by: Siddhant Rai <47355538+siiddhantt@users.noreply.github.com>
Co-authored-by: Pavel <pabin@yandex.ru>
Co-authored-by: Alex <a@tushynski.me>
Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Co-authored-by: Mariam Saeed <69825646+Mariam-Saeed@users.noreply.github.com>
Co-authored-by: Rahul <rahulgithub96@gmail.com>
2026-03-12 14:46:26 +00:00
Alex df57053613 feat: improve crawlers and update chunk filtering (#2250) 2026-01-06 00:52:12 +02:00
Alex 9e7f1ad1c0 Add Amazon S3 support and synchronization features (#2244)
* Add Amazon S3 support and synchronization features

* refactor: remove unused variable in load_data test
2025-12-30 20:26:51 +02:00
Alex 98e949d2fd Patches (#2218)
* feat: implement URL validation to prevent SSRF

* feat: add zip extraction security

* ruff fixes
2025-12-24 17:05:35 +02:00
Alex e0a9f08632 refactor and deps (#2184) 2025-12-10 23:53:59 +02:00
Alex 03452ffd9f feat: add GitHub access token support and fix file content fetching logic (#2032) 2025-10-07 16:53:14 +03:00
ManishMadan2882 e24a0ac686 (test:parsers) github, reddit 2025-09-29 20:33:05 +05:30
ManishMadan2882 8c91b1c527 (tests:parsers) remote 2025-09-29 19:39:24 +05:30