Commit Graph
504 Commits
Author SHA1 Message Date
Alex f6400cd736 feat: per-source RAG configuration (retrieval strategies, chunking, exposure, prescreen)
Introduces a per-source config contract that makes RAG behavior strategy-dispatched instead of a single hardcoded path. Every source gains a validated JSONB config; an empty/absent config reproduces current behavior byte-for-byte, and the whole path is gated by PER_SOURCE_RETRIEVAL_ENABLED.

Foundation: sources.config JSONB column + migration 0022_source_config; SourceConfig/ChunkingConfig/RetrievalConfig pydantic models (strict on write, lenient on read); ChunkerCreator and RetrieverCreator.register registries; config threaded through the upload routes, ingest/remote/connector workers, and reingest.

Retrieval: a Dispatcher groups sources by retriever key (all-classic collapses to today's single ClassicRAG under one shared token budget; non-classic retrievers get their own instance), removing the previous single-global-retriever collapse in stream_processor. Per-source chunks, score_threshold (honored for pgvector/mongodb, safely ignored elsewhere), and rephrase_query toggle. New PATCH /api/sources/<id>/config with team-aware (effective_write_owner) authz and a requires_reingest signal.

Chunking strategies: recursive, markdown, parent_child (selectable per source; re-ingest to apply). Search exposure: per-source prefetch vs agentic_tool for agentic/research agents. Map-reduce prescreen: optional LLM relevance pre-filter implemented as a composable post-retrieval stage that wraps any retriever.

Backend and frontend (shared Retrieval options panel + edit modal) with tests; backend suite and frontend vitest green. Excludes the wiki and GraphRAG flagships.
2026-06-20 21:54:23 +01:00
Alex 1b7f84d981 Merge remote-tracking branch 'origin/main' into team 2026-06-16 17:34:27 +01:00
Alex 26d741ab94 feat: notifications on share 2026-06-16 14:50:55 +01:00
Alex d37e6bec51 feat: better tags and dropdowns for sharing 2026-06-16 12:14:28 +01:00
Alex d21f4758be feat: teams 2026-06-16 09:21:21 +01:00
Alex c772ebc60f fix: namespace 2026-06-15 11:44:02 +01:00
Alex 9a349401f4 feat: admin dashboard 2026-06-15 11:30:02 +01:00
Alex 82cd7b7d49 fix: tests 2026-06-14 22:38:28 +01:00
Alex 42678e95cd feat: admin role rbac 2026-06-14 21:36:07 +01:00
Alex ee4abf1038 fix: keep agent published on update 2026-06-14 14:44:27 +01:00
Alex ab3337ced8 fix: stale numbers and slug reassignments 2026-06-14 14:24:18 +01:00
Alex 3e968be6cc Merge origin/main into agent-exports; renumber agent-slug migration to 0019
main added migration 0018_tool_attempts_attribution (revises 0017_oidc_scim), which collided with the feature's 0018_agent_slug. Renumbered the agent-slug migration to 0019 (revises 0018_tool_attempts_attribution) so the Alembic chain stays linear (single head). Auto-merge was conflict-free — models.py, the frontend API layer, and all 7 locale files merged additively.
2026-06-13 14:42:50 +01:00
Alex f8e2fb715e Merge pull request #2534 from arc53/logs-and-anal-revamp
revamp analytics & logs — per-agent attribution, unified log timeline
2026-06-13 14:25:32 +01:00
Alex cd6f40471a fix: sanitise on request 2026-06-13 13:55:13 +01:00
Alex 5f29a384e2 feat: agent import / export 2026-06-13 13:31:46 +01:00
Pavel 9fb59785b2 Second fixes batch 2026-06-13 15:39:16 +04:00
Pavel 59704b0f73 sec fixes 2026-06-12 18:19:15 +04:00
Alex b8eaa5a68d fix: mini tool call issues 2026-06-11 16:49:52 +01:00
Alex 636d4d35d1 feat(prompts): revamp preset prompts, tool naming, and prompt templating
Rewrite the default/creative/strict presets (classic + agentic) into
structured sections: grounding and cite-by-title guidance, insufficient-
context behavior, current date, respond-in-user-language, scoped mermaid
usage, an untrusted-content guardrail, and a conditional XML-tagged
document context block. A memory directory listing is injected at render
time via the template prefetch mechanism so the model starts oriented
without burning a tool call.

Fixes along the way:
- Agentic preset swap was dead code: _get_prompt_content cached the
  classic preset before create_agent's swap check ran, so agentic and
  research agents always got the classic preset. The swap now happens
  inside _get_prompt_content.
- Jinja autoescape corrupted document content in custom prompts
  (< -> &lt;); prompts are not HTML, autoescape is now off.
- Literal {summaries} leaked into the prompt when no docs were
  retrieved; the placeholder is now stripped.
- Agentic/research prompts referenced tool names from a dropped naming
  scheme (search_internal, reason_think); they now reference the real
  names (search, reason).
- The strict preset told the model to "be very creative and use your
  imagination" right after "never make up information".
- extract_tool_usages recorded intermediate attribute chains as
  bare-tool usages, which meant "run all actions" at prefetch; only
  maximal chains are recorded now.
- Headless runs retrieved docs but never rendered them into the
  prompt; the prompt is now rendered like the streaming path.
- Default tools were unreachable by name in prompt templates
  (prefetch results were keyed by synthetic id only); defaults now
  claim the name key unless an explicit row shadows it.

Tool layer: memory/notes/todo actions are namespaced (memory_view,
note_overwrite, todo_create, ...) with legacy unprefixed names still
accepted via prefix stripping; duplicate action names across tools are
disambiguated with the owning tool's name instead of numeric suffixes;
thin tool descriptions rewritten (brave, duckduckgo, telegram, ntfy,
cryptoprice, read_webpage, internal_search, think).

Docs are now wrapped per chunk in <document index>/<source>/<content>
tags for citation-by-title support.
2026-06-11 11:18:46 +01:00
Pavel 64e19f4d11 revamp analytics & logs — per-agent attribution, unified log timeline
- Move analytics endpoints to Postgres with agent filtering that matches both stamps (api_key for external traffic, agent_id for owner/headless)
- Add tool & schedule analytics, token grouping (model/agent/source) and side-channel toggle
- Merge chat/system/webhook/workflow/schedule events into one logs timeline with level/type/search filters
- Stamp user/agent on tool_call_attempts at propose time (migration 0018) and backfill via parent message
- Frontend: revamped Analytics charts and Logs page
2026-06-11 13:35:57 +04:00
Alex 10cc16eab0 Merge pull request #2530 from arc53/oidc-login
feat: oidc login
2026-06-10 17:36:36 +01:00
Alex 321d88d5be fix(oidc): close re-review gaps in the hardening commit
Follow-up to the OIDC security hardening, from a max-effort re-review:

- Refresh now re-checks the denylist immediately before minting, against the
  (possibly remapped) identity but anchored on the original session `iat` — so a
  back-channel logout / SCIM deny that lands during the IdP grant, or one
  targeting the refreshed sub/sid, still blocks renewal instead of being escaped
  by the renewed token's fresh iat. Completes the watermark revocation fix.
- SCIM PUT `userName` immutability check is now case-insensitive, matching the
  case-insensitive list/create — a differently-cased userName echo no longer
  400s "userName is immutable" and blocks deprovision. Completes the SCIM
  case-insensitivity fix.
- Drop the unconditional state-cookie deletion on every callback exit: it let
  one tab's callback clear another in-flight tab's cookie, breaking concurrent
  logins. The cookie self-expires (max_age) and the Redis state is single-use,
  so the delete wasn't needed.

Tests added for the refresh revocation re-check and the case-insensitive PUT.
2026-06-10 14:15:32 +01:00
Alex a3d09e8890 fix(oidc): harden session security from PR review (CSRF, revocation, refresh)
Address the high/medium correctness findings on the OIDC/SCIM PR:

- Login CSRF / session fixation: bind `state` to a Secure/HttpOnly/SameSite=Lax
  cookie at login and require the callback to echo it, so a code+state captured
  from another browser can't silently sign a victim into the attacker's account.
- Require `exp` on session JWTs under AUTH_TYPE=oidc (require_exp), so an
  exp-less HS256 token signed with JWT_SECRET_KEY can't authenticate forever or
  outlive the denylist.
- Denylist now keys revocation on an `iat` watermark instead of a deletable
  flag: a fresh login (newer iat) self-supersedes a revocation without clearing
  it, so sessions revoked on other devices stay revoked. Drops the
  login/SCIM-reactivation denylist-clearing paths (allow_user/allow_idp_sub).
- Refresh: gate the disabled-account check on the post-grant identity (not just
  the old sub); attempt the IdP grant before consuming the refresh token and
  return a retryable 503 (restoring the token) on transient IdP errors instead
  of force-logging-out a live session.
- Gate the oidc blueprint at request time on AUTH_TYPE=oidc, so non-oidc
  deployments cleanly 404 these routes instead of 500-ing on an unset
  OIDC_ISSUER (mirrors SCIM_ENABLED).
- Surface revocation write failures: back-channel logout returns 502, and SCIM
  deactivation rolls back and returns 503, when the denylist write fails — so
  the IdP retries instead of recording a logout/deprovision that didn't revoke.
- Back-channel logout: require `jti`, run the replay check unconditionally, and
  reject stale `iat` beyond the replay-cache window.
- Make migration 0017 idempotent (IF NOT EXISTS) so re-apply can't wedge startup.
- SCIM userName matching is case-insensitive (caseExact=false) for the list
  filter and create-dedup.

Tests added/updated across test_oidc.py, test_scim.py, test_auth.py,
test_app_routes.py and the SCIM integration test.
2026-06-10 13:58:58 +01:00
Alex 160736f075 feat: visibility settings 2026-06-10 01:40:53 +01:00
Alex 054f0f1b8b feat: oidc renewal, groups, logout, scim 2026-06-10 00:57:18 +01:00
Alex 8bc777a428 feat: oidc login 2026-06-09 23:24:44 +01:00
Alex 28d62f4199 feat: better tool call followup on compat mode 2026-06-09 18:20:14 +01:00
Alex 82bacfdb2b feat: conv visibility 2026-06-08 23:46:28 +01:00
Alex 73ed1cf607 feat: async events on asgi 2026-06-08 22:02:21 +01:00
Alex 1f06f31aa3 feat: better shutdowns 2026-06-08 15:07:31 +01:00
Alex 59e9523666 fix: revoke stale UI surfaces when durable state goes terminal
The "Tool approval needed" toast (and several sibling surfaces) could
linger after the state they represent was already gone. User-scoped SSE
events (tool.approval.required, schedule.autopaused, attachment.queued,
…) are durable and replayed on reconnect, but no terminal path emitted a
matching clearing event and the reconciler only wrote operator-facing
stack_logs — so a failed/expired message replayed its approval prompt
with nothing to act on, and the toast trusted event presence over the
actual message state.

Backend — emit a user-facing event on every terminal path:
- reconciler deletes pending_tool_state and publishes
  tool.approval.cleared when a stuck message is failed;
  cleanup_pending_tool_state does the same for TTL-reaped rows
- reconciler now publishes source.ingest.failed (stalled ingest),
  schedule.run.failed (timeout/pending) and schedule.completed (once)
- schedules PATCH-resume / DELETE publish schedule.resumed / .cancelled
  so a stale schedule.autopaused can't outvote them on replay
- store_attachment gains the on_poison terminal-event hook the ingest
  tasks already have; mcp_oauth_task gains a soft/hard time limit so a
  hung flow self-reports mcp.oauth.failed
- tighten the reconciler exemption to the (conversation_id, user_id)
  composite key

Frontend — stop trusting event presence over truth:
- notificationsSlice.resolveToolApproval evicts the matching
  tool.approval.required and persists its (stable) id dismissed so the
  backlog replay stays suppressed
- ToolApprovalToast drops approval events older than the resumable TTL
  window as a backstop for a lost clearing event
- schedulesSlice handles schedule.resumed / .cancelled / .completed
- upload dismissals now outlive the SSE backlog retention window

Tests cover the new clearing events at the reconciler, slice and
dispatch layers.
2026-06-07 13:17:50 +01:00
Alex 37417f99c8 fix(v1): address review feedback (multimodal provider gate, save default, leak scope, continuation sampling)
- Drop multimodal content for non-OpenAI-family providers (Google/Anthropic
  raised on OpenAI image_url parts) so multimodal requests degrade to text
  instead of returning 500.
- Stateless continuations (no conversation_id) default save_conversation to
  False, avoiding an orphan conversation with an empty question on every tool
  round (stateful continuations still default True).
- Only strip leaked reasoning from content for structured requests
  (response_format / json_schema / json_object); legitimate answers that mention
  the marker text are no longer corrupted.
- Forward sampling params (temperature, max_tokens, ...) on continuation turns.
- An explicit json_object request clears an agent-configured json_schema so it
  isn't silently overridden.
- Drop max_tokens when max_completion_tokens is also sent (OpenAI rejects both).
- Don't build a chatcmpl-None completion id from the placeholder "None" id event.
2026-06-04 12:55:47 +01:00
Alex 9eb7878ad3 feat(v1): support multimodal content arrays (text + image_url)
OpenAI-compatible clients send multimodal user turns as a `content` array of
typed parts. translate_request previously assigned the array straight to the
question, breaking the string-only retrieval / token-budgeting / history paths
(HTTP 500). Now:

- content_to_text() extracts text from content arrays for the question,
  history and system prompt, so the string paths work unchanged.
- The full content array (text + image_url parts) is preserved as
  `multimodal_content`, threaded to the agent and emitted as the final user
  message so images reach the model. Token budgeting uses the text only.

The content array (incl. image_url) now reaches the LLM call intact; images
render for vision-capable models. A text-only upstream model will reject the
image_url variant, as expected.
2026-06-04 12:23:08 +01:00
Alex 534ffc4e5b feat(v1): stateless tool continuation + sampling-param passthrough
- Stateless tool continuation. OpenAI-compatible clients (opencode, etc.)
  resend the full messages array — system, user, assistant(tool_calls),
  tool(results) — but no conversation_id, so the prior
  "conversation_id required for tool continuation" 400 broke every tool call.
  When no conversation_id is present, rebuild the agent + pending tool calls +
  tool results directly from the resent messages
  (StreamProcessor.build_continuation_from_messages) instead of loading
  server-side pending_tool_state, and call gen_continuation.

- Forward OpenAI sampling params (temperature, max_tokens,
  max_completion_tokens, top_p, frequency_penalty, presence_penalty, stop,
  seed) from the request to the LLM gen call; the agent otherwise uses its
  configured defaults.
2026-06-04 12:17:27 +01:00
Alex b1f220ec05 fix(v1): make docsgpt extension SSE frames valid chat.completion.chunks
Strict OpenAI clients (e.g. opencode / the Vercel AI SDK) validate every
streaming `data:` frame as a chat.completion.chunk. The id / source /
tool_call / tool_calls_pending events were emitted as bare `{"docsgpt": ...}`
objects with no `choices`, which those clients reject ("expected array,
received undefined" on `choices`).

Wrap the extension in an otherwise-empty chunk
(choices:[{index:0,delta:{},finish_reason:null}]) with a top-level `docsgpt`
field that OpenAI clients ignore, and skip the placeholder "None"
conversation_id frame emitted when the call is not persisted.
2026-06-04 12:07:30 +01:00
Alex 472918f577 feat(v1): support response_format / response_schema on /v1/chat/completions
Make the OpenAI-compatible Chat Completions endpoint honor per-request
Structured Outputs and keep it OpenAI-compatible.

- translate_request now forwards the request's `response_format` (json_schema)
  or a `response_schema` convenience field to the agent as its json_schema,
  overriding the agent-configured schema for that request.
- Honor `response_format.json_schema.strict` (default true); strict:false
  passes the schema through without forcing additionalProperties:false /
  all-required (OpenAI's lenient mode).
- Support `response_format {"type":"json_object"}` via the provider's native
  JSON mode.
- Keep `content` clean: some models echo their reasoning into content as
  stringified `{'type': 'thought', ...}` reprs when response_format is set.
  Strip those from content and reroute them to `reasoning_content` (OpenAI
  never puts reasoning in content), for both streaming and non-streaming.

Flow: translate_request -> StreamProcessor._configure_agent ->
Agent.json_schema / json_schema_strict / json_object -> _llm_gen ->
prepare_structured_output_format(schema, strict).
2026-06-04 11:51:13 +01:00
Alex fbd37e627c fix: issues with local tool calling on non openai endpoint 2026-06-03 22:50:03 +01:00
Alex 94bc4d1eb6 Fix remote-device tool timing out on scheduled runs (Redis-backed broker) (#2511)
* fix: route remote-device tool through Redis so scheduled runs reach the device

The remote-device tool worked interactively but timed out on every scheduled
run. DeviceBroker was an in-process, in-memory singleton, but scheduled runs
execute in the Celery worker — a different process from the gunicorn web tier
that holds the device's SSE session — so a worker-side dispatch never reached
the device and the tool always hit its deadline.

Make the broker Redis-backed so every hop crosses the process boundary:
- queued commands       -> Redis list   dev:cmd:{device_id}
- output chunks         -> Redis stream dev:out:{invocation_id}
- invocation metadata   -> Redis hash   dev:inv:{invocation_id}
- SSE upgrade tickets    -> Redis key    dev🎫{device_id}
Per-connection SSE session state stays in the web process. Reuses the existing
get_redis_instance()/CACHE_REDIS_URL; no new infrastructure. Also makes the web
tier safe to scale past one worker.

Concurrency hardening (from adversarial review + real-Redis e2e):
- XADD the output/control chunk before flipping completed=1, and have
  drain_output do a final non-blocking flush after observing completion, so a
  reader can't see completion and stop before the control chunk lands (this had
  reintroduced the false "device did not respond (timed out)" under a race).
- _collect_result builds the result from drained chunks, checks the deadline
  only after capturing a chunk, and falls back to the authoritative snapshot
  (before cleanup) when no control chunk was observed.
- Audit outcome is written from locally-known fields so it survives the worker
  racing to delete the invocation; a denied command now records a terminal
  "denied" outcome instead of staying "dispatched".
- cmd-queue TTL raised to 900s (>= max drain deadline); dispatch-failure and
  reaped-invocation cleanup; UTF-8 byte counts.

Tests: new tests/devices/{conftest (FakeRedis double), test_broker_cross_process,
test_broker_race, test_submit_output_audit}; drain/cleanup/ticket tests rewritten
for the Redis contract. The race tests fail against the pre-fix code. ruff clean;
device + tool-executor suites green.

* fix: log instead of silently passing on failed-dispatch cleanup

Addresses the code-quality lint on the best-effort hash delete in
dispatch_invocation's failure path: replace the bare `except: pass` with a
logger.debug carrying the invocation_id. No behavior change — cleanup stays
best-effort and still returns a failed Invocation.
2026-05-29 13:43:02 +01:00
Alex 418e2d989a feat: add model id on token usage (#2507) 2026-05-28 19:33:01 +01:00
Alex 8e0f7ced5f feat: remote device connection as tool (#2506)
* feat: remote device connection as tool

* fix: test

* fix: fix issues

* fix: ruff

* fix: mini issues

* fix: bugs

* fix: tests and mini issues
2026-05-28 18:23:25 +01:00
Alex b2d1d7c2c3 Deepseek (#2499)
* feat: add deepseek

* feat: deepseek reasoning tool fix
2026-05-25 23:49:25 +01:00
Manish Madan 6cfdc4fbae Feature to search conversations (#2471) 2026-05-22 16:19:54 +01:00
Alex 0507b32221 feat: default tools (#2485)
* feat: default tools

* fix: tests

* feat scheduler

* fix: scheduler UI

* Agent switch breadcrumbs

* fix: minor fixes to scheduler

* fix: tests
2026-05-22 16:05:03 +01:00
Alex 1de82ca040 fix: batch limits and failed task reque limit (#2484) 2026-05-18 22:22:43 +01:00
Alex 8f7742c937 fix: better source upload status and fix reconciliation issue (#2482)
* fix: better source upload status and fix reconciliation issue

* fix: mini issues

* chore: locale coverage
2026-05-18 14:22:03 +01:00
Alex e167cf8247 fix: broken syncs (#2480)
* fix: broken syncs

* fix: mini fixes
2026-05-17 23:58:28 +01:00
Alex e351f45d88 Feat notification system (#2472)
* feat: SSE notification system

Adds a per-user SSE pipe (GET /api/events) plus a per-message
chat-stream reconnect endpoint (GET /api/messages/<id>/events).

Backend substrate:
- application/events/ — durable journal (Redis Streams) + live
  pub/sub for user-scoped events, with publish_user_event() as
  the worker-side entrypoint.
- application/streaming/ — broadcast_channel for pub/sub fanout
  and event_replay for the per-message snapshot+tail path.
- application/storage/db/repositories/message_events.py +
  alembic 0007 — Postgres journal for chat-stream events.
- application/worker.py — ingest/reingest/remote/connector/
  attachment/mcp_oauth tasks publish queued/progress/completed/
  failed envelopes alongside their existing status updates.

Frontend client:
- frontend/src/events/ — connect/reconnect, Last-Event-ID cursor,
  backoff with jitter. Each tab runs its own connection; no
  cross-tab dedup (future work).
- frontend/src/notifications/ — recentEvents ring, cursor
  tracking, tool-approval toast.
- frontend/src/upload/uploadSlice.ts — extraReducers for
  source.ingest.* and attachment.* events.

Coverage: 132 SSE tests across events substrate, replay, journal,
routes, and worker publishes.

* refactor(attachments): remove polling, SSE-only

frontend/src/components/MessageInput.tsx no longer runs a 2s
setInterval against getTaskStatus for every processing
attachment. The attachment.* SSE reducers in uploadSlice.ts are
now the sole driver of attachment state transitions.

* feat(connector): consume source.ingest.* SSE, remove polling

frontend/src/components/ConnectorTree.tsx now mirrors FileTree's
slice-walking pattern: it watches notifications.recentEvents
for source.ingest.{completed,failed} envelopes matching the
sync's source id, and no longer polls /task_status every 2s.

* refactor(source-ingest): remove polling, SSE-only

frontend/src/upload/Upload.tsx and
frontend/src/components/FileTree.tsx no longer run getTaskStatus
polling fallbacks. The source.ingest.* SSE reducers in
uploadSlice.ts and FileTree's slice walk are now the sole
drivers of upload/reingest state transitions.

* refactor(mcp-oauth): carry authorization_url in SSE, remove polling

application/worker.py::mcp_oauth now publishes
authorization_url on the mcp.oauth.awaiting_redirect envelope.
frontend/src/modals/MCPServerModal.tsx consumes it from SSE
instead of polling /oauth_status/<task_id> every 1s.

The URL is generated inside DocsGPTOAuth.redirect_handler when
the FastMCP client triggers OAuth. The worker now plumbs a
publish callback through tool_config -> MCPTool -> DocsGPTOAuth
so the awaiting_redirect publish fires from inside the handler
at the exact point the URL becomes known. The legacy Redis
mcp_oauth_status setex writes and the GET
/api/mcp_server/oauth_status/<task_id> endpoint are kept as
belt-and-suspenders; nothing in the frontend reads them now.

* feat(source-ingest): plumb limited flag through SSE for token-cap UX

application/worker.py::ingest_worker and remote_worker now publish
``limited: bool`` on the source.ingest.completed envelope.
uploadSlice routes ``payload.limited === true`` to a failed status
with a ``tokenLimitReached`` flag, and UploadToast surfaces the
translated tokenLimit i18n string. No worker code path sets
limited=true today; this is a forward-looking contract so when
token-cap detection lands, the UX is already wired.

* refactor(mcp-oauth): read status from SSE journal, drop polling endpoint

MCPOAuthManager.get_oauth_status now walks the per-user SSE Streams
journal (user:{user_id}:stream) for the latest mcp.oauth.* envelope
matching the task id, returning the status string derived from the
event type suffix and the payload fields. The worker is the single
source of truth — its publish_user_event calls write the same
record the SSE client receives live.

Removed:
- /api/mcp_server/oauth_status/<task_id> route in
  application/api/user/tools/mcp.py
- mcp_oauth_status worker function and mcp_oauth_status_task Celery
  wrapper
- All mcp_oauth_status:{task_id} Redis setex writes (4 in mcp_oauth,
  2 in DocsGPTOAuth.redirect_handler / callback_handler)
- The update_status closure in mcp_oauth that wrote the polling
  payload

Tests updated:
- get_oauth_status now takes (task_id, user_id); new coverage walks
  a fake xrevrange response for the completed envelope, the no-match
  case, and a Redis-down case
- Removed TestMCPOAuthStatus route tests and TestMcpOauthStatusTask
  celery-wrapper test
- Removed the two oauth_status methods from the integration runner

mcp_oauth:auth_url/state/code/error Redis keys remain — they are
the OAuth flow's own state (not the dropped polling payload).

* chore(mcp-oauth): delete orphaned getMCPOAuthStatus client

The /api/mcp_server/oauth_status/<task_id> endpoint was removed in
the prior commit; the corresponding userService method and the
MCP_OAUTH_STATUS endpoint constant had no remaining callers in the
frontend, so they're deleted along with it.

* fix(events): drop live publish when journal write fails

application/events/publisher.py returned an envelope to live
pubsub subscribers even when the XADD to the durable journal
failed. The envelope had no ``id`` field, which bypassed the SSE
route's dedup floor and broke ``Last-Event-ID`` semantics for any
reconnecting client.

Best-effort delivery means dropping consistently, not delivering
inconsistent state. Now: if the journal write fails the publisher
returns None and skips the live publish entirely.

* fix(notifications): dedupe sseEventReceived against immediate dupes

Snapshot replay + live tail can both deliver the same id when the
live pubsub frame and the replay XRANGE overlap. The route's own
dedup floor catches the common case, but consumers walking
``recentEvents`` (FileTree, ConnectorTree, MCPServerModal,
ToolApprovalToast) would otherwise act on the same envelope
twice when a duplicate slipped through.

Belt-and-suspenders: short-circuit when the most recent id in
the ring matches the incoming one.

* fix(events): skip replay budget INCR when no snapshot work possible

_allow_replay incremented the per-user counter on every
/api/events GET, including no-op connects from a fresh client
with no cursor against an empty backlog. React StrictMode dev
double-mounts plus a few tabs trivially tripped the default
30-per-60s budget on idle reconnects.

XLEN pre-check: when last_event_id is None and the user stream
is empty, the connect can't do snapshot work — return True
without INCR. Cursor-bearing connects still INCR unconditionally
(probing the cursor's relationship to stream contents would
require a redundant XRANGE).

* fix(streaming): tighten journal contract + recover from seq collisions

Two related fixes to application/streaming/message_journal.py.

1. record_event now rejects non-dict payloads at the gate. The
   live path (base.py::_emit) wrapped non-dicts as
   {"value": payload}; the replay path in event_replay synthesized
   {"type": event_type}. A reconnecting client would receive a
   different envelope than the one originally streamed. Now both
   paths see byte-identical envelopes because non-dicts can't be
   journaled at all. The corresponding event_replay fallback is
   replaced with a warn-and-skip for any legacy rows.

2. record_event handles IntegrityError on (message_id, sequence_no)
   collisions by reading latest_sequence_no and retrying once with
   latest+1. The most likely cause is a stale seq seed on a
   continuation retry where the route read MAX(seq) from a
   separate connection before another writer committed past it.
   Previously the error was swallowed and the event silently
   dropped from the journal; now it lands at the next available
   seq. The live pubsub publish uses the materialised seq so the
   journal row and the live frame agree.

* perf(streaming): batch message_events INSERTs per stream

complete_stream previously opened a fresh db_session() per yielded
event, doing one Postgres INSERT + commit per chunk on the WSGI
thread. Streaming answers emit ~100s of answer chunks per response,
so the route was paying ~100 PG roundtrips per stream serialized on
commit latency.

New BatchedJournalWriter in application/streaming/message_journal.py
accumulates rows per stream and flushes on three triggers:
- size: buffer reaches 16 entries
- time: 100ms elapsed since the last flush
- lifecycle: close() at end-of-stream

Live pubsub publishes still fire synchronously per record(), so
subscribers see events in real time — only the durable journal write
is amortized. On bulk INSERT IntegrityError the writer falls back to
per-row record() with the existing seq+1 retry so a single colliding
seq doesn't drop the rest of the batch.

complete_stream wires journal_writer.close() into every exit path
(happy end, tool-approval-paused end, GeneratorExit, error handler)
so the terminal event is committed before the generator returns —
otherwise a reconnecting client could snapshot up to the last flush
boundary and live-tail waiting for an end that's still in memory.

Repository gets bulk_record() — one SQLAlchemy executemany INSERT
for the bulk path. All-or-nothing on collision (Postgres aborts the
whole batch); the writer's per-row fallback handles recovery.

* chore(upload): drop dead UploadTask.lastEventAt field

The lastEventAt field on UploadTask had no remaining consumers — the
matching Attachment.lastEventAt was cleaned up earlier. Remove the
field declaration and the slice write site.

* chore(frontend): drop orphaned getTaskStatus client

After the polling-removal sweep no caller in frontend/src/ references
userService.getTaskStatus or endpoints.USER.TASK_STATUS. The backend
route /api/task_status itself stays — agents, webhooks, e2e specs,
and the public docs still depend on it.

* docs(repo): remove stale planning docs from repo root

notification-channel-design.md, plan.md, and reminder-tool-design.md
were leftover Claude planning artifacts from the SSE substrate work
that landed accidentally. CLAUDE.md prohibits creating planning docs
unless asked — delete them.

* docs(message-events): clarify repo vs wrapper payload contract

MessageEventsRepository.record accepts any JSONB-compatible value; the
streaming wrapper record_event tightens this to dicts only because the
live and replay paths reconstruct non-dict payloads differently. Spell
the split out so the next reader of the repo method doesn't assume the
wrapper's contract applies here.

* refactor(events): raise on malformed stream id instead of lex fallback

stream_id_compare's lex-fallback branch was a footgun: a malformed id
that sorts lex-greater than a real one would pin live-tail dedup
forever, dropping every subsequent legitimate event silently. Both
current callers in application/api/events/routes.py pre-validate
inputs against _STREAM_ID_RE before calling, so changing the function
to raise ValueError is a no-op on the happy path and turns the future-
caller footgun into a loud failure.

* test(tasks): cover cleanup_message_events task body

Adds skipped-when-no-POSTGRES_URI and happy-path coverage for the
Celery janitor. The skipped path returns the documented short-circuit
shape without touching the repo. The happy path seeds a backdated
row, runs the task against the pg_conn fixture, and asserts the
retention window's row is deleted while in-window rows survive.
Mirrors the TestCleanupPendingToolState pattern.

* fix(notifications): treat /c/new as no current conversation

useMatch('/c/:conversationId') treats the literal URL /c/new as a
real conversation id, so the toast suppression check confused
'user is on /c/new' with 'user is on the conversation needing
approval'. Explicit guard: when the matched id is 'new', fall
through to the no-match case so approval toasts still surface.

* docs(events): enumerate publish_user_event None-return paths

The function returns Optional[str] today, with None conflating five
distinct outcomes (missing args / push disabled / unserialisable /
Redis down / XADD failed). Every current call site is fire-and-
forget and ignores the return, so the right move is to document the
five cases rather than promote to an enum return — keeps the API
small while making the diagnostic surface (logs) obvious. If a
future caller needs to react differently per reason, promote then.

* refactor(sources): move source-id derivation out of worker module

application/api/user/sources/upload.py imported _derive_source_id
from application.worker — pulling the entire Celery worker module
into the API process at import time just for a two-line helper.

Move DOCSGPT_INGEST_NAMESPACE and the derivation function to a
new application/storage/db/source_ids.py module that both layers
can import without that dependency edge. worker.py re-exports the
old names (_derive_source_id, DOCSGPT_INGEST_NAMESPACE) for
backward-compatible imports from tests and any other in-tree
callers; new code should import from the new module directly.

* fix(cache): enable Redis health_check_interval to surface half-open TCP

Without health_check_interval, a half-open TCP socket (NAT silently
dropped state, ELB idle-close) can leave pubsub.get_message hanging
past the SSE generator's keepalive cadence — the kernel never
surfaces the dead socket because no payload is in flight. Setting
health_check_interval=10 makes redis-py ping every 10s when
otherwise idle, so the next get_message after the dead window
raises and the SSE loop falls into its reconnect path instead of
silently freezing on the user.

* chore(events): rename attachment.processing.progress to attachment.progress

The event-type taxonomy was inconsistent: source ingest emits
source.ingest.progress (three segments) while attachments emitted
attachment.processing.progress (four segments). Drops the
.processing. infix for parity. Worker publish sites, the slice
reducer's match, and the worker tests all flip together.

No external consumers — the event type is purely internal between
the publisher and the in-tab slice; safe to rename in one commit.

* feat: events cleanup

* fix: better docs

* fix: e2e tests
2026-05-15 12:23:31 +01:00
Alex b4c4ab68f0 feat: durability and idempotency keys (#2450)
* feat: durability and idempotency keys

* feat: more durable frontend

* fix: tests

* fix: mini issues

* fix: better json validation

* fix: tests
2026-05-04 23:25:41 +01:00
Alex 552bfe016a fix: better token counting and fixes cache 2026-04-28 01:47:53 +01:00
Alex 318de18d43 feat: BYOM (#2433) 2026-04-27 22:09:33 +01:00