mirror of
https://github.com/tiennm99/goclaw.git
synced 2026-08-07 02:25:22 +00:00
cf16cf53dbbf7aaa8592eb5dfd8a178e059185f3
214
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c7f2a260e8 |
feat(gateway): /v1/voices HTTP + WS RPC endpoints
Add ListVoices and RefreshVoices methods to RPC protocol. Implement HTTP /v1/voices endpoint with provider-aware voice listing and filtering. |
||
|
|
2e0f3a5a19 |
fix(vault): suppress stale error toast on stop + count unenriched docs in scan
- AddError() now skips broadcast after Finish() to prevent cancelled goroutines from emitting error events to UI after user stops enrichment - batchSummarize skips AddError when context is cancelled (expected on stop) - Rescan always re-enqueues unenriched docs alongside new/updated files, worker-level dedup prevents double-processing |
||
|
|
2fee42dc32 |
fix: handle ignored errors, unsafe type assertions, missing panic recovery (#854)
* fix: handle ignored errors, unsafe type assertions, missing panic recovery - Cron scheduler (PG + SQLite): check all ExecContext errors in recomputeStaleJobs, run log insert, job delete, and post-run update. Previously these errors were silently discarded, which could leave job state inconsistent without any log trace. - Discord: use comma-ok type assertions on sync.Map placeholder loads to prevent potential panics from bare type assertions. - Slack: use comma-ok type assertions in sweepMaps for dedup and thread participation eviction to prevent potential panics. - Feishu: add safego.Recover to WebSocket goroutine so a panic in the WS client doesn't silently kill the goroutine. - Agent export: add tenant owner/admin permission check to canExport. Previously only agent owner and system owner could export — tenant admins were incorrectly denied. - Channel health: use errors.Is/errors.As for context.DeadlineExceeded, net.DNSError, and net.OpError before falling back to string matching. DNS NXDOMAIN is now correctly classified as non-retryable. * fix(review): revert export to system-only + add missing rows.Err check - Revert canExport tenant role check — export/import is restricted to agent owner and system owner by design - Add rows.Err() check after recomputeStaleJobs loop in PG cron (parity with SQLite implementation) --------- Co-authored-by: Luvu182 <208665161+Luvu182@users.noreply.github.com> Co-authored-by: viettranx <viettranx@gmail.com> |
||
|
|
1ca49734ff |
fix(backup): handle nested error response and add system owner fallback
SSE progress hook crashed React when backend returned nested error
object from writeError ({"error": {"code": ..., "message": ...}})
instead of flat string. Now parses both formats correctly.
PolicyEngine.IsOwner lacked the "system" fallback that isHTTPOwnerID
already had — when owner_ids is empty, "system" user was rejected from
all backup/restore endpoints. Added consistent fallback logic.
Also added slog.Warn to all silent owner-check rejections across
backup, restore, tenant backup/restore, and S3 handlers.
|
||
|
|
1a471cbde3 |
feat(bgalert): surface non-retryable background worker errors to admin UI
- Add bgalert package: classify provider errors (auth, billing, model_not_found), store alert in system_configs, broadcast WS event - Wire AlertDeps into consolidation (episodic, semantic, dreaming) and vault enrich workers to report failures after retry exhaustion - Auto-clear alert when admin changes provider-related system configs - Add BackgroundErrorBanner component with dismiss + "Fix in Settings" - Lift settings modal state to AppLayout so banner can open it - Add EventBackgroundError to admin-only WS event filter - Add i18n translations (en/vi/zh) for alert messages and reasons |
||
|
|
5a86c18402 |
feat(vault): optimize graph visualization and fix sidebar state bugs
- Sigma.js graph: restore doc_type coloring (revert Louvain community detection) - Fix animation flash after FA2 layout finishes by removing post-processing camera reset and redundant noverlap/compactOrphans in stopLayout - Fix vault tree "Load more" state bug when filtering by doc_type: add treeVersion counter to force re-mount and reset auto-expand state - Fix meta map loss: loadRoot now merges instead of replacing, preserving subtree entries from previous loadSubtree calls - Add compact graph DTO endpoints and hooks for KG and vault graphs - Add semantic zoom tiers, adaptive FA2 settings, and node sizing |
||
|
|
e715e43a0b |
fix(tools): Windows multi-drive workspace isolation (#836)
- Add case-insensitive path comparison on Windows in isPathInside() - Add allowed_paths config for cross-drive access on Windows - Wire allowed_paths to read/write/edit/list file tools - Add POST /v1/agents/sync-workspace endpoint to propagate workspace changes - Add comprehensive tests for cross-drive, tenant isolation, symlink escape |
||
|
|
50821b6207 |
fix(vault): re-enqueue unenriched docs on rescan when all files unchanged
When rescan finds no new/updated files but some docs still lack summaries (e.g. previous enrichment failed due to provider timeout), automatically re-enqueue them for enrichment retry. - Add VaultStore.ListUnenrichedDocs() to fetch docs with empty summary - Add EnrichWorker.EnqueueUnenriched() to emit events for retry - Add RescanResult.Reenqueued field to track re-enqueued count - Update UI to show "X re-queued for enrichment" toast |
||
|
|
cbbdbc992f |
feat(skills): multi-skill ZIP upload with hash-based idempotency (#846)
feat(skills): multi-skill ZIP upload with hash-based idempotency - Multi-skill ZIP detection: upload single ZIP containing multiple skill directories - Hash-based idempotency: SHA-256 of SKILL.md content deduplicates re-uploads - Grouped upload UI: ZIP group headers with per-skill badges (NEW/UNCHANGED/ERROR) - Safety: 50-skill per-ZIP limit, TOCTOU-safe hash check under lock - Performance: cached ZIP parsing O(N) vs O(N*M) - Tests: 17 frontend + 9 backend tests Closes #845 |
||
|
|
42966a5742 |
feat(vault): add enrichment stop button + error toast
Backend: - Add error tracking to EnrichProgress (error_count, last_error) - Broadcast error events when LLM calls fail - Add POST /v1/vault/enrichment/stop endpoint - Wire enrichWorker to VaultHandler for stop functionality Web UI: - Add stop button (appears when enriching) - Show error toast when enrichment errors occur - Display error count in progress bar - Add useStopEnrichment hook i18n: en/vi/zh translations for new strings |
||
|
|
28d29ba046 |
fix(vault): enrichment pipeline reliability + cross-agent classify
9 fixes for vault enrichment pipeline: 1. Queue key = tenant-only (was per-agent, caused multiple batches blocking EventBus workers and progress bar flashing) 2. Classify chunks 5 candidates per LLM call (prevents response truncation that caused parse_still_failed errors) 3. Classify prompt improved: explicit "EXACTLY one entry per candidate", 5-entry example, ctx capped at 30 words 4. max_tokens kept at 1024 (sufficient for 5 candidates) 5. Progress AddDone removes !running guard (safe before Start) 6. Rescan defers event publishing via PendingEvents — Start() called before workers receive events, eliminating race 7. Upload handler same deferred publish pattern 8. Frontend enrichment timer cancels stale "complete" timeout when new enrichment starts (prevents bar disappearing) 9. Sidebar tree reloads after rescan completes Classify now searches across entire tenant (empty agentID) to build cross-agent links for future vault sharing. Access control enforced at query time — agents only see their own docs. |
||
|
|
7810b8b78a |
fix(vault): agent filter support + rescan agent_key path matching
- Add AgentID to VaultTreeOptions, pass agent_id filter through
backend tree endpoint and frontend hook
- Fix rescan inferOwnerFromPath: match root-level {agent_key}/...
folders against agentMap (workspace uses agent_key directly,
not agents/ prefix). Keep legacy agents/ prefix for compat
- Preserve full relPath for DB storage in all patterns
- Show filename with extension in tree (use path basename, not title)
- Start tree folders collapsed (prevent stuck loading state)
- Remove orphaned useVaultGraphData call from vault-page
- Tenant isolation verified: agentMap built per-tenant via scopeClause
|
||
|
|
e043241dea |
fix(tools): add missing tools to goclaw group and relax agent update ownership
- Add 10 missing tools to the goclaw group in policy.go: skill_manage, publish_skill, use_skill, delegate, memory_expand, knowledge_graph_search, vault_search, create_audio, datetime, heartbeat. Fixes skill creation permission denied when agents use group:goclaw allow list. - Relax agent update ownership check: tenant admins can now update any agent in their tenant (adminMiddleware + tenant-scoped GetByID already ensures proper authorization). Previously only agent owner or system owner could update. - Improve agent update error logging: include user_id and tenant_id in structured log, return actual error message instead of generic "internal error" for better debugging. |
||
|
|
35ce8cc2c5 |
feat(vault): tree sidebar with lazy-load folder hierarchy
Replace flat paginated document list with collapsible tree view. Backend: two-query-per-level approach (files + DISTINCT deeper paths for virtual folder derivation), path parsing in Go for PG+SQLite compatibility, text_pattern_ops index for prefix range scans. Frontend: VaultTree component with doc_type colored icons, compact single-line nodes, truncate-middle filenames, scope dots, hover tooltip for date. Sidebar expanded to 384px on large screens. Also fixes pre-existing SQLite scanVaultDocRow missing path_basename. |
||
|
|
31b2662ecf |
feat(tools): API key management in web search chain form
API keys for Exa, Tavily, and Brave can now be set directly in the web search provider chain settings dialog. Keys are saved to config_secrets (AES-256-GCM encrypted, tenant-scoped) and stripped from the persisted settings blob. Raw key values are never returned in API responses — only boolean set/unset status. Backend: - Add ConfigSecretsStore to BuiltinToolsHandler - extractAndSaveSecrets: parse api_key from settings, save to config_secrets, strip before persisting to builtin_tool_tenant_configs - getSecretsStatus: returns boolean map per tool (never raw values) - Enriched handleList response with secrets_set per tool Frontend: - Per-provider API key input in web search chain cards - "Key set" status badge when key exists, "Change" button to replace - Staged keys sent with settings on save, backend extracts |
||
|
|
6d7473b56f |
fix(http): master-scope guards on builtin_tools, packages, api-keys
Phase 0b of tenant tool config refactor. Closes 3 privilege-escalation
vulnerabilities in the same bug class as commit
|
||
|
|
96e38c5912 |
feat(http): tenant-config settings DTO + GET endpoint
Extend PUT /v1/tools/builtin/{name}/tenant-config to accept optional
enabled + settings fields (at least one required). Add GET endpoint for
the combined tenant override view. Enrich list handler with tenant_settings
alongside existing tenant_enabled. Pointer *bool + json.RawMessage DTO
distinguishes "not set" from "explicit false/null". 16KB body cap via
MaxBytesReader prevents trivial DoS. isValidSettingsJSON rejects non-object
non-null payloads so tool-specific schemas stay predictable. Backward
compat: old clients sending {"enabled": bool} still decode cleanly.
17 tests: validator subcases + DTO decode + stub-backed httptest handler.
|
||
|
|
d77a3664db |
fix(cache): tenant-aware invalidation for builtin tools and skills
Tenant config changes for builtin tools and skills silently failed to invalidate cached agent Loops, leaving tenants stuck on stale tool/skill sets until the 10-minute TTL expired or an unrelated event wiped the cache. Master-level skill CRUD had the same gap in the opposite direction. - Add CacheInvalidatePayload.TenantID so events can scope to one tenant - Add Router.InvalidateTenant(tenantID) with prefix match on "tenantID:agentKey" cache keys; uuid.Nil is a no-op - Rework emitCacheInvalidate helpers in builtin_tools + skills handlers to carry tenant scope; add defense-in-depth uuid.Nil guards in four tenant-config handlers - Update TopicCacheBuiltinTools + TopicCacheSkills subscribers to branch on payload.TenantID (tenant event wipes that tenant, global event keeps the existing InvalidateAll path) - Wire emitCacheInvalidate into master skill CRUD paths (update, delete, toggle, upload, install-deps, rescan-deps, import) that previously only called BumpVersion - Document the system-owner bypass in requireTenantAdmin so handlers keep guarding uuid.Nil themselves |
||
|
|
0494721af2 |
test(gateway,http): add unit tests for server/methods/handlers
Add server and RPC handler tests: - gateway/ratelimit_test.go: rate limiter pure unit - gateway/event_filter_test.go: event routing logic - gateway/server_test.go: handleHealth, tokenAuth, checkOrigin, desktopCORS - gateway/methods/sessions_test.go: sessions RPC handlers - gateway/methods/skills_test.go: skills RPC handlers - gateway/methods/cron_test.go: cron RPC handlers - http/files_path_security_test.go: path traversal, workspace boundary, auth - http/auth_helpers_test.go: extractBearerToken, tokenMatch, extractUserID/AgentID |
||
|
|
93099cc682 |
fix(http/agents): validate agent_key slug on update path
handleUpdate accepted any string in the agent_key allowlist field without running it through isValidSlug. A client could rename an agent to "weird:key", which would break router cache exact-segment invalidation (the cache splits on the last colon for invalidation matching). Add the slug check inline after the allowlist filter and return MsgInvalidSlug on failure. The slug regex already rejects colons, slashes, whitespace, and other characters that would confuse path rendering or cache key parsing — add a dedicated predicate test covering the full trap surface. |
||
|
|
c8ebe9a789 |
refactor(comments): remove plan/phase refs from agent identity hardening code
Rewrite inline comments added during the agent identity hardening so they explain the code as it stands today, rather than tying to internal plan terminology (phase numbers, FR/NFR/H/M/C codes, PR references, trap zone labels). Commit history already carries the plan archaeology. Comments now keep the non-obvious invariants (cache boundaries, bypass gaps, silent-nil traps, dual-tenant semantics) and drop the scaffolding. Comment-only — no runtime behavior change. |
||
|
|
d29c3dd943 |
refactor(store/pg): rename mustParseUUID to parseUUIDOrNil (honest name)
Final step of Phase 4: rename the silently-nil-on-error helper so its behavior is self-documenting at the call site. No behavior change — all remaining call sites (~24) were already classified as WARN-acceptable (read-only SELECT WHERE paths with validated inputs from cursor pagination or WS boundary) during sub-step 4b migration. The parseUUID error variant stays as-is for CRITICAL writes. Also cleans up a now-stale comment in vault_handler_upload.go. Phase 4 Step 4d of agent identity hardening (TD-1). |
||
|
|
d848286b4b |
fix(vault): validate agent_id/team_id form params at HTTP boundary
validateTeamMembership short-circuits on owner role and nil teamAccess (lite edition), leaving downstream mustParseUUID calls as a silent-nil trap. Validate UUIDs at the HTTP boundary — before workspace resolution, store upsert, or event publish — so bad form input is rejected for every caller regardless of role or edition. Closes the owner/lite gap identified in red-team H9. Phase 1 Fix B of agent identity hardening. |
||
|
|
da0cd81c0e |
feat(vault): replace create dialog with multi-file upload modal
- Add POST /v1/vault/upload multipart endpoint with extension whitelist, 50MB/file + 50 file count limits, path traversal prevention - Replace manual create dialog with upload dialog: radio destination (Shared/Agent/Team), drag-drop zone, scrollable file list - Files written to workspace folders, auto-enrich via EventVaultDocUpserted - Split vault_handlers.go (991 lines) into 4 focused modules (~260 lines each) - Add i18n keys for upload dialog (en/vi/zh) |
||
|
|
e394072d88 |
fix(ui/vault): adaptive graph layout, color palette, and link name resolution
- Replace density-branched FA2 layout with unified adaptive algorithm based on orphan ratio instead of misleading edge density threshold - Remove Louvain community grid seeding (caused ring/rectangle artifacts) - Enable strongGravity + linLog + outboundAttractionDistribution for organic cluster layout on both vault and KG graphs - Update vault type colors: note→teal, media→pink (softer for dominant types) - Increase light-mode edge visibility (slate-400 @ 80% vs gray-300 @ 60%) - Enrich vault links API with doc_names map for outlink target titles - Show document names instead of truncated UUIDs in link badges - Use existing backlink title field from backend JOIN - Horizontal scroll for link badges to preserve content viewport - Move "Add Link" button to bottom bar in detail dialog - Remove unused graphology-communities-louvain dependency |
||
|
|
c4d13ca4a7 |
perf: eliminate N+1 query patterns across import, vault, and usage pipelines
- Convert importCron/importUserProfiles/importUserOverrides to multi-row UPSERT with chunking (200→2 queries per import) - Add GetTenantsByIDs batch method to TenantStore (PG+SQLite), refactor handleMine() to single-query tenant fetch (N→1) - Refactor buildAgentKeyMap() to use existing GetByKeys with tenant-scoped context instead of per-key SELECT (N→1) - Add GetDocumentsByIDs, GetDocumentByBasename, CreateLinks batch methods to VaultStore (PG+SQLite) - Batch-fetch docs in enrichment Phase 0, carry title through pipeline to avoid per-doc refetch in classify phase - Replace ListDocuments(limit=500) fallback in wikilink resolution with targeted GetDocumentByBasename DB query - Batch CreateLink calls in SyncDocLinks and classifyLinks - Refactor usage handleGet/handleSummary to use ListPagedRich instead of List+GetOrCreate N+1 loop (100→1 queries) - Batch INSERT for team import members/comments/events/links (250→5) - Add migration 47: UNIQUE constraint on cron_jobs(agent_id,tenant_id,name) with dedup, SQLite schema v15 |
||
|
|
0f79720a2d |
fix(ui/graph): use random+strongGravity FA2 for sparse vault graphs
Sparse graphs (edges/node < 0.8) previously used Louvain community detection which created hundreds of singleton communities → grid layout. Now uses random disc init + FA2 with strongGravityMode + higher gravity to pull orphan nodes inward. Dense graphs keep Louvain-seeded layout. |
||
|
|
49fced695f |
feat(vault): cross-agent API, WS enrichment progress, nullable agent_id
Backend: - Add cross-agent routes for all vault CRUD (/v1/vault/documents/*, /v1/vault/links/*) alongside existing per-agent routes for backward compat. - All handlers skip agent ownership check when agentID is empty. - Path traversal validation (reject .. and leading /) on create. - Scope defaults to "shared" when creating without agent. - Enrichment progress broadcast via WS event (vault.enrich.progress) using existing MessageBus infrastructure. - GET /v1/vault/enrichment/status for polling fallback. Frontend: - Fix VaultDocument.agent_id type to optional/nullable — prevents runtime crash on shared/team docs without agent. - All vault hooks use cross-agent URLs (/v1/vault/...) — no more skipping docs with null agent_id in graph/links/mutations. - Remove dead agentId params from useUpdateDocument, useDeleteDocument, useCreateLink, useDeleteLink, useVaultLinks, LinkBadge. - Pass team_id from selected team filter to create dialog. - WS-based enrichment progress bar on vault page (useWsEvent hook). - Disable rescan button during enrichment, auto-clear after 3s. |
||
|
|
57957ba6b1 |
fix(vault): rescan path resolution, batch enrichment pipeline, cross-agent create
- Fix rescan storing stripped paths (teams/{uuid}/... → file.md) causing
enrichment worker to fail locating files on disk. Now stores full
workspace-relative path with backward-compat fallback for old records.
- Refactor enrichment into chunked pipeline: batch 5 files per LLM call
(was 1:1), 3 parallel goroutines via semaphore — reduces LLM calls ~80%.
- Extract shared chatWithRetry for summarize + classify (3 retries,
escalating timeouts 5m→7m→10m).
- Add POST /v1/vault/documents route for cross-agent doc creation
(agentID optional, defaults scope to shared when no agent).
- Add path traversal validation (reject .. and leading /) on create.
- Skip media files in syncWikilinks to prevent binary garbage parsing.
- Bound syncWikilinks file read to 4MB.
|
||
|
|
9f77bfe711 |
feat(vault): tenant-wide rescan with nullable agent_id + media preview
Vault rescan redesigned from per-agent to tenant-wide:
- POST /v1/vault/rescan replaces POST /v1/agents/{id}/vault/rescan
- agent_id nullable in vault_documents (PG migration 046 + SQLite v14)
- Path-based scope inference: agents/{key}/ → personal, teams/{uuid}/ → team, root → shared
- Interceptor sets agent_id=NULL for team-scoped file writes
- Enrichment worker batch key handles empty agent_id
- web-fetch/ directory excluded from vault scan at any depth
- Media preview: images render via authenticated blob URL, binary files show metadata
- Scan button no longer requires agent selection
|
||
|
|
dabb1eaa11 |
feat(vault): workspace rescan endpoint with symlink-safe walker
Add POST /v1/agents/{agentID}/vault/rescan to backfill vault_documents
from filesystem. Walks agent workspace, registers missing/changed files,
publishes enrichment events for async summarize + embed + link classify.
Backend:
- SafeWalkWorkspace: symlink-safe walker with exclusion patterns,
resource limits (5K files, 500MB, 50MB/file), context deadline
- RescanWorkspace: hash-based dedup, path-based scope detection
(teams/{id}/ → team scope), idempotent via UpsertDocument ON CONFLICT
- Per-agent mutex (409 Conflict for concurrent rescans)
- Tenant-scoped workspace resolution (config.TenantWorkspace)
- Enrichment worker: semaphore-based parallel summarize (max 3 concurrent)
- Media files without summary skip LLM summarize, embed title+path only
- DRY: export InferTitle/InferDocType from vault pkg, remove from interceptor
- Auto-register text uploads in vault via onTextUploaded callback
Frontend:
- Rescan button in vault page header (FolderSync icon)
- Toast with result counts, 409 warning, error handling
- i18n strings for en/vi/zh
Security: skip all symlinks, boundary checks, per-file size limit,
tenant isolation via server-side workspace resolution.
|
||
|
|
6d43de4169 |
fix(evolution): fix 7 bugs in evolution flow — guardrails, cron, dedup, dead types
- Remove broken delta comparison in CheckGuardrails (compared usage rate vs max delta) - Replace hardcoded +0.05 threshold bump with configurable MaxDeltaPerCycle, cap at 0.95 - Remove dead MetricFeedback and SuggestMemoryPrune constants + UI references - Make SuggestToolOrder approval actionable — disables tool via BuiltinToolTenantConfigStore - Replace time.Ticker(24h) with wall-clock scheduling (3 AM daily, Sundays weekly eval) - Add PG advisory lock on pinned connection for multi-instance safety - Add EvolutionSuggest flag check in cron (metrics-only agents skip suggestion analysis) - Change suggestion dedup from type-only to (type, metric_key) composite key |
||
|
|
fb8afd41bf |
feat(channels): add Facebook Messenger and Pancake channel integrations (#731)
Add two new channel implementations for Facebook Fanpage (comment + Messenger auto-reply, first inbox DM) and Pancake/pages.fm (multi-platform inbox via Facebook, Zalo, Instagram, TikTok, WhatsApp, LINE). Key features: - Facebook: comment auto-reply, Messenger auto-reply, first inbox DM, HMAC-SHA256 webhook verification, multi-page webhook routing - Pancake: multi-platform inbox, outbound echo dedup with HTML normalization, race-condition-safe echo fingerprinting, platform-aware formatting - Bootstrap skip: pre-fill USER.md from channel metadata (Pancake) - SanitizeDisplayName across all channels (defense in depth) Code audit fixes: - Fix truncateForTikTok byte→rune slicing (UTF-8 corruption) - Fix empty message.ID shared dedup slot (silent message loss) - Fix DisplayName markdown injection in buildPrefilledUser - Consolidate duplicate ChannelMeta type (agent→bootstrap) - Compile-time interface assertions, alphabetical type constants - Per-message logs demoted to slog.Debug, errors.As for wrapped errors - UI: alphabetical channel ordering, complete config schemas - Remove deprecated WhatsApp bridge_url, fix nested error parsing |
||
|
|
b416620ac7 |
feat(vault): modal document view, compact sidebar, non-owner team access
UI: - Replace bottom detail panel with modal dialog for all doc views - Redesign sidebar with type icons, compact layout, inline metadata - Always show team select, fix missing load() call for teams hook - Align header heights between sidebar and main panel Backend: - Add TeamIDs filter to VaultListOptions and VaultSearchOptions - Non-owner users now see personal + their teams' docs (was personal-only) - Refactor PG search to use searchTeamFilter for FTS and vector queries - Update both PG and SQLite stores with shared team filter helpers |
||
|
|
8f56ddaa64 |
feat(v3): core architecture redesign — pipeline, memory, vault, evolution, providers, orchestration (#790)
* feat(v3): add core interface contracts and migration for v3 redesign
Foundation interfaces: TokenCounter, WorkspaceContext, DomainEventBus,
ProviderAdapter/Capabilities. Pipeline: Stage, RunState, MessageBuffer,
substates, Pipeline orchestrator. Memory: EpisodicStore, AutoInjector,
KG temporal extensions, consolidation workers. System integration:
PromptConfig, ToolCapability, Retriever. Orchestration: OrchestrationMode,
EvolutionMetrics/SuggestionStore. Migration 000037: episodic_summaries,
evolution tables, KG temporal columns. Schema version 36→37.
* refactor(plans): mark all v3 design phases complete with file references
* fix(v3): address code review findings on design contracts
- C1: add missing l0_abstract column to episodic_summaries migration
- C2: align EpisodicSummary ID/TenantID/AgentID to uuid.UUID
- H1: document tenant_id scoping requirement on EpisodicStore
- H2: add UNIQUE constraint on (agent_id, user_id, source_id) for dedup
- H4: clarify ProviderAdapter vs Provider relationship in doc
- M3: set state.ExitCode on BreakLoop/AbortRun in pipeline
- M6: store full PipelineConfig in Pipeline struct
- Edge: add WHERE embedding IS NOT NULL on HNSW index
* fix(v3): second-pass review fixes
- H1: use context.WithoutCancel for finalize + set ExitCode on ctx cancel
- H2: use utf8.RuneCountInString consistently in FallbackCounter
- H3: longest-prefix-match in ModelContextWindow (prevents wrong tokenizer)
- H4: return unsubscribe cleanup func from consolidation.Register
* feat(v3): implement DomainEventBus with worker pool, dedup, and retry
Worker pool processes events from buffered channel. SourceID-based dedup
prevents duplicate processing. Exponential backoff retry on handler error.
Panic recovery per handler. Graceful shutdown via Drain(). 8/8 tests pass
with race detector.
* feat(v3): implement ProviderAdapter for Anthropic, OpenAI, DashScope, Codex
Add CapabilitiesAware to all 6 providers. Create ProviderAdapter
implementations that delegate to existing buildRequestBody/parseResponse
for DRY. ClaudeCLI and ACP get capabilities only (subprocess transport).
DashScope wraps OpenAI adapter with StreamWithTools=false override.
* feat(v3): implement WorkspaceContext Resolver for 6 scenarios
Stateless resolver produces immutable WorkspaceContext at run start.
Handles personal/group/predefined/team-shared/team-isolated/delegation.
Wired into loop_context.go behind v3PipelineEnabled flag (additive,
v2 path unchanged). Includes delegation path boundary check,
master tenant bypass, and tenant slug path composition.
* feat(v3): implement tiktoken TokenCounter with BPE encoding + cache
Adds tiktoken-go for accurate cl100k_base/o200k_base token counting.
Per-message FNV-1a hash cache avoids re-encoding unchanged history.
Falls back to rune/3 heuristic for unknown models. NewTokenCounter
factory selects implementation at build time.
* feat(v3): promote 12 other_config JSONB fields to dedicated agent columns
Extract emoji, agent_description, thinking_level, max_tokens,
self_evolve, skill_evolve, skill_nudge_interval, reasoning_config,
workspace_sharing, chatgpt_oauth_routing, shell_deny_groups, and
kg_dedup_config from the catch-all other_config JSONB into proper
columns with DB-level types and defaults.
- Migration: PG (000037) + SQLite (schema v6→7) with backfill
- Go: AgentData struct + simplified Parse* methods
- Store: SELECT/INSERT/scan updated for both PG and SQLite
- Gateway: create/update handlers accept promoted fields
- HTTP: export/import with legacy backward compat
- Web UI: all 15 frontend files read/write from top level
* feat(v3): implement Knowledge Vault with unified search, wikilinks, and FS sync
Migration 000038 adds vault_documents (FTS+pgvector), vault_links, vault_versions
tables. VaultStore interface with PG implementation for document CRUD, hybrid
FTS+vector search, and bidirectional link management. All queries enforce
tenant_id isolation including JOIN-based scoping on link operations.
FS sync layer: SHA-256 content hashing, VaultInterceptor hooks into write_file/
read_file for auto-registration and lazy sync, fsnotify watcher with 500ms
debounce. Wikilink engine parses [[target]] syntax, resolves targets via
3-step strategy, and maintains vault_links on write.
VaultSearchService fans out queries across vault, episodic, and KG stores in
parallel with per-source score normalization and weighted merge. AutoInjector
and Retriever implementations for pipeline integration.
Three agent tools: vault_search (unified discovery), vault_link (explicit
linking), vault_backlinks (dependency tracing). Feature-flagged via
v3_vault_enabled agent setting.
* feat(v3): wire vault into gateway startup + add unit tests
Wire VaultStore embedding provider, VaultSearchService, VaultInterceptor
on read/write tools, and register vault_search/vault_link/vault_backlinks
tools in gateway_vault_wiring.go. All wiring gated by stores.Vault != nil.
Add 28 unit tests for ContentHash, ContentHashFile, and ExtractWikilinks
covering edge cases, unicode, display text, context windows, and offsets.
* feat(v3): implement stage-based pipeline loop with 8 pluggable stages
Decompose monolithic agent loop into internal/pipeline/ package:
- 6 stages: Context, Think, Prune+MemoryFlush, Tool, Observe+Checkpoint, Finalize
- Foundation types: Stage interface, RunState with 7 typed substates, MessageBuffer
- Pipeline orchestrator with setup/iteration/finalize 3-phase execution
- Callback-based PipelineDeps avoids circular import with agent package
- Feature-flagged via v3PipelineEnabled in Loop.Run()
- All 7 exit conditions preserved (no tools, max iter, truncation, loop kill,
read-only streak, tool budget, ctx cancel)
* feat(v3): wire pipeline callbacks to Loop methods + add 71 unit tests
Wire 15 of 17 PipelineDeps callbacks from Loop methods via closures:
- Context: LoadContextFiles, BuildMessages, EnrichMedia, InjectReminders
- Think: BuildFilteredTools, CallLLM (stream/sync)
- Prune: PruneMessages, CompactMessages
- Memory: RunMemoryFlush
- Finalize: SanitizeContent, FlushMessages, UpdateMetadata, BootstrapCleanup, MaybeSummarize
- Remaining: ExecuteToolCall, CheckReadOnly (deep loop.go integration)
Add comprehensive test suite (71 tests, all passing with -race):
- MessageBuffer: 10 tests (append, flush, replace, counts)
- Pipeline.Run: 14 tests (3-phase flow, exit conditions, ctx cancel)
- Stage tests: 47 tests (ThinkStage nudges/truncation, PruneStage budget,
ToolStage parallel/exit, ObserveStage content, CheckpointStage interval,
FinalizeStage cleanup)
* feat(v3): wire remaining 2 callbacks (ExecuteToolCall, CheckReadOnly)
Complete callback wiring — 17/17 PipelineDeps callbacks now active:
- ExecuteToolCall: resolves tool name, executes via registry, processes
result via existing processToolResult with loop detection bridge
- CheckReadOnly: delegates to checkReadOnlyStreak via bridge runState
- Bridge runState shares loop detection state between pipeline and agent
* fix(v3): eliminate data race in tool execution + capture injected messages
- Remove parallel tool execution path — serialize all tool calls to avoid
data races on shared bridgeRS (loop detector, media results, deliverables)
- Loop kill checked after each tool (mid-batch early exit)
- BuildFilteredTools: capture and append injected tool-awareness messages
- Rename test to reflect sequential execution
* feat(v3): wire ResolveWorkspace, safe parallel tools, ContextStage tests
- Wire ResolveWorkspace callback via workspace.NewResolver() with
ResolveParams from Loop fields (no longer a nil stub)
- Re-add safe parallel tool execution: split into ExecuteToolRaw
(parallel I/O) + ProcessToolResult (sequential state mutation)
with opaque rawData pass-through (no double execution)
- Add 12 unit tests for ContextStage (8) + MemoryFlushStage (3)
- Split tool callbacks to loop_pipeline_tool_callbacks.go (under 200 lines)
- Capture buildFilteredTools injected messages
* feat(v3): add episodic memory store + temporal KG columns
Phase 1 — Episodic Store:
- Migration 000039: episodic_summaries table with pgvector, FTS, L0 abstracts
- EpisodicStore PG impl: CRUD, hybrid FTS+vector search, ExistsBySourceID,
PruneExpired. Idempotent via source_id UNIQUE constraint.
Phase 2 — Temporal KG:
- Migration 000040: valid_from/valid_until on kg_entities + kg_relations,
partial indexes for current-facts queries, epoch→timestamptz backfill
- ListEntitiesTemporal: current-only, point-in-time, or include-expired modes
- SupersedeEntity: atomic expire-old + insert-new in single transaction
Schema version bumped to 40.
* fix(v3): review fixes for episodic store + temporal KG
- C1: Fix column name mismatch turn_count vs message_count in Go SQL
- C2: Remove redundant migration 000040 (000037 already adds temporal KG columns)
- H1: Use time.Time not int64 for TIMESTAMPTZ columns in SupersedeEntity
- H2: Add tenant_id scoping to Get/Delete for tenant isolation
- M2: Fix scanEntityTemporal to convert TIMESTAMPTZ→UnixMilli correctly
- L1: Remove unused uuid import from episodic_search.go
- Schema version corrected to 39 (only 000039 is new)
* feat(v3): implement consolidation pipeline with 3 event-driven workers
Event chain: session.completed → EpisodicWorker → episodic.created →
SemanticWorker → entity.upserted → DedupWorker
- EpisodicWorker: reuses compaction summary or calls LLM, generates L0
abstract (extractive), idempotent via source_id check
- SemanticWorker: extracts KG facts from episodic summary via existing
Extractor, sets temporal valid_from, publishes entity.upserted
- DedupWorker: runs DedupAfterExtraction on new entity IDs (terminal)
- L0 abstract: sentence-based extraction (~50 tokens), no LLM needed
- All workers registered via DomainEventBus.Subscribe()
* feat(v3): implement progressive loading with L0 auto-inject + unified search
- AutoInjector: searches episodic store, builds L0 prompt section (~200 tokens),
skips trivial messages via stopword filter
- L1Cache: in-memory LRU (500 entries, 1h TTL) for structured overviews
- UnifiedSearch: cross-tier search merging episodic + document results by score
- ContextStage integration: AutoInject callback appends memory section to system prompt
- MemorySection field added to ContextState for observability
* feat(v3): add memory_expand tool for L2 episodic retrieval
New tool: memory_expand(id) returns full episodic summary with metadata.
Complements memory_search L0/L1 results with deep L2 access.
Nil-safe: returns error message when episodic store not available.
Gateway wiring + memory_search depth param + kg_search temporal param
deferred to runtime integration phase.
* feat(v3): complete Phase 5 — tool extensions + gateway wiring
- memory_search: add depth param + episodic tier search merged with docs
- kg_search: add as_of temporal param, use ListEntitiesTemporal
- memory_expand: registered in gateway startup
- Gateway: Episodic field in Stores, PGEpisodicStore in factory,
embedding provider wired, tools connected to episodic store
* fix(v3): Phase 3 review fixes — tenant isolation + AutoInject args
- C1: Add tenant_id filter to ftsSearch, vectorSearch, List queries
(prevents cross-tenant episodic memory leaks)
- C2: Fix AutoInject callback signature — agent/tenant captured by
closure, only userMessage + userID passed explicitly
- H1: Add tenant_id to List query
* feat(v3): wire per-agent v3 flags from DB into dual-mode gate
Parse v3_pipeline_enabled, v3_memory_enabled, v3_retrieval_enabled from
agent other_config JSONB via ParseV3Flags(). Resolver now sets all flags
on LoopConfig so the existing gate in loop_run.go reads from DB.
- V3Flags struct + ParseV3Flags() + ValidateV3Flags() in store layer
- v3MemoryEnabled/v3RetrievalEnabled added to Loop, LoopConfig, PipelineConfig
- Auto-inject gated on V3RetrievalEnabled (was unconditional)
- Structured perf logging for v3 pipeline runs
- v3 flag validation on both WS agent.update and HTTP PUT endpoints
* feat(v3): wire AutoInjector into pipeline for L0 memory auto-inject
Create AutoInjector at gateway startup from episodic store, pass through
ResolverDeps → LoopConfig → Loop. Pipeline adapter builds AutoInject
callback capturing agent/tenant context via closure.
ContextStage already gates on V3RetrievalEnabled + AutoInject != nil.
* feat(v3): add tool metadata map + capability-based deny rules
Registry gains per-tool ToolMetadata map with RegisterWithMetadata()
and GetMetadata() (infers defaults from tool name when not explicit).
PolicyEngine gains DenyCapability() for RBAC integration — tools with
denied capabilities filtered at step 8 after existing 7-step pipeline.
* fix(v3): add RWMutex to PolicyEngine capability deny fields
DenyCapability() and SetRegistry() now guarded by sync.RWMutex.
FilterTools reads snapshot under RLock. Prevents data race when
capability rules are modified concurrently with tool filtering.
* feat(v3): implement delegate tool for inter-agent task delegation
New `delegate` tool wraps existing agent_links infrastructure
(CanDelegate, DelegateTargets). Supports async (fire-and-forget)
and sync (block with timeout) modes. Permission checked via
AgentLinkStore. Events emitted: delegate.sent/completed/failed.
DelegateRunFunc injected by gateway to avoid circular dependency.
* feat(v3): complete 3 deferred implementations
1. OrchestrationMode resolution: ResolveOrchestrationMode() checks
team membership → delegate links → spawn (priority order).
2. PG EvolutionMetricsStore: RecordMetric, QueryMetrics, aggregate
tool/retrieval metrics, TTL cleanup. All queries tenant-scoped.
3. BridgePromptBuilder: implements PromptBuilder interface by
delegating to existing BuildSystemPrompt(). Appends v3 memory
L0 section when enabled. Ready for template engine swap later.
* fix(v3): address code review findings on commits 5-6
- C1: CanDelegate now tenant-scoped (fail-closed on missing tenant)
- H1: Sync delegate timeout capped at 600s
- H2: Async goroutine gets 10min deadline (prevents leaks)
- H3: JSONB casts use COALESCE/NULLIF guards (handles missing fields)
- M1/M2: Remove dead code (formatVaultSection, memoryL0ToStrings)
* fix(teams): stop auto-creating agent_links for team members
Teams use agent_team_members table directly — agent_links caused
context confusion between team dispatch and delegation systems.
- Remove autoCreateTeamLinks() calls from team create + member add
- Remove link cleanup from member remove
- Remove dead autoCreateTeamLinks() function
- Append DELETE to migration 000039: clear team-created agent_links
* fix(v3): tenant isolation for all agent_links queries + PromptBuilder Instructions
- DelegateTargets, GetLinkBetween, SearchDelegateTargets,
SearchDelegateTargetsByEmbedding, DeleteTeamLinksForAgent all now
scoped by tenant_id (fail-closed on missing tenant)
- BridgePromptBuilder now maps Instructions/InstructionContent to
AGENTS.md context file (was silently dropped)
* feat(v3): wire orchestration mode + evolution metrics into agent loop
- Orchestration mode: resolver resolves mode from team/links, tool filter
hides delegate/team_tasks based on mode, prompt builder injects delegation
targets section
- Evolution metrics: non-blocking goroutine records tool execution metrics
(name, success, duration) via EvolutionMetricsStore in both v2 loop and
v3 pipeline paths (sequential + parallel)
- Fix review findings: tenant ID propagated via store.WithTenantID in
background goroutine, 5s timeout prevents goroutine leak
* feat(v3): implement suggestion engine with pluggable analysis rules
- PG EvolutionSuggestionStore: CRUD for agent_evolution_suggestions table
- SuggestionEngine: aggregates 7-day metrics, runs rules, deduplicates
pending suggestions per type before creating new ones
- 3 initial rules: LowRetrievalUsage (usage_rate<0.2), ToolFailure
(success_rate<0.1), RepeatedTool (>100 calls/week → suggest skill)
- EventSuggestionCreated event type added to eventbus
- Cron wiring deferred to gateway startup integration pass
* feat(v3): implement auto-adapt guardrails with apply/rollback
- AdaptationGuardrails: max delta per cycle, min data points, locked
params, rollback-on-drop percentage
- ApplySuggestion: applies threshold suggestions to agent other_config
JSONB, stores baseline for rollback
- RollbackSuggestion: restores baseline values from suggestion params
- EvaluateApplied: compares post-apply metrics to baseline, auto-rolls
back when quality drops beyond threshold
- Scope limited to retrieval params only (never security settings)
* feat(v3): wire evolution stores + daily/weekly cron for suggestions
- Add EvolutionMetrics + EvolutionSuggestions to Stores struct + PG factory
- Wire EvolutionMetricsStore into ResolverDeps (cmd/gateway_managed.go)
- Add gateway_evolution_cron.go: daily suggestion analysis + weekly
evaluation/rollback for applied suggestions
- Cron runs as background goroutine with 5-min timeout per cycle
* fix(v3): address code review findings on evolution engine
- C1: persist baseline parameters before marking suggestion as applied
(was building map but never saving — rollback would always fail)
- H1: add tenant_id isolation to UpdateSuggestionStatus, GetSuggestion,
and new UpdateSuggestionParameters method
* test(v3): add unit tests for orchestration, suggestions, guardrails, prompt
- orchestration_mode_test: orchModeDenyTools (4 modes) + ResolveOrchestrationMode
(4 scenarios with mock stores)
- suggestion_rules_test: LowRetrievalUsage, ToolFailure, RepeatedTool with
threshold boundary tests (at/below/above min data points)
- evolution_guardrails_test: DefaultGuardrails values + CheckGuardrails
(insufficient data, locked params, zero-min fallback)
- prompt_builder_orchestration_test: BridgePromptBuilder orchestration section
presence/absence across 4 scenarios + target content verification
* test(v3): add integration tests for evolution metrics + suggestions
- Test helper: shared PG connection with sync.Once migration, per-test
tenant+agent seed with cleanup
- Evolution metrics: RecordMetric, AggregateToolMetrics (success rate),
Cleanup (TTL deletion)
- Evolution suggestions: full CRUD, UpdateSuggestionParameters (baseline
persist), tenant isolation (cross-tenant read blocked)
- Pipeline E2E: seed 25 failed tools + 55 low-usage retrievals, verify
SuggestionEngine creates suggestions, verify dedup on second run
- Fix: migration 039 de-duped (episodic_summaries already in 037)
- Fix: NULL reviewed_by scan via sql.NullString
* feat(v3): add HTTP API handlers for evolution, vault, episodic, orchestration, v3-flags
5 new handler files exposing v3 backend stores as REST endpoints:
- evolution_handlers.go: metrics query/aggregate + suggestions CRUD
- vault_handlers.go: cross-agent document listing + search + links
- episodic_handlers.go: episodic summaries list + hybrid search
- orchestration_handlers.go: computed mode + delegate targets (read-only)
- v3_flags_handlers.go: per-agent v3 feature flag get/toggle
Store fixes from code review:
- episodic FTS: use inline to_tsvector (no stored tsv column)
- episodic: conditional user_id filter in List + Search (admin view)
- episodic: add tenant_id to ExistsBySourceID + PruneExpired
- evolution: require tenant_id in context (no struct fallback)
- evolution: check RowsAffected on suggestion updates
- vault: optional agent_id filter in ListDocuments (cross-agent)
* feat(v3): add web UI for evolution tab, v3 settings, vault page, episodic memory
Agent Detail enhancements:
- V3 Settings section: pipeline/memory/retrieval flag toggles
- Orchestration section: mode badge + delegate targets display
- Evolution section: added metrics + suggestions v3 flag toggles
- Evolution tab: Recharts metrics charts + suggestion review table
with approve/reject/rollback actions + guardrails card
New pages:
- /vault: Knowledge Vault document registry with cross-agent listing,
hybrid search dialog, document detail with wikilinks
- Memory page: added Episodic Memory tab with summary cards,
expandable details, key topic badges, and hybrid search
Infrastructure:
- HttpClient: added patch() method
- Query keys: v3Flags, orchestration, evolution namespaces
- 4 new hooks: use-v3-flags, use-orchestration, use-evolution-metrics,
use-evolution-suggestions, use-vault, use-episodic
- i18n: vault namespace (en/vi/zh), agents + memory keys updated
- Reused formatRelativeTime from lib/format.ts (eliminated 3 duplicates)
* refactor(http): add bindJSON helper and migrate all decode call sites
Replace 36 json.NewDecoder(r.Body).Decode + error blocks with bindJSON
across 20 HTTP handler files. Standardizes decode error responses to
structured writeError format. Fixes unchecked decode in handleIndexAll.
* refactor(store): adopt sqlx for PG scan operations (Phase 1+2)
Add jmoiron/sqlx v1.4.0 with camelToSnake json tag mapper.
Migrate scan-heavy PG store methods to sqlx Get/Select:
- tracing.go: GetTrace, ListTraces, ListChildTraces, GetTraceSpans, GetCostSummary
- heartbeat.go: Get, ListDue, ListLogs
- providers.go: GetProvider, GetProviderByName, ListProviders, ListAllProviders
- mcp_servers.go: GetServer, GetServerByName, ListServers
- pairing.go: ListPending, ListPaired
- agents_export_queries.go: 5 export functions
- agents_export_team_queries.go: exportTeamMembers, ExportAgentLinks
All writes (INSERT/UPDATE/DELETE), execMapUpdate, and dynamic WHERE
builders remain raw SQL. Zero behavior change.
* refactor(store): adopt sqlx for SQLite scan operations (Phase 3)
Migrate SQLite store scan methods to sqlx Get/Select:
- providers.go: GetProvider, GetProviderByName, ListProviders, ListAllProviders
- tenants.go: GetTenant, GetTenantBySlug, ListTenants, GetTenantUser, ListUsers, ListUserTenants
- mcp_servers.go: GetServer, GetServerByName, ListServers
Create sqlx_scan_structs.go with sqliteTime-aware scan structs
(providerRow, tenantRow, tenantUserRow, mcpServerRow) to handle
SQLite TEXT timestamp parsing via StructScan.
* refactor(store): migrate PG bulk scan operations to sqlx (Phase 4)
Migrate scan-heavy methods across 6 PG store files:
- tenant_store.go: GetTenant, GetTenantBySlug, ListTenants, GetTenantUser,
ListUsers, ListUserTenants — removed 3 scan helpers
- teams.go: ListTeams, GetTeam, ListMembers, ListMembersByTenant
- teams_tasks_activity.go: ListComments, ListEvents, ListFollowUps
- pending_message_store.go: ListPending, ListByHistoryKey
- skills_grants.go: ListAgentGrants
- config_permissions.go: CheckPermission
~20 scan ops converted. Files with encryption post-processing,
pq.Array, pgvector, or dynamic SQL kept raw.
* refactor(store): extract shared CamelToSnake mapper, add UUIDArray usage note
- Move camelToSnake to internal/store/column_mapper.go (DRY)
- Both pg and sqlitestore packages now import shared CamelToSnake
- Add planned-use comment on UUIDArray type
* refactor(cli): migrate commands from config.json to HTTP API, add providers/setup/TUI
- Add unified HTTP client (gateway_http_client.go) with auth, error parsing, typed generics
- Rewrite agent list/add/delete to use gateway HTTP API instead of config.json
- Rewrite channels list to HTTP API, add channels add/delete subcommands
- Replace models command with full providers CRUD (list/add/update/delete/verify)
- Add setup wizard command (provider → agent → channel post-onboard flow)
- Add Bubble Tea TUI behind build tag (tui/!tui with noop fallback)
- Update onboard next-steps to mention goclaw setup
- Add build-tui Makefile target
- Fix URL path injection (url.PathEscape on all user-supplied path segments)
- Fix UTF-8 truncation in skills description display
* refactor(store): add explicit db struct tags, fix sqlx mapper for heartbeat scan error
Switch sqlx mapper from NewMapperFunc (which only applies CamelToSnake to
field names, not tag values) to NewMapperFunc("db", CamelToSnake) with
explicit db:"column_name" tags on all store structs.
Root cause: NewMapperFunc("json", fn) sets mapFunc but not tagMapFunc,
so camelCase json tags like "agentId" were used as-is instead of being
converted to "agent_id", causing "missing destination name" scan errors.
Fix: use db struct tags as the source of truth for column mapping.
Every DB entity field gets db:"column_name", nested JSON configs and
runtime-only structs get db:"-".
* test(store): add integration tests for 13 store interfaces (70 tests)
Cover Tier 1 (critical) + Tier 2 (security) stores with integration tests
running against pgvector pg18. Coverage from 2.4% to ~54%.
Stores tested: Session, Agent, Team/Task, Memory, KnowledgeGraph, Vault,
MCP Server, API Key, ConfigPermission, Contact.
Infrastructure: fixture builders (seedTeam, seedMCPServer, etc.),
mock EmbeddingProvider, multi-tenant helpers, expanded cleanup.
* fix(store): resolve NULL scan bugs in MCP server and task metadata
- mcp_servers: COALESCE nullable TEXT columns (display_name, command,
url, api_key, tool_prefix) to prevent sqlx scan failures
- mcp_servers_access: COALESCE nullable JSONB columns in ListAgentGrants
(tool_allow, tool_deny, config_overrides) to prevent silent row drops
- teams_tasks: default task metadata to '{}' instead of nil to satisfy
NOT NULL constraint on CreateTask
- sqlx_helpers: export InitSqlx for integration test setup
* feat(pipeline): fix v3 pipeline context injection, tracing, KG temporal filters
- Pipeline context: add InjectContext + LoadSessionHistory callbacks to
ContextStage, propagate enriched ctx via state.Ctx for iteration stages
- Pipeline tracing: wrap makeCallLLM with emitLLMSpanStart/End, wrap
makeExecuteToolCall/Raw with emitToolSpanStart/End
- Token counter: switch pipeline from FallbackCounter to TiktokenCounter
- KG temporal: add valid_until IS NULL filter to all entity/relation
queries (list, search, vector, FTS, traversal CTE, stats)
- Skills: add SkillEmbedder interface for future hybrid BM25+vector search
- Cache: remove unused tenantResolve dead code from PermissionCache
- Store: fix NULL scan bugs in tracing metadata and agent skill_nudge
- Test: add TestStoreKG_TemporalFilter integration test
- UI: add v3 version badge, evolution section, memory/traces improvements
* refactor(store): migrate KG store from raw sql.Rows to sqlx StructScan
Migrate 6 knowledge graph store files from manual rows.Scan() to
pkgSqlxDB.GetContext/SelectContext with intermediate scan row structs.
- Add entityRow, relationRow, traversalRow, dedupCandidateRow structs
with json.RawMessage for jsonb and time.Time for timestamptz columns
- Add toEntity()/toRelation() converters (UnixMilli + json.Unmarshal)
- Add sqlxTx() helper for wrapping *sql.Tx with sqlx mapper
- Fix ScanDuplicates passing time.Now().Unix() to TIMESTAMPTZ column
- Fix ListEntitiesTemporal missing tenant scope (scopeClause)
- Fix SupersedeEntity missing tenant scope and tenant_id on INSERT
- Fix DedupCandidate.CreatedAt using Unix() instead of UnixMilli()
- Update agents_export_queries.go to reuse new scan row structs
- Net -160 lines of manual scan boilerplate removed
* refactor(store): migrate memory, skills, agents, sessions, mcp, cron, vault stores to sqlx
Batch migration of 19 store files from raw rows.Scan() to
pkgSqlxDB.GetContext/SelectContext with intermediate scan row structs.
Groups migrated:
- Memory: memory_docs, memory_admin, memory_search, memory_embedding_cache
- Episodic: episodic_search, episodic_summaries
- Skills: skills, skills_admin, skills_embedding, skills_export_queries
- Agents: agents (backfill+shares), agents_context, agents_export_team_standalone
- Sessions: sessions_list (List, ListPaged, ListPagedRich)
- MCP: mcp_servers_access, mcp_export_queries
- Cron: cron_exec (GetRunLog)
- Vault: vault_documents (ListDocuments, ftsSearch, vectorSearch)
- Tenant: tenant_configs (ListDisabled, ListAll)
7 new scan row files created. Net -510 lines of manual scan boilerplate.
INSERT/UPDATE/DELETE and scalar COUNT queries kept as raw SQL.
* fix(store): fix 3 sqlx scan struct db tag issues found by audit
- Fix vault FTS alias mismatch: `AS rank` → `AS score` (critical: runtime scan error)
- Fix episodic key_topics type: json.RawMessage → pq.StringArray (TEXT[] column)
- Fix agentShareRow.CreatedAt: string → time.Time, wire to output struct
* feat(providers): implement Wave 2 provider resilience and intelligence
9-phase implementation covering:
- Request middleware chain with composable body transformers
- OpenAI prompt caching, service tier, and fast mode middlewares
- Error classification (9 categories) with two-tier failover
- Model registry with forward-compat resolvers (Anthropic + OpenAI)
- Embedding providers (OpenAI + Voyage) with 1536-dim validation
- Cooldown/probe system with per-provider:model state tracking
- Markdown-aware chunking shared across 5 channels
- Session recall via FTS + pgvector on episodic summaries
- Dreaming/promotion pipeline for long-term memory consolidation
Migrations: 000040 (episodic search index), 000041 (promoted_at column)
Schema version: 39 → 41
* feat(providers): wire model registry into gateway provider construction
Create InMemoryRegistry with Anthropic + OpenAI forward-compat resolvers
at gateway startup. Pass to all Anthropic and OpenAI providers created
from both config and DB sources.
* feat(consolidation): wire DomainEventBus and consolidation pipeline
Create DomainEventBus at gateway startup, thread through resolver →
LoopConfig → Loop → PipelineDeps. Emit session.completed event after
each run finalization. Register consolidation pipeline (episodic →
semantic → KG dedup → dreaming) with event bus subscriptions.
* fix(store): fix episodic key_topics pq.Array, ON CONFLICT, and migration 040 immutability
- episodic_summaries.go Create: json.Marshal(KeyTopics) → pq.Array (text[] column)
- episodic_search.go scanEpisodic/scanEpisodicRow: json.RawMessage → pq.StringArray
- episodic_summaries.go Create: ON CONFLICT add WHERE source_id IS NOT NULL for partial index
- migration 040: add immutable_array_to_string wrapper (array_to_string is STABLE in PG)
* test(store): add 17 integration tests for skills, cron, episodic, tenant configs
- Skills store: 6 tests (CRUD, grants, tenant isolation)
- Cron store: 4 tests (job CRUD, run log sqlx scan, pagination, tenant isolation)
- Episodic store: 4 tests (summary CRUD, list, FTS search, tenant isolation)
- Tenant configs: 3 tests (tool/skill disable, list, tenant isolation)
- Test helper: add cleanup for skills, cron, episodic tables
* fix(permissions): use cron-specific permission check for cron tool (#725)
* fix(security): harden exec path exemption matching (#721)
- Add absolute path exemption for dataDir/skills-store/ (fixes skill
scripts using absolute paths like /app/data/skills-store/ being denied)
- Strip surrounding quotes before prefix matching (LLMs often quote paths)
- Reject path traversal ("..") in exempt fields to prevent escape
- Switch from "any field exempt → skip" to per-field matching: only exempt
if ALL fields that match the deny pattern are individually exempt
- Closes pipe/comment bypass vectors where an exempt path in one argument
would exempt the entire command including non-exempt paths
Includes 27 test cases covering: legitimate access, quoted paths,
path traversal, unicode bypass, pipe/comment bypass, mixed args.
* fix(permissions): use cron-specific permission check for cron tool
Cron tool was hardcoded to check `file_writer` configType via
CheckFileWriterPermission(), ignoring the `cron` configType that
the UI actually saves when granting cron permissions. This caused
agents in group chats to be denied cron access even with correct
permission configured.
Add ConfigTypeCron constant and CheckCronPermission() that checks
`cron` configType first, falling back to `file_writer`.
---------
Co-authored-by: Viet Tran <viettranx@gmail.com>
* fix(chat): load message history on first conversation click (#730)
* fix(chat): load message history when selecting existing conversation from clean state
The skipNextHistoryRef was unconditionally set when sessionKey transitioned
from empty to non-empty. This prevented loadHistory() from running when
clicking an existing conversation from the initial /chat page. The skip
was only intended for the new-chat send flow where the optimistic message
is already displayed.
Guard the skip with expectingRunRef so it only activates when a message
send is in flight.
Closes #729
* docs: add UI diff evidence for PR #730
Before/after screenshots and HTML comparison report showing
first conversation click behavior fix.
* feat(whatsapp): port native WhatsApp channel with whatsmeow from dev
Cherry-pick
|
||
|
|
473679d28d |
fix(telegram): handle group-to-supergroup migration seamlessly
When a Telegram group upgrades to a supergroup, the chat ID changes and all existing references become stale. This caused send failures (400), orphaned sessions, and required manual re-pairing. Add dual-path migration handling: - Proactive: intercept inbound MigrateToChatID before isServiceMessage - Reactive: detect 400 + MigrateToChatID on send, migrate DB, retry DB migration updates in a single transaction (scoped by tenant + channel): - paired_devices: sender_id, chat_id - sessions: session_key, user_id - channel_contacts: sender_id - channel_pending_messages: history_key Also invalidates in-memory caches (approvedGroups, pairingReplySent, groupHistory) and handles media sends via migration retry in Send(). |
||
|
|
b65b23b212 |
fix(pool): strip stale member names from settings on save
When a pool member provider is deleted/disabled, its name persists in the owner's extra_provider_names config. Validation already skips stale refs, but the names accumulated as garbage in the DB. Now stripStalePoolMembers() runs after validation passes, removing member names that don't exist in the active provider set before the settings are written to the database. |
||
|
|
e9733e08c4 |
fix(pool): improve pool management UX — clickable affordance, stale ref handling, managed-by banner (#671)
* fix(secure-cli): resolve ambiguous column in LookupByBinary JOIN query (#641) LookupByBinary uses LEFT JOIN with secure_cli_user_credentials but SELECT columns lacked table alias prefix, causing PostgreSQL error: "column reference 'id' is ambiguous (SQLSTATE 42702)" This silently broke ALL credentialed CLI exec — commands fell through to regular shell exec without injected env vars. Fix: use b.-prefixed column names for JOIN queries. Also add diagnostic logging to lookupCredentialedBinary for future debugging. * fix(agent): defer warning messages after parallel tool results (#644) When parallel tool calls trigger loop detection warnings, the warning messages (role="user") were inserted between tool result messages (role="tool"). This breaks the Anthropic API when routed through OpenAI-compatible proxies (e.g. LiteLLM): the proxy groups consecutive tool messages into a single user message with tool_result blocks, but an intervening user warning splits the group, causing orphaned tool_results and HTTP 400 "tool_use ids without tool_result blocks". Fix: accumulate warning messages during parallel result processing and append them after all tool results, preserving the consecutive grouping. Closes #642 * fix(docker): resolve @rollup/rollup-linux-arm64-musl missing on Alpine (#647) Added ui/web/.npmrc with supportedArchitectures for musl+glibc/arm64+x64. Updated Dockerfile to use --no-frozen-lockfile so pnpm fetches native rollup binding compatible with Alpine's musl libc. Lockfile still pinned by copy order. * docs(README): add history stars (#462) * fix(pool): skip stale pool member references during validation Unknown pool member references (deleted or disabled providers) now continue instead of returning an error. Prevents stale data from blocking provider saves. Closes #670 * fix(ui): redesign pool member selector and add managed-by banner Pool member selector: - Replace invisible outline button with custom element using dashed primary border, + icon badge, and "Click to add" hint text - Visible in both light and dark themes; hover transitions to solid border with shadow; active press scales down for tactile feedback Managed-by banner: - Show "Pool Defaults" section on pool members with info banner explaining which provider owns the pool, plus a Link navigation - Previously this section was completely hidden with no explanation i18n: add poolManagedByDescription and clickToAdd keys (en/vi/zh) * docs: add before/after UI evidence for PR #671 Annotated screenshots with red callout borders marking review areas. Self-contained HTML comparison report with dark/light theme toggle. * feat(ui): add pool discovery badges and setup wizard Replace verbose info banner with per-card "Pool available" badge on unpooled ChatGPT OAuth providers. Clicking the badge opens a new pool setup wizard dialog where users select owner, members, and strategy in one step. * docs: update UI evidence with pool discovery before/after * fix(ui): hide pool members from provider selector in agent forms Pool member providers are managed via the pool owner's routing config. Showing them as standalone options in the agent Provider dropdown is confusing — users may select a member directly instead of the owner, bypassing pool routing entirely. Filter out providers that exist in ownerByMember from the enabled providers list in ProviderModelSelect. * fix(ui): hide pool members from provider selector and add Pool badge Pool member providers are filtered out of the agent Provider dropdown in both the Create Agent dialog and the shared ProviderModelSelect component. Pool owners display a "Pool" badge so users know the provider routes to multiple accounts automatically. * docs: add provider selector before/after evidence * fix: revert stale merge in secure_cli.go and fix hardcoded i18n strings - Revert secureCLISelectColsAliased: b.agent_id → b.is_global (agent_id was dropped in migration 36, stale merge conflict artifact) - Replace hardcoded "Pool" badge text with t("providers:list.poolBadge") in provider-model-select and agent-identity-and-model-fields - Replace hardcoded "Disabled" with t("common:disabled") in pool wizard - Add list.poolBadge key to en/vi/zh locale files --------- Co-authored-by: Viet Tran <viettranx@gmail.com> Co-authored-by: Plateau Nguyen <nguyennlt.ncc@gmail.com> Co-authored-by: DNT <ducconit@gmail.com> |
||
|
|
156b2dd96c |
feat(secure-cli): per-agent grants with setting overrides
Replace agent_id column on secure_cli_binaries with is_global flag
and new secure_cli_agent_grants table for per-agent access control
with optional deny_args, deny_verbose, timeout_seconds, tips overrides.
- Migration 000036: create grants table, migrate agent-specific rows,
dedup binaries, drop agent_id, add is_global
- Store layer: SecureCLIAgentGrantStore interface + PG implementation,
LookupByBinary with LEFT JOIN grant merge, ListForAgent
- HTTP API: CRUD endpoints at /v1/cli-credentials/{id}/agent-grants
- Agent loop: buildCredentialCLIContext uses ListForAgent for scoped
system prompt (agents only see authorized CLIs)
- Web UI: grants dialog with card list + inline form, is_global toggle
replaces agent dropdown, i18n for en/vi/zh
|
||
|
|
2b1180f59d |
fix(http): allow localhost URLs for local provider types (Ollama, Claude CLI) (#674)
SSRF validation in validateProviderURL() blocked all localhost/loopback addresses, preventing local providers like Ollama from being configured with http://localhost:11434/v1. Introduce localProviderTypes map to skip SSRF checks for inherently local provider types (ollama, claude_cli, acp). Closes #673 |
||
|
|
c2bd5349e5 |
fix(ui,kg): fix credential user picker, KG filters, graph limits, and traversal depth
- Fix MCP/CLI credential dialogs: separate search text from selected user ID to prevent keystroke-as-selection bug (use onSelect pattern) - Fix Combobox: suppress value-sync after selection to prevent label flicker - Fix MCP dialog: prevent full-page spinner flash when switching users - Fix KG page filter: replace session-based scope with stats-based user_ids from KG entities (more reliable user list) - Increase KG graph visualization limit from 50 to 200 nodes with optimized D3 force layout breakpoints for 100+/150+ nodes - Cap graph API server-side limit at 500 to prevent unbounded queries - Increase KG traversal max_depth from 3 to 5 (HTTP API + agent tool) - Add dedicated Zustand store for KG detail dialog (isolated from main graph) - Add depth selector UI (2-5 hops) in entity detail dialog |
||
|
|
80c2c0e0ee |
fix(providers): normalize Ollama api_base URL to include /v1 suffix
- Write-time normalization via normalizeOllamaAPIBase() on create/update - Read-time safety net in resolveAPIBase() for pre-existing DB records - Covers both ProviderOllama and ProviderOllamaCloud - Stop hardcoding +"/v1" in registerInMemory/startup registration Closes #654 |
||
|
|
78a2152053 |
refactor(http): split provider_models.go into focused modules
- provider_models.go (154L): types, handler, reasoning helpers - provider_models_fetch.go (191L): remote API fetchers (Anthropic, Gemini, OpenAI, Ollama) - provider_models_catalog.go (103L): hardcoded model lists All files under 200-line guideline. No logic changes. |
||
|
|
e8579f5e1c |
feat(providers): add rich Ollama model listing via /api/tags (#659)
- Ollama/OllamaCloud providers now use native /api/tags for richer metadata - Display name built from family + parameter_size + quantization_level - Strips /v1 suffix, DockerLocalhost support, graceful fallback on error - 7 comprehensive tests covering all edge cases Closes #656 |
||
|
|
99720892a7 |
fix(storage): hide tenant roots from master storage API
Master tenant's base dir is the global data dir, which contains the tenants/ cross-tenant isolation root. Without this fix, the Storage page exposes tenant slugs, file names, and sizes. - Extract topLevelPath() helper, remove dead code - Add isHiddenPath() scoped to master tenant only - Guard handleList, handleRead, handleSize with hidden path checks - handleSize uses filepath.SkipDir to skip entire tenant subtree - Add regression tests for list, read, size, subpath, and non-master tenant isolation |
||
|
|
484d434f6c |
fix(security): harden sandbox, auth, and shell deny patterns
Sandbox: add noexec/nosuid/nodev to tmpfs mounts, remove SETUID/SETGID/CHOWN caps, add PidsLimit 256 default, keep no-new-privileges from base. Auth: reject X-GoClaw-User-Id header spoofing in dev mode (no gateway token), use full 32-byte HMAC for file tokens instead of truncated 16-byte. Shell: add NFKC Unicode normalization + zero-width character stripping before deny pattern matching, add 5 export-prefixed env var deny patterns, fix exemption logic to check per-argument prefix instead of whole-command substring (prevents bypass via comments while preserving skill store access). |
||
|
|
19d8faf28f | refactor(knowledge-graph): minor UI and handler adjustments | ||
|
|
10cd911dc7 |
feat(secure-cli): check-binary endpoint, agent select dropdown, fix ambiguous column
Backend: - Add POST /v1/cli-credentials/check-binary to resolve binary via exec.LookPath - Security: safeBinaryNameRe regex prevents filesystem probing - Fix ambiguous column in LookupByBinary LEFT JOIN (secureCLISelectColsAliased) - Add diagnostic logging to lookupCredentialedBinary Frontend: - Replace agent ID text input with Select dropdown (reuse useAgents hook) - Add Check Binary button next to binary name (auto-fills resolved path) - Add binary path hint explaining auto-detection from PATH - Update i18n keys (en/vi/zh) for new UI elements |
||
|
|
173f1cea53 |
fix(cli-credentials): show existing env keys on edit and support merge/remove (#629)
fix(cli-credentials): show existing env keys on edit and support merge/remove - Expose env_keys on secure CLI list/get and per-user credential list (values stay hidden) - PUT merges env with stored secrets: empty value keeps existing, removed keys dropped - Generic error messages to prevent env key enumeration - Audit logging for user credential set/delete - DRY: reuse envKeysFromDecryptedJSON helper - Client-side env key name validation (en/vi/zh i18n) |
||
|
|
899c8eb0b4 |
fix(skills): fix dep detection false positives and all-or-nothing install
- Filter Python stdlib modules at scan time to prevent false positives when the runtime checker fails (e.g. pip:argparse, pip:sys) - Install pip/npm packages one-by-one instead of batch so partial success is preserved when one package fails - Persist missing deps to DB after install via StoreMissingDeps() so reload reflects actual state instead of stale data - Use explicit master tenant context in handleInstallDeps for consistency with rescanAndUpdate() |