The bridge exposed a static BridgeToolNames subset that drifted from the
tool registry: use_skill, datetime, knowledge_graph_search and skill_manage
were never added, while the system prompt's skill-loading protocol requires
agents to call use_skill. claude_cli agents following the protocol hit a
nonexistent tool and could fabricate results.
Implement the structural fix recommended in #1373 triage:
- register the full bridge-capable surface (registry minus hard exclusions
spawn/create_forum_topic) instead of the static list
- gate BOTH tools/list (new WithToolFilter) and tools/call through one
shared predicate bridgeToolAllowed:
* callers WITH a verified agent policy get exactly the policy-filtered
surface (same WouldAllow check the call path always enforced)
* callers WITHOUT one (anonymous, or agent without tools_config) keep the
legacy conservative BridgeToolNames set - no exposure widening
- downgrade the per-call denial log Warn->Info; list filtering makes probes
of denied tools rare and the call gate is the intended enforcement point
Fixes#1373
Cron job execution set the tenant ID on its context but not the tenant
slug. Tenant-scoped filesystem paths (skills-store, workspace, media via
config.TenantScopedDir) key off the slug and fall back to an id-based
path when it is absent — a different directory than where HTTP/WS skill
upload materialized the files (which sets the slug). As a result a cron
agent turn in a non-master tenant saw NONE of its tenant's managed
skills: skill_search returned 0 results and the agent, unable to run the
skill, produced an ungrounded answer.
Add cronTenantContext() which resolves the tenant slug via TenantStore
and sets both WithTenantID and WithTenantSlug. Master tenant and
nil-store/lookup-failure paths fall back to id-only (prior behavior).
Thread TenantStore into makeCronJobHandler and runCommandCronJob.
Tested: added unit tests for cronTenantContext (slug injected for
non-master; master skips lookup; nil store and lookup error fall back to
id-only). Verified end-to-end on a live tenant: before, a daily-agenda
cron guessed an empty day; after, it read the real event from the DB.
Note: other background executors that build a context from a tenant ID
(e.g. heartbeat) likely share this gap and are worth an audit.
* fix: unify NO_REPLY detection (#1233)
Co-authored-by: GoClaw Operator <operator@goclaw>
* feat(acp): surface session/update tool_call notifications at Info level (#1141)
session/update notifications carrying a ToolCall (or inline tool_call /
tool_call_update kind) were only dumped at Debug level inside the params
blob. Operators could see `security.tool_granted` (permission granted) but
had no way to tell whether the tool actually executed successfully — both
"permission granted then failed silently" and "permission granted and
succeeded" looked identical in journalctl.
Add a structured Info log emitting toolCallId, name/title, status, and a
content preview (truncated at 400 chars) whenever the notification
contains tool-call state. This is what made it possible to diagnose the
recent .goclaw/-path-deny regression — `status=failed` immediately after
`security.tool_granted` revealed the gap that the granted-only log hid.
No behavior change beyond logging volume; preview is bounded.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(exec): exempt venv python interpreter from .goclaw/ path deny (#1140)
The ExecTool path-deny rule blocks any token containing `.goclaw/` unless
it matches one of the AllowPathExemptions prefixes (skills-store, tenants).
This silently rejected legitimate commands invoking the goclaw-managed
Python interpreter via its absolute path:
/home/user/.goclaw/venv/bin/python3 .../script.py
The first token `/home/user/.goclaw/venv/bin/python3` matched the deny
pattern but no exemption, so the entire command was denied.
Naive exemption (".goclaw/venv/bin/") does not work: matchesAnyPathExemption
resolves both tokens and exemption candidates via EvalSymlinks, and the
venv's python3 is a symlink into the host's python cellar (e.g. linuxbrew).
The token canonicalizes to /home/linuxbrew/.../python3.14 while a literal
".goclaw/venv/bin/" prefix never gets touched.
Fix: resolve venv/bin/python3 once at startup and exempt the dirname of
the resolved target. Failure to resolve (no venv present) silently falls
through.
Without this, ACP-driven agents either fail outright or work only via
fragile heuristics (cwd-local symlinks generated on the fly by the LLM).
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(cron): invert stateless gate to match UI label (#1139)
The cron job handler reset the session when stateless=false and skipped
reset when stateless=true — the opposite of what every UI locale labels
the field (en/ko/zh/vi all describe stateless as "each run starts fresh
without loading previous messages").
The buggy gate caused stateless=true crons to silently accumulate session
history across every execution, leading to context bloat and increasing
the chance of LLMs short-circuiting tool calls in favor of replaying
prior assistant turns. One affected daily ETL cron grew to 38 messages
over 18 days before the regression was noticed.
Fix: gate the Reset on `if job.Stateless` so the runtime matches the UI
contract. No DB migration is required — existing values were set by users
based on the UI label, so they already encode the intended behavior.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(providers): retry transient Codex response failures (#1332)
* fix(providers): retry transient Codex response failures
* test(cron): tolerate nil session store in handler tests
* fix(background): clear stale provider alerts (#1362)
Co-authored-by: Collective Developer <man@collective.dev>
---------
Co-authored-by: Duy /zuey/ <duy@wearetopgroup.com>
Co-authored-by: nguyenha935 <nguyenthanhha935@gmail.com>
Co-authored-by: GoClaw Operator <operator@goclaw>
Co-authored-by: codebit0 <34156842+codebit0@users.noreply.github.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Co-authored-by: Zezae Oh <zezaeoh@gmail.com>
Co-authored-by: Collective Developer <man@collective.dev>
* fix(mcp): return bare tool names with descriptions for already-connected MCP servers, fixing empty descriptions/schemas in agent prompts
* fix(mcp): resolve live tool descriptions/schemas in ListToolsForAgent for legacy prefixed tool_allow entries
The previous fix (d612da8e) only corrected the admin "browse/select tools"
endpoint. The actual runtime path building the LLM system prompt's tool list,
Manager.ListToolsForAgent, still looked up tool_cache/registry entries keyed
by the raw stored tool_allow name. Agents with grants captured before that
fix have tool_allow persisted as registered (prefixed) names instead of bare
MCP tool names, so every lookup missed and every tool in the prompt preview
came back with an empty description and schema.
Add bareMCPToolName to normalize both legacy-prefixed and current bare
tool_allow/tool_deny entries before matching, and resolveMCPToolInfo to
prefer the live registry's current description/schema (falling back to the
settings tool_cache) instead of trusting whatever was persisted verbatim.
Both the explicit-allow and cache-enumeration branches now share this single
resolution path.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Adds an operator opt-in env var (comma-separated hostnames/IPs) to
whitelist additional local hosts for ollama/acp provider-type URL
validation, on top of the existing hardcoded localhost allowlist.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
* feat(skills): allow editing SKILL.md content for database-managed skills
Adds the ability to edit a skill's actual SKILL.md content directly
from the web UI, restricted to database/tenant-managed skills only --
bundled/system skills remain read-only.
Backend: new PUT /v1/skills/{id}/files/{path...} endpoint
(handleWriteFile in skills_versions.go), writing to the current
version directory of a non-system skill only (403 for is_system
skills), with the same path-traversal/symlink/ownership guards as the
existing read endpoint. Bumps skill store version and emits
cache-invalidate + audit event on success.
Frontend: Content tab in the skill detail dialog gets an "Edit
content" button, disabled with a tooltip for system skills. Editing
fetches the RAW file content via the existing read endpoint (which
does not strip YAML frontmatter, unlike the stripped preview shown in
the Content tab by default) to avoid silently destroying frontmatter
on save. Save/cancel follow the established catch-and-toast error
pattern (no uncaught promise rejections, matching PR #1346's config
save button fix).
i18n keys added to all 4 locales (en, ko, vi, zh).
* fix(skills): bump version on content edit, fix scroll regression in content view
handleWriteFile previously overwrote the current skill version in place
instead of creating a new immutable version, inconsistent with the
skill_manage tool's patch action. Now creates a new version directory
and updates the DB pointer, matching that convention.
Also fixes a scroll regression in the skill detail dialog's Content
tab introduced by the edit-mode textarea -- both the read-only preview
and edit textarea are now properly scrollable within the dialog.
* fix(skills): refresh version display after save, fix dialog height constraint for scrolling
Version bump now correctly reflects in the UI immediately after save
without requiring a page reload. Fixed the actual root cause of the
scroll issue: the dialog's height wasn't bounded, so overflow-y-auto
on inner content had no effect since nothing constrained the dialog's
total height in the first place.
---------
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Browser pairing authenticates users as RoleOperator, but every provider
and OAuth route was wrapped in requireAuth(RoleAdmin). The /setup flow reads
GET /v1/providers, so paired browsers got 403 and hung on the setup page.
Split provider/OAuth route auth by operation:
- providers.go: GET list/get/models/status/embedding/codex-activity now use a
readAuth wrapper (method-derived min → Viewer); POST/PUT/DELETE and
reconnect/verify stay Admin. Responses already mask API keys and queries
stay tenant-scoped, so no secrets are exposed.
- oauth.go: GET status/quota use readAuth (they return only connection state
and quota, never tokens); start/callback/logout stay Admin (preserving the
#450 admin-only-management decision for mutations).
Updates the two auth tests to assert the new contract: an Operator can read
providers/OAuth status (200) but still cannot mutate them (403).
`goclaw migrate up` (and every other migrate subcommand) failed on Windows
with `create migrator: failed to open source, "file:///D:/.../migrations":
open .: The filename, directory name, or volume label syntax is incorrect.`
golang-migrate's file source driver mis-parses absolute drive-letter file://
URLs; the drive-letter formatting in absoluteToFileURI produced a URL the
driver could not open. Replace the file:// URL with an iofs source over
os.DirFS, which uses native OS path handling and works identically on every
platform. Drop the now-unused absoluteToFileURI/migrationsSourceURL helpers
and their URL-shape tests, and add a DB-free regression test that opens the
real migrations directory via newMigrationSource().
Verified end-to-end on Windows: `migrate version` and `migrate up` now run.
Closes#1076, #1338.
Fix A — tenant backup aborted with SQLSTATE 42703 "column id does not
exist" because exportQuery() hardcoded ORDER BY id. Several tenant-scoped
tables have composite PKs and no id column. Add a TableDef.OrderBy field,
honor it in exportQuery(), and set it for every id-less registry table:
config_secrets, agent_team_members, tenant_hook_budget, system_configs,
builtin_tool_tenant_configs, skill_tenant_configs, user_agent_profiles.
Fix B — hooks, tenant_hook_budget and webhook config were missing from
the backup registry, silently dropping their data on backup/restore. Add
hooks, tenant_hook_budget, webhooks (preserve) plus the hook_agents junction
(via ParentJoin through hooks, composite PK), and mark hook_executions and
webhook_calls as ephemeral in the skipped list.
Fix C — restoring on a fresh server failed with "N active DB connection(s)
detected" because the gateway's own pool connections were counted as active
clients. Tag pool connections with application_name='goclaw' (pg.OpenDB) and
exclude them in CheckActiveConnections; genuine external clients still block.
Adds unit + integration regression tests, including an export-over-every-
registered-table test that surfaced the additional id-less tables.
Deleting a team failed with SQLSTATE 23514 whenever it owned a
team-scoped vault document. The vault_documents.team_id FK is
ON DELETE SET NULL; when team_id became NULL the
vault_docs_team_null_scope_fix() trigger (migration 000043)
unconditionally set scope='personal'. Team docs have agent_id IS NULL,
so the resulting (personal, agent_id NULL) row violates the
vault_documents_scope_consistency CHECK and aborts the whole delete.
PostgreSQL: migration 000089 replaces the trigger function to pick a
valid target scope by ownership, and only rewrite genuinely team-scoped
rows so scope='custom' docs that merely carry a team_id are preserved:
agent_id IS NOT NULL -> 'personal' (defensive: legacy dirty rows)
agent_id IS NULL -> 'shared' (normal team docs)
SQLite: the same class of bug exists (FK SET NULL leaves scope='team'
with a NULL team_id and aborts on the CHECK), but SQLite fires no
trigger during the FK action, so a DB-level fix is impossible.
SQLiteTeamStore.DeleteTeam now converts team-scoped docs in the same
transaction before removing the team, scoped by tenant_id in the
tenant path to keep tenant isolation.
Adds regression tests on both engines and bumps RequiredSchemaVersion
to 89. SQLite needs no schema migration since the fix is in Go code.
Ollama has two separate request-building code paths: OllamaProvider
(native /api/chat, used for num_ctx control) and OpenAIProvider
(OpenAI-compat /v1/chat/completions). An earlier fix disabled thinking
mode by hardcoding think=false, but only in the OpenAI-compat path --
OllamaProvider.buildRequest() never set the think field at all, so
reasoning-capable models (qwq, deepseek-r1) defaulted to visible
chain-of-thought reasoning regardless of that fix. Confirmed live via
a docker-engineer agent streaming full reasoning traces despite the
existing disable.
Replaced the hardcoded always-off behavior with a provider-level
tri-state setting (llm_providers.settings.thinking_enabled: unset =
default off, explicit true/false overrides), configurable via the
provider's Advanced settings dialog. Both OllamaProvider.buildRequest()
and OpenAIProvider.buildRequestBody() now read and respect this same
setting, so the toggle works regardless of which Ollama code path a
given deployment routes through.
Added tests for setting parsing (unset/true/false/malformed) and both
provider request-builders' handling of the override.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
fix(zalo): correctly forward inbound reference images to the agent
Fixes#1352.
Three compounding bugs prevented zalo_personal from using photos as vision/image-generation references:
1. Zalo CDN URLs often lack usable extensions, so downloads landed as .bin and were misclassified as documents (not images). fixDownloadedExt sniffs content and renames correctly.
2. Photo captions (user's actual instruction) were dropped — now appended after the media tag.
3. Quoted photos were never downloaded — extractQuoteMedia + ParseAttachment now fetch the actual image.
Tests cover fixDownloadedExt and ParseAttachment helpers. Live-verified against production Zalo traffic.
Codex native image_generation output (state.Observe.AssistantImages) was
persisted to workspace/media/ and attached to the assistant message's
MediaRefs for session history, but never added to state.Tool.MediaResults
— the only field that flows into RunResult.Media and from there into the
outbound channel message. Generated images showed up in the web UI (which
reads session history) but chat channels (Telegram, Zalo, Slack, ...)
received text-only replies with no attachment.
Append the persisted image refs to state.Tool.MediaResults right after
they're attached to the assistant message, so they reach RunResult.Media
via the existing pipeline plumbing and cmd/gateway_consumer_normal.go's
appendMediaToOutbound — the same path create_image/send_file already use.
No new delivery path needed.
bufio.Scanner has a fixed max-token size (SSEScanBufMax, 1MB). Codex's
native image_generation streams a full base64 image in a SINGLE SSE
data line (response.image_generation_call.partial_image and the final
image_generation_call item), which regularly exceeds 1MB for anything
larger than a thumbnail. Scanner.Scan() then fails with:
stream read error: bufio.Scanner: token too long
killing the whole turn after the model already generated the image.
Replace bufio.Scanner with bufio.Reader.ReadString, which grows to fit
a line of any length instead of erroring past a fixed cap. SSEScanBufMax
is now unused and removed; SSEScanBufInit is reused as the reader's
initial buffer size.
maybeSummarize spawns a background goroutine that logs via slog.Info
before calling the LLM provider -- correct, intentional async design,
not a bug. The test read the shared log buffer (buf.String(), inside
captureSlog) immediately after maybeSummarize returned, i.e. right
after the goroutine was merely spawned, not after it finished logging
-- an unsynchronized concurrent read/write on bytes.Buffer, caught by
go test -race. This was surfaced by an unrelated PR's CI run
(nextlevelbuilder/goclaw#1346), confirmed unrelated to that PR's
actual changes.
Moved the existing done-channel wait inside captureSlog's closure, so
the buffer read only happens after the background goroutine has
signalled it reached provider.Chat -- which happens strictly after
the slog.Info write on this path. Proper happens-before via the
channel, no sleep/poll, no production code changed.
Verified 10/10 passes under go test -race -count=10, plus full
package regression run clean.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Surface cache-read and cache-creation input tokens (plus the
prompt_tokens_include_cached_segments flag) in webhook usage responses,
so callers can distinguish cached from non-cached input tokens.
The data already flowed end-to-end via providers.Usage and
agent.RunResult.Usage; the two webhook envelope structs (webhookLLMUsage
for the sync LLM webhook, callbackUsage for the async callback payload)
copied only 3 of the fields, dropping the cache data. Add the 3 cache
fields (mirroring providers.Usage JSON tags exactly, all omitempty) and
populate them. The async callbackUsage mapping is extracted into a pure
newCallbackUsage helper for unit testing.
Additive and non-breaking: omitempty keeps responses without caching
byte-identical. No schema change.
The gateway binary (/app/goclaw) has embedded operator/admin subcommands
(agent list, doctor, skills list, sessions list, providers list,
channels list, cron list, traces list, etc) that the bundled goclaw
skill teaches agents to use for self-administering the gateway. But
/app was never added to PATH -- confirmed live via docker exec:
`which goclaw` empty, while `/app/goclaw --help` works fine directly.
Non-sandboxed agent exec runs in the same container as the gateway
itself, so /app/goclaw is a real, reachable file for those agents --
the only blocker was PATH resolution for the bare `goclaw` command.
Prepend /app to PATH in docker-entrypoint.sh, ahead of the existing
runtime-installed package directories, so bare `goclaw` commands
resolve correctly for both operators (docker exec) and agents (exec
tool) alike.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Start() bound the shared *rod.Browser to a context built with
context.WithTimeout(ctx, 15s) and a deferred cancel, so the browser's
context was canceled the moment Start() returned. Any later reuse of the
shared handle then failed immediately with "context canceled" — most
visibly m.browser.Incognito() in tenantBrowserLocked, so a successful
"start" action was followed by a failing tab creation for scoped
(tenant|user|agent) agent sessions. The getPage path masked this via its
auto-reconnect (reconnectLocked uses a background context), but the
incognito path has no such recovery.
Connect with the default background context, the same way reconnectLocked
already does. The local launcher path stays bounded by launchCtx and
remote reachability by resolveRemoteCDP, so no meaningful connect-time
guard is lost.
buildRequestBody emitted prompt_cache_key and prompt_cache_retention for
every Codex request. The ChatGPT subscription OAuth backend
(chatgpt.com/backend-api) rejects prompt_cache_retention with HTTP 400
"Unsupported parameter", so every LLM call from an agent using a ChatGPT
OAuth Codex provider failed at the first think step.
Gate both params on isOpenAINativeEndpoint(apiBase), the same native-only
policy already documented and applied by CacheMiddleware. This covers both
CodexProvider and CodexAdapter, since the adapter delegates to the same
builder. Server-side prefix caching still works on the OAuth backend
without these params.
Update the two prompt-cache tests to assert both cases: native endpoints
include the params, the OAuth backend omits them.
prefilter() had zero logging -- matcher regex failures, tool-name
mismatches, CEL compile errors, and CEL evaluating to false were all
completely silent, making a misconfigured hook impossible to debug.
A user's PreToolUse script hook silently failed to block an
unauthorized file write with no way to determine why.
- Added slog.Debug/Warn logging throughout prefilter() and the
dispatch chain: hooks.prefilter.considered, matcher_checked,
condition_evaluated, error, and hooks.dispatch.decision -- covering
matcher match/miss, CEL compile/eval result, and final decision
with reasoning.
- console.log/console.error output from script hooks (captured via
the existing goja sandbox stdout buffer, previously discarded after
execution) now surfaces via slog.Debug and persists into
HookExecution.ConsoleOutput / Metadata["console_output"] in the
audit record, using the existing metadata JSON column (no
migration needed).
- Fail-closed, narrowly scoped: when a hook's matcher successfully
matches a tool call but CEL condition or script execution then
errors, the dispatcher now blocks that specific tool call
(previously: silently treated as non-match, call proceeded
allowed). Hooks that legitimately don't match remain unaffected
pass-through -- this does not turn one broken hook into a global
outage, it only blocks the specific gated action that couldn't be
confidently evaluated.
Added tests: matched-then-CEL-errors blocks the call and never
invokes the handler; matcher-doesn't-match proceeds normally
unaffected (confirms narrow scope); console.log output is captured
in the audit record.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
* fix(prompt): remove duplicate Team Members section from system prompt
The TEAM.md context file already provides the Members section with better
formatting. The system prompt section was redundant and inconsistently
formatted.
- Remove buildTeamMembersSection() call from system prompt building
- Remove unused buildTeamMembersSection() function
* fix(mcp): wire store to manager for prompt preview tool visibility
MCP manager needs database access to query configured servers for tool visibility
in prompt preview.
- Add SetStore() method to MCP manager
- Wire pgStores.MCP to manager after initialization
- Add debug logging for MCP initialization flow
* fix(systemprompt): hide Tooling section when agent has no tools
The Tooling section header and boilerplate were displayed even when the
agent had no tools available. Add early return in buildToolingSection()
to skip the section entirely when toolNames is empty.
This reduces prompt noise for agents with no tool access.
* feat(mcp): cache tool descriptions for prompt preview visibility
MCP tools now show descriptions in prompt preview without requiring a live
server connection.
- Add CacheToolDescriptions() method to MCPServerStore (PG + SQLite)
- Cache tool descriptions in settings['tool_cache'] when server connects
- Use cached descriptions in ListToolsForAgent (fallback: hints → cache → global)
- Descriptions are auto-populated from live server manifest on connection
- Admins can still override via tool_hints in server settings
* test(mcp): add CacheToolDescriptions to MCPServerStore test fakes
Commit 5ce410b6 added CacheToolDescriptions() to the MCPServerStore
interface but missed updating test mock implementations, breaking
go vet across internal/mcp, internal/agent, internal/http, and
internal/channels/bitrix24.
Add no-op implementations matching each fake's existing style.
* fix(providers): enforce per-agent tool policy for Claude CLI provider
The Claude CLI provider (stdio+MCP bridge) was not enforcing per-agent
tool policy, unlike other providers where the policy-filtered tool list
already drives the system prompt's Tooling section.
Two gaps closed:
1. --disallowedTools was previously skipped entirely when no MCP config
path was resolved, letting the CLI subprocess run with its full
native toolset (Bash, Edit, Read, Write, Glob, Grep, WebFetch,
WebSearch) regardless of agent policy. It's now unconditional and
derived from the agent's actual allowed-tools list (state.Tool.AllowedTools),
mapped to Claude CLI's native tool names.
2. The MCP bridge server executed any tool call without checking the
calling agent's policy. It now resolves the agent's policy from
context (via the existing HMAC-verified agent lookup) and denies
calls to tools outside that agent's allowed set, logging
security.mcp_bridge_denied on denial.
Both gaps were closed using existing plumbing (PolicyEngine.WouldAllow,
AgentData.ParseToolsConfig, the bridge context middleware) — no new
cross-cutting mechanism was introduced.
* feat(web): show MCP/tool schemas in system prompt preview dialog
The prompt-preview API response includes a separate `tools` field
(the actual JSON schemas sent to the LLM as the tools API parameter)
alongside `prompt` (the system prompt text), but the web UI only
rendered `prompt`, silently dropping the tools list.
Add a collapsible Tools section to both the full-screen System Prompt
dialog and the inline agent-detail preview, showing tool count, name,
description, and expandable parameter schema per tool. i18n keys added
to en/vi/zh locales.
* fix(prompt): render pinned skills on bootstrap turns
Pinned skills are documented (web UI copy) as "always inlined in the
system prompt", but the entire Skills section was gated behind
!cfg.IsBootstrap, so pinned skill XML never appeared on bootstrap
turns (first message of a session) despite the promise.
Separate pinned-skill rendering from bootstrap-suppressed guidance:
- Bootstrap + pinned skills present: render pinned XML only, no
search/manage guidance (which stays suppressed as before)
- Non-bootstrap: unchanged behavior
- Minimal/none modes: pinned skills always render regardless of
bootstrap state
Add regression tests covering all four prompt modes on bootstrap
turns, plus a non-bootstrap guard confirming existing behavior is
preserved.
* fix(skills): resolve managed skills directory per-tenant, not master-only
skills.Loader was wired at startup to scan a single fixed directory
(the master tenant's managed-skills dir), making any skill belonging
to a non-master tenant invisible to both pinned-skills prompt
resolution and skill_search/use_skill, regardless of DB visibility
settings.
- Loader now resolves the calling tenant's managed-skills directory
per-call via context (store.TenantIDFromContext), never enumerating
other tenants' directories
- Skill cache is now tenant-keyed to prevent slug collisions and
cross-tenant cache leaks across tenants using the same skill slug
- gateway_setup.go passes the root data dir instead of a pre-resolved
master-tenant path
Write-side tooling (skill_manage, publish_skill) was already correctly
tenant-scoped per-operation — no changes needed there.
Added TestLoader_ManagedSkills_TenantIsolation proving two tenants
with same-slug/different-content skills never see each other's
content, including after cache population from a different tenant's
lookup.
Known follow-up (not in this commit): skill_search's BM25 index is
still a single process-global index shared across tenants, which is
a related but separate cross-tenant search-result leak requiring its
own scoped fix (per-tenant index maps + threading tenant context
through ensureIndex/rebuildIndex).
* fix(tools): scope skill_search BM25 index per-tenant
SkillSearchTool held a single process-global BM25 index built once
from whichever tenant's context first triggered ensureIndex, then
reused for all subsequent Execute() calls regardless of caller —
leaking one tenant's skill search results into another's, the
search-path counterpart to the managed-directory bug fixed in
7b4668ad.
- index/lastVersion are now keyed per-tenant (map[uuid.UUID]*tenantIndexState)
- ensureIndex resolves the calling tenant from context and only
builds/reads that tenant's index entry, never touching another
tenant's cached state
- Builtin/bundled skills remain visible in every tenant's index
(Loader already merges those tiers correctly per 7b4668ad)
Loader.Version() remains a single global counter — a version bump in
one tenant causes unnecessary rebuilds in others but does not cause
cross-tenant leakage, an acceptable tradeoff to avoid scope creep.
Added TestSkillSearchTool_TenantIsolation proving two tenants with
same-slug/different-content skills never see each other's search
results, including after cache population from a different tenant.
* fix(tools): fix group-spec expansion in tool policy engine
PolicyEngine.registry was only ever set via SetRegistry(), which was
never called in production (only in one test) — so pe.registry was
permanently nil in production. Every group-expansion helper
(applyProfile, intersectWithSpec, unionWithSpec, subtractSpec,
expandSpec, matchDenySpec, filterByCapability) silently dropped any
"group:*" spec entry instead of expanding it when registry was nil.
Concretely: "group:mcp" (auto-injected into agentToolPolicy.AlsoAllow
for any agent with MCP tools) never resolved to real tool names, so
MCP tools connected successfully and appeared in prompt text (which
reads the registry directly, bypassing PolicyEngine) but were never
included in the actual ChatRequest.Tools payload sent to the LLM —
confirmed live via mcp.agent.tools_loaded tools=6 immediately followed
by mcp.filtered_tools mcp_defs_count=0 in the same request. This
affects any agent relying on group-based grants, not just MCP.
PolicyEngine is a shared/global singleton used concurrently across
all agents (constructed once at gateway startup), so mutating a
registry field per-call would be a data race. Fix instead threads the
registry as an explicit parameter from FilterTools down through all
internal group-expansion helpers, and adds IsDenied/WouldAllow
registry parameters, removing the dead SetRegistry() mechanism
entirely.
Also fixes group expansion for the per-user-MCP-tools path: FilterTools
is sometimes called with a userToolOverlay wrapping a *Registry rather
than a *Registry directly; added Unwrap() to userToolOverlay so the
new registry-resolution logic works for both cases.
Added 4 tests proving group expansion works via the threaded parameter
alone (no SetRegistry): plain registry allow, userToolOverlay allow,
deny-side group expansion, and the WouldAllow bridge-server path.
Blast radius note: this restores intended access for every agent
configured with group:* specs (group:mcp, group:vault, group:goclaw,
group:coding, etc.) that were silently inert before. Existing agent
configs relying on group grants will gain the tool access they were
nominally already configured for.
* fix(mcp): cache tool descriptions from pool-connected servers too
5ce410b6 added tool-description caching (for prompt-preview visibility)
only inside connectServer. connectViaPool — the separate connect path
used when MCP connections go through the shared pool — never got the
same caching hook, even though it shares the same underlying
connectAndDiscover wire handshake.
Confirmed live: cloudflare/docker connected via connectViaPool and
received real descriptions over the wire, but prompt preview still
showed blank descriptions because this path never wrote to the cache.
Also removes the temporary mcp.connect.raw_tool debug log added
earlier this session for diagnosing the same issue — no longer needed
now that the root cause is fixed.
* chore(skills): remove temporary pinned-skills diagnostic logging
Confirmed live: pinned skills (caveman, infra-ansible-knowledge) now
resolve correctly end-to-end for tenant-scoped agents. Debug logging
added to trace the resolution chain is no longer needed.
* fix(tools): deny always wins over AlsoAllow group grants
AlsoAllow's unionWithSpec could reintroduce a tool explicitly listed
in Deny, since it added tools back from allTools without re-checking
deny specs. Previously masked because AlsoAllow's group-expansion was
also broken (fixed in f7af95de this session) — group specs silently
expanded to nothing, so this ordering bug never manifested. Now that
group expansion works, an admin-denied tool that's also reachable via
a group:* AlsoAllow entry (e.g. group:mcp) would silently reappear.
Re-apply deny-spec subtraction as a final step after AlsoAllow union,
for both global and per-agent policy, so deny always wins regardless
of which allow mechanism tries to add a tool back.
Added tests proving global and per-agent Deny correctly override an
overlapping AlsoAllow group grant, while sibling non-denied tools in
the same group remain allowed.
* fix(mcp): enumerate cached tools instead of wildcard placeholder in prompt preview
ListToolsForAgent (the prompt-preview path) collapsed any server with
an empty ToolAllow (unrestricted grant — the common case) into a
single "server__*" placeholder entry, even when tool_cache already
had every real tool name and description from connect time (5ce410b6,
8ffd67b9). This meant agents with unrestricted MCP server access never
saw individual tool names or descriptions in prompt preview, forcing
trial-and-error tool usage.
When ToolAllow is empty and tool_cache is populated, enumerate every
cached tool (skipping any explicitly denied) and emit one
MCPToolPreviewInfo per tool, matching the construction logic already
used for the ToolAllow-non-empty case. Falls back to the single
placeholder only when tool_cache is also empty (server never
connected).
The live (non-preview) conversation path, buildMCPToolDescs, does not
have this bug — it resolves tool identity from the live connected
registry, never from ToolAllow, so no placeholder shortcut exists
there.
Added tests covering both the cache-populated enumeration case and
the no-cache placeholder fallback.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(prompt): filter alias re-injection and add MCP schemas to preview ToolDefs
Prompt-preview's tools: schema array (PreviewResult.ToolDefs) had two
bugs, both preview-only — confirmed live conversations already use
correctly-filtered tool payloads via buildToolsPayload/PolicyEngine.FilterTools,
unaffected by either bug:
1. Alias re-injection iterated ALL registry aliases globally with no
check against the deny-filtered toolNames list, letting denied
tools reappear via their alias name (e.g. a denied canonical tool
still showing up as its Claude-Code-compat alias like Bash/Edit/Write/Read).
Now skips any alias whose canonical tool isn't in the filtered set.
2. MCP tool descriptions (from the store-based, connection-free
ListToolsForAgent path) only ever fed the prompt TEXT section,
never got converted into ToolDefinition schema objects — so the
tools: array never showed MCP tools at all, even when the MCP
section of the prompt text correctly listed them. Now appends a
ToolDefinition per MCP tool with name+description populated and a
placeholder {"type":"object"} Parameters schema, documented as
preview-only (real parameter schemas require a live MCP connection,
only available during an actual conversation turn).
Added tests proving denied-tool-alias exclusion and MCP tool inclusion
in preview ToolDefs.
* feat(skills): inline full content for pinned skills instead of pointer-only
Web UI documented "Pinned skills are always inlined in the system
prompt", but BuildPinnedSummary was just BuildSummary with an
allowlist filter — same as the general searchable skill list:
name/description(truncated)/location pointer only, requiring
use_skill+read_file round trips to get actual content. No code path
inlined real SKILL.md content for pinned skills specifically.
BuildPinnedSummary now reads and inlines full SKILL.md content
(frontmatter stripped) per pinned skill inside <skill_instructions>
tags. Per-skill (10000 bytes) and total (30000 bytes) size caps fall
back to the original pointer-only format with a note when a skill is
too large to inline, so oversized skills degrade gracefully instead
of blowing the prompt budget.
Added tests: full-content inlining, size-cap fallback, and tenant
isolation for the new inline path (mirroring the existing managed-
skills tenant isolation test).
* fix(mcp): real parameter schemas and full policy enforcement in prompt preview
Three interconnected fixes to prompt preview, none affecting the live
conversation path (which was already correct):
1. Real MCP parameter schemas instead of a useless empty placeholder.
tool_cache previously stored only name+description; extended to
also capture the real JSON Schema from the MCP server's tools/list
response (CachedToolInfo{Description, Parameters}) at connect time,
for both direct-connect and pool-connect paths. Preview now shows
complete, real input schemas instead of {"type":"object"} with no
properties — verified against actual wire data from a live
cloudflare MCP server showing genuinely rich schemas (zone_id,
type, name, content, ttl, proxied, all typed with descriptions and
correct required arrays) that were previously being discarded.
Backward-compatible: old-shape cache entries degrade gracefully to
description-only rather than crashing, self-healing on next
connect.
2. Global tool deny now enforced in preview. BuildPreviewPrompt
previously hand-rolled a partial policy reimplementation
(per-agent deny only, explicitly skipping the full PolicyEngine
"because runtime state isn't available in preview") — but
PolicyEngine.WouldAllow already handles this per-tool-name without
needing channel context. A tool denied via the global config (not
per-agent) would appear in preview despite being correctly denied
in every real conversation. Preview now calls WouldAllow per
candidate tool, with a graceful per-agent-deny-only fallback when
no PolicyEngine is wired (e.g. in tests).
3. MCP tools now also subject to the same policy check — previously
the MCP tool supplement (store-based, connection-free tool listing)
added MCP tool names unconditionally, bypassing WouldAllow entirely,
so a denied MCP tool could still appear in preview.
Added tests for all three: real-schema presence, global-deny exclusion
for both core and MCP tools, and backward-compat cache handling.
* fix(prompt): remove redundant per-tool MCP enumeration from prompt text
Now that MCP tool schemas in the tools: API parameter are real and
complete (1290d4f1), the ## MCP Tools prompt-text section's per-tool
"- mcp_x__y: description" enumeration is pure duplication with zero
added value — the model already gets each tool's real schema
(including description) via tools:.
buildMCPToolsInlineSection now keeps only the behavioral instructions
that aren't expressible via JSON schema and thus aren't duplicated:
prefer-MCP-over-core-tools guidance, and the optional-parameter
guidance (don't guess/fill optional fields). The per-tool name+
description enumeration loop is removed. Section still only appears
when the agent has MCP tools (len(cfg.MCPToolDescs) > 0, unchanged
gate).
Updated tests to assert the enumeration is gone while the behavioral
instructions remain; ToolDefs assertions are now the authoritative
check for MCP tool allow/deny filtering behavior (prompt text no
longer enumerates names at all).
* test(agent): update TeamContextInjection test for removed Team Members section
TestBuildSystemPrompt_TeamContextInjection asserted the presence of a
'Team Members' prompt-text section that was intentionally removed in
71d33180 (duplicate of the canonical TEAM.md-context-file Members
section, which has better formatting). The test was never updated to
match, causing it to fail on every run since. Moved the assertion
from wantIn to wantNotIn for the 3 affected subtests -- BuildSystemPrompt
correctly no longer renders team-member roster info directly; that
info now comes exclusively from the TEAM.md context-file mechanism,
outside this test's isolated scope.
* fix(prompt): resolve real registry for WouldAllow calls in preview
BuildPreviewPrompt's two WouldAllow calls hardcoded reg=nil, silently
breaking group:* expansion (e.g. group:mcp) needed to resolve the
AlsoAllow grant production actually uses to grant MCP tool access
(resolver_helpers.go's agentToolPolicyWithMCP injects
AlsoAllow: ["group:mcp"]). With reg=nil, WouldAllow could match
literal tool names fine (the 18 core/static tools) but could never
resolve group-based grants, so every MCP tool silently failed
WouldAllow and was excluded from preview -- confirmed live via curl:
18 tools returned, zero mcp_* ones, for an agent with genuinely
working MCP access in real conversations.
Live conversations were never affected -- internal/mcp/bridge_server.go's
WouldAllow call already correctly passes a real registry.
Fix resolves a real *tools.Registry from deps.ToolLister via
tools.ResolveConcreteRegistry (the same helper used at the live call
site), passing it to both WouldAllow calls instead of nil. Falls back
to nil gracefully for test mocks that don't implement the full
ToolExecutor interface, preserving existing test behavior.
Added a test proving an MCP tool granted via the exact production
AlsoAllow: ["group:mcp"] pattern is now correctly included in preview
ToolDefs, where the old reg=nil bug would have silently excluded it.
* fix(prompt): use literal deny check for MCP tools in preview, not group expansion
The MCP-tools policy gate added in 1290d4f1 called WouldAllow with a
real registry (per 9b5fd2eb), which requires group:mcp expansion
against that registry to grant access via the production
AlsoAllow: ["group:mcp"] pattern. But MCP tools are only ever
registered into ephemeral per-agent registry clones at live connection
time (manager_connect.go) -- never into the shared/global registry
preview uses. group:mcp always resolved empty in preview's
connection-free context, so WouldAllow denied every MCP tool --
confirmed live via tool-name diff: live conversations correctly
included all 6 MCP tools, preview included zero.
MCP access-granting is already correctly handled by
ListToolsForAgent's own per-server tool_allow/tool_deny grant logic
(confirmed working correctly earlier this session). The preview gate
only needs to catch the narrower case of a literally-denied tool name
via global/per-agent policy config -- it never needed group
expansion. Replaced WouldAllow with IsDenied(nil, name, agentPolicy),
which forces a pure literal-name match with zero registry dependency,
matching the existing usage pattern already established elsewhere in
policy.go. This class of bug cannot recur: there's no registry-passing
code path left in this check to silently reintroduce group-expansion
dependence.
Added a test proving MCP tool inclusion in preview is independent of
group-expansion outcome (no AlsoAllow: group:mcp needed for a
non-denied tool to appear).
Also reverts the temporary loop.filtered_tool_names/
preview_prompt.filtered_tool_names diagnostic logging used to capture
the live-vs-preview tool-name comparison that diagnosed this bug.
* fix(http): preserve real MCP parameter schemas through HTTP preview adapter
mcpPreviewAdapter.ListToolsForAgent (the HTTP-layer glue converting
mcp.MCPToolPreviewInfo to agent.MCPToolPreviewInfo for BuildPreviewPrompt)
only copied RegisteredName and Description, silently dropping
Parameters -- a bug present since this adapter was introduced
(2499d0be/7e250244), unrelated to today's other MCP preview fixes.
This was masked until d5fc6344 fixed MCP tools being excluded from
preview entirely (a separate bug) -- once MCP tools started appearing
again, this pre-existing adapter gap became visible: tools showed up
correctly, but always with the bare {"type":"object"} placeholder
instead of their real cached schema (confirmed live: update_dns_record
missing its 7 real properties).
One-line fix: copy Parameters through in the adapter's struct literal.
Added a regression test constructing a real *mcp.Manager with
populated tool_cache, asserting the adapter's output preserves
specific real schema properties (not just non-nil Parameters) --
verified this test fails without the fix and passes with it.
---------
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
fix(usage): repair cost analytics and display precision
- Automatic OpenRouter pricing sync with cost backfill for traces/snapshots/events
- Live usage data merging for current-hour dashboard accuracy
- 2-decimal API cost formatting across usage/overview pages
- Comprehensive test coverage (PG, SQLite, HTTP, UI)
Merged by github-maintain automation.
When an agent is deleted, usage_snapshots rows are set to agent_id=NULL via FK ON DELETE SET NULL.
If aggregate rows already exist for the same (bucket_hour, provider, model, channel, tenant_id),
the unique index collision causes a constraint error.
Pre-aggregate the agent's usage into NULL-agent rows before deletion, eliminating the collision.
Token counts are summed into existing aggregate rows or new rows are inserted.
Fixes: duplicate key value violates unique constraint "idx_usage_snapshots_unique"
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
* fix(mcp): show configured MCP tools in prompt preview
MCP tools were only visible in prompts if currently loaded in registry
(from active sessions). Prompt preview showed empty MCP tools section.
Solution: Query MCP store for configured servers and tool lists.
Show tools that would be available, not just currently loaded.
- internal/mcp/manager.go: add ListToolsForAgent() for store-based tool discovery
- internal/agent/preview_prompt.go: supplement live registry with store tools
- internal/http/agents.go: add mcpPreviewMgr field and setter
- internal/http/agents_prompt_preview.go: create adapter bridging mcp.Manager to preview
- cmd/gateway.go: wire up MCPPreviewAdapter during setup
* fix(mcp): load MCP servers from database at startup, not config file
Single source of truth: MCP servers are now loaded from mcp_servers table
at gateway startup, instead of from the config file. This ensures that
MCP servers configured via web UI are automatically available without
requiring config file updates.
- cmd/gateway_setup.go: remove config fallback, add initMCPFromDB() helper
- cmd/gateway.go: call initMCPFromDB() after store initialization
- internal/mcp/manager.go: add SetConfigs() for runtime configuration
Now mcpMgr will be non-nil when MCP servers are in the database,
enabling MCP tools in prompt preview.
---------
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Root cause: internal/agent/loop_context.go re-applied tenant scoping
(TenantLayer/TenantID) on top of l.dataDir, which was already
tenant-scoped upstream in internal/agent/resolver.go, producing paths
like /app/workspace/tenants/<slug>/tenants/<slug>/teams/<teamID>
instead of the correct /app/workspace/tenants/<slug>/teams/<teamID>.
Also consolidates all previously-duplicated tenant-path-joining logic
(config.TenantDataDir, config.TenantWorkspace, tools.TenantLayer,
workspace.tenantPath) into a single canonical config.TenantScopedDir
function in internal/config/tenant_paths.go, with all other
implementations now delegating to it.
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>
Give agents direct access to the CRM entity, Openline session, task,
and workgroup binding of an inbound chat without paying an extra RPC
per turn. Everything is derived from CHAT_ENTITY_* fields Bitrix
already ships in the ONIMBOTMESSAGEADD webhook.
New metadata keys emitted alongside the existing bitrix_chat_entity_*
pair (only present when non-empty, so DMs / plain groups pay no cost):
Every chat with a title/type:
bitrix_chat_title, bitrix_chat_type
Openline (ChatEntityType=LINES), decoded from ENTITY_ID + DATA_1:
bitrix_ol_connector_code (facebook / synity_zalo_oa_chat / ...)
bitrix_ol_line_id
bitrix_ol_external_uid
bitrix_ol_bitrix_user_id
bitrix_ol_line_config_id
bitrix_ol_session_started_at
bitrix_ol_active_crm_type / bitrix_ol_active_crm_id
CRM linkage, decoded from CHAT_ENTITY_DATA_2 for Openline
(schema "LEAD|<id>|COMPANY|<id>|CONTACT|<id>|DEAL|<id>") or from
ENTITY_ID for native CRM-integrated chats ("CONTACT|1780"):
bitrix_crm_lead_id / company_id / contact_id / deal_id
bitrix_crm_type + bitrix_crm_id
Single-token ENTITY_ID cases:
bitrix_task_id (TASKS_TASK)
bitrix_workgroup_id (SONET_GROUP)
bitrix_mail_id (MAIL)
Parsing is best-effort: malformed input (short ENTITY_ID, odd DATA_1
token count) logs a WARN and falls through to a partial map rather
than crashing the handler. Unknown ChatEntityType values pass raw
without noise so future surfaces (CALENDAR, LIVECHAT, ...) don't spam
logs before the parser knows how to decode them.
Verified live against tamgiac.bitrix24.com with three connectors
(Facebook, Zalo OA, Zalo personal) plus CRM/Task/Workgroup/Mail
sample events. Table-driven test covers 13 shapes including malformed
input and the NONE/0 edge cases.
Surface parity:
- Gateway server: parser + wire in handle.go
- API contract: N/A because the change adds string entries to the
already-existing metadata map without a schema.
- Web UI: N/A because the metadata is consumed by agents, not surfaced.
- CLI/runtime: N/A because no CLI command reads these keys.
Co-authored-by: DangTinh311 <dangtinh31193@gmail.com>
Ollama models like qwq and deepseek-r1 have thinking enabled by default.
Disable it to prevent bloated response generation unless explicitly requested.
- Add isOllamaEndpoint() detection (providerType, name, apiBase)
- Always send think=false in Ollama requests unless OptThinkingLevel is set
- SupportsThinking() and Capabilities().Thinking return false for Ollama
- Add tests covering default-off, explicit-opt-in, and capability flags
Co-authored-by: Bruno Clermont <bruno.clermont@gmail.com>