8 Commits
Author SHA1 Message Date
thotam 3f90057c24 fix(pipeline): stop aborting runs on heuristic context budget estimates (#1587)
Runs on models without a registered tokenizer (e.g. 9router brand models)
ended with the generic "Agent couldn't generate a response" fallback even
though the real request used about 55% of the context window.

PruneStage counted history with TokenCounter, which falls back to a
chars/2 heuristic for unregistered models and overcounted about 1.8x.
Once over budget it ran memory flush (~35s, invisible in traces), then
mid-loop compaction, which cannot summarize a history made only of tool
call/result pairs. The callback reported the untouched history as
compacted, PruneStage still saw it over budget and returned AbortRun
before any LLM call, and FinalizeStage replaced the empty reply with the
fallback.

- PruneStage and ContextStage overhead count with the request guard's
  BudgetCounter. PruneStage no longer controls loop flow; the final
  request guard in ThinkStage decides.
- CompactMessages returns ErrNotCompacted when history is unchanged.
  Callers stop counting it as a compaction and do not retry it in the
  same run, while post-run summarization still sees the pressure.
- When the guard exhausts every reduction step, ThinkStage stops the run
  with a localized chat.context_budget_exceeded notice instead of an
  error, so the run's tool results are still persisted. The stop reason
  marks the trace and agent span as error; team tasks, cron and
  heartbeat treat it as a failure via RunOutcome.Failure().
- Memory flush and mid-loop compaction emit event spans.
- Web and desktop UIs treat an unset context_pruning as enabled (the
  backend default since 7639a8c0), keep it unset when untouched, and can
  re-enable pruning after it was turned off.
2026-09-29 18:04:08 +07:00
ナムandCollective Developer 6b738924d5 fix(discord): qualify group titles across surfaces (#1420)
Co-authored-by: Collective Developer <man@collective.dev>
2026-07-10 22:25:45 +07:00
MToan 7844df74f5 fix(cron): scope cron context to tenant slug so managed skills resolve (#1378)
Cron job execution set the tenant ID on its context but not the tenant
slug. Tenant-scoped filesystem paths (skills-store, workspace, media via
config.TenantScopedDir) key off the slug and fall back to an id-based
path when it is absent — a different directory than where HTTP/WS skill
upload materialized the files (which sets the slug). As a result a cron
agent turn in a non-master tenant saw NONE of its tenant's managed
skills: skill_search returned 0 results and the agent, unable to run the
skill, produced an ungrounded answer.

Add cronTenantContext() which resolves the tenant slug via TenantStore
and sets both WithTenantID and WithTenantSlug. Master tenant and
nil-store/lookup-failure paths fall back to id-only (prior behavior).
Thread TenantStore into makeCronJobHandler and runCommandCronJob.

Tested: added unit tests for cronTenantContext (slug injected for
non-master; master skips lookup; nil store and lookup error fall back to
id-only). Verified end-to-end on a live tenant: before, a daily-agenda
cron guessed an empty day; after, it read the real event from the DB.

Note: other background executors that build a context from a tenant ID
(e.g. heartbeat) likely share this gap and are worth an audit.
2026-07-07 10:52:17 +07:00
nguyenha935andGoClaw Operator 969a5879ae fix: unify no reply detection on dev (#1303)
Co-authored-by: GoClaw Operator <operator@goclaw>
2026-06-29 11:25:13 +07:00
Zezae Oh c0531e1e07 fix(cron): reset session for stateless cron jobs, not stateful ones (#1262)
The per-run session reset was gated on `!job.Stateless`, inverting the
flag: stateless jobs (the token-saving default) skipped the reset and
accumulated unbounded history, while stateful jobs were wiped every run.

Reset for `job.Stateless` instead, and clear BOTH session layers — the
goclaw session store AND the Claude CLI on-disk .jsonl. claude-cli
resumes its own session by a deterministic per-key UUID, so without
clearing the .jsonl a "stateless" run still replayed the entire
accumulated history (and silently grew it run after run).
2026-06-23 11:09:01 +07:00
Zezae Oh 2bf0f9e172 feat(cron): per-job LLM provider/model override (#1234)
makeCronJobHandler builds the agent RunRequest with no provider/model, so cron
jobs always run on the agent default provider. High-frequency scheduled jobs
(e.g. triage) can't be routed to a cheaper model the way heartbeats already can
via agent_heartbeats.provider_id/model.

Add provider_id/model to cron_jobs (Postgres migration 000074; SQLite schema +
migration v42→43, both mirroring agent_heartbeats) and resolve them into
RunRequest.ProviderOverride/ModelOverride in the handler, reusing the existing
per-run override plumbing that heartbeats use. NULL/unset → agent default, so
existing jobs are unaffected.

Settable today via UpdateJob patch (CronJobPatch.ProviderID/Model) or directly
in the DB; tool/RPC/dashboard exposure left as a follow-up.

Verified: go build ./... and go build -tags sqlite ./... (exit 0), go vet,
gofmt, and cron unit tests on both pg and sqlite stores.
2026-06-21 00:19:39 +07:00
Goon cd7c812b3d fix(cron): suppress no-reply deliveries by token 2026-06-13 09:01:56 +07:00
Duy /zuey/ f9440baca2 fix(cron): preserve SecureCLI credential context
Closes #54
2026-05-24 09:42:44 +07:00