Runs on models without a registered tokenizer (e.g. 9router brand models)
ended with the generic "Agent couldn't generate a response" fallback even
though the real request used about 55% of the context window.
PruneStage counted history with TokenCounter, which falls back to a
chars/2 heuristic for unregistered models and overcounted about 1.8x.
Once over budget it ran memory flush (~35s, invisible in traces), then
mid-loop compaction, which cannot summarize a history made only of tool
call/result pairs. The callback reported the untouched history as
compacted, PruneStage still saw it over budget and returned AbortRun
before any LLM call, and FinalizeStage replaced the empty reply with the
fallback.
- PruneStage and ContextStage overhead count with the request guard's
BudgetCounter. PruneStage no longer controls loop flow; the final
request guard in ThinkStage decides.
- CompactMessages returns ErrNotCompacted when history is unchanged.
Callers stop counting it as a compaction and do not retry it in the
same run, while post-run summarization still sees the pressure.
- When the guard exhausts every reduction step, ThinkStage stops the run
with a localized chat.context_budget_exceeded notice instead of an
error, so the run's tool results are still persisted. The stop reason
marks the trace and agent span as error; team tasks, cron and
heartbeat treat it as a failure via RunOutcome.Failure().
- Memory flush and mid-loop compaction emit event spans.
- Web and desktop UIs treat an unset context_pruning as enabled (the
backend default since 7639a8c0), keep it unset when untouched, and can
re-enable pruning after it was turned off.
Cron job execution set the tenant ID on its context but not the tenant
slug. Tenant-scoped filesystem paths (skills-store, workspace, media via
config.TenantScopedDir) key off the slug and fall back to an id-based
path when it is absent — a different directory than where HTTP/WS skill
upload materialized the files (which sets the slug). As a result a cron
agent turn in a non-master tenant saw NONE of its tenant's managed
skills: skill_search returned 0 results and the agent, unable to run the
skill, produced an ungrounded answer.
Add cronTenantContext() which resolves the tenant slug via TenantStore
and sets both WithTenantID and WithTenantSlug. Master tenant and
nil-store/lookup-failure paths fall back to id-only (prior behavior).
Thread TenantStore into makeCronJobHandler and runCommandCronJob.
Tested: added unit tests for cronTenantContext (slug injected for
non-master; master skips lookup; nil store and lookup error fall back to
id-only). Verified end-to-end on a live tenant: before, a daily-agenda
cron guessed an empty day; after, it read the real event from the DB.
Note: other background executors that build a context from a tenant ID
(e.g. heartbeat) likely share this gap and are worth an audit.
The per-run session reset was gated on `!job.Stateless`, inverting the
flag: stateless jobs (the token-saving default) skipped the reset and
accumulated unbounded history, while stateful jobs were wiped every run.
Reset for `job.Stateless` instead, and clear BOTH session layers — the
goclaw session store AND the Claude CLI on-disk .jsonl. claude-cli
resumes its own session by a deterministic per-key UUID, so without
clearing the .jsonl a "stateless" run still replayed the entire
accumulated history (and silently grew it run after run).
makeCronJobHandler builds the agent RunRequest with no provider/model, so cron
jobs always run on the agent default provider. High-frequency scheduled jobs
(e.g. triage) can't be routed to a cheaper model the way heartbeats already can
via agent_heartbeats.provider_id/model.
Add provider_id/model to cron_jobs (Postgres migration 000074; SQLite schema +
migration v42→43, both mirroring agent_heartbeats) and resolve them into
RunRequest.ProviderOverride/ModelOverride in the handler, reusing the existing
per-run override plumbing that heartbeats use. NULL/unset → agent default, so
existing jobs are unaffected.
Settable today via UpdateJob patch (CronJobPatch.ProviderID/Model) or directly
in the DB; tool/RPC/dashboard exposure left as a follow-up.
Verified: go build ./... and go build -tags sqlite ./... (exit 0), go vet,
gofmt, and cron unit tests on both pg and sqlite stores.