mirror of
https://github.com/tiennm99/DocsGPT.git
synced 2026-10-04 14:12:58 +00:00
* fix: route remote-device tool through Redis so scheduled runs reach the device
The remote-device tool worked interactively but timed out on every scheduled
run. DeviceBroker was an in-process, in-memory singleton, but scheduled runs
execute in the Celery worker — a different process from the gunicorn web tier
that holds the device's SSE session — so a worker-side dispatch never reached
the device and the tool always hit its deadline.
Make the broker Redis-backed so every hop crosses the process boundary:
- queued commands -> Redis list dev:cmd:{device_id}
- output chunks -> Redis stream dev:out:{invocation_id}
- invocation metadata -> Redis hash dev:inv:{invocation_id}
- SSE upgrade tickets -> Redis key dev🎫{device_id}
Per-connection SSE session state stays in the web process. Reuses the existing
get_redis_instance()/CACHE_REDIS_URL; no new infrastructure. Also makes the web
tier safe to scale past one worker.
Concurrency hardening (from adversarial review + real-Redis e2e):
- XADD the output/control chunk before flipping completed=1, and have
drain_output do a final non-blocking flush after observing completion, so a
reader can't see completion and stop before the control chunk lands (this had
reintroduced the false "device did not respond (timed out)" under a race).
- _collect_result builds the result from drained chunks, checks the deadline
only after capturing a chunk, and falls back to the authoritative snapshot
(before cleanup) when no control chunk was observed.
- Audit outcome is written from locally-known fields so it survives the worker
racing to delete the invocation; a denied command now records a terminal
"denied" outcome instead of staying "dispatched".
- cmd-queue TTL raised to 900s (>= max drain deadline); dispatch-failure and
reaped-invocation cleanup; UTF-8 byte counts.
Tests: new tests/devices/{conftest (FakeRedis double), test_broker_cross_process,
test_broker_race, test_submit_output_audit}; drain/cleanup/ticket tests rewritten
for the Redis contract. The race tests fail against the pre-fix code. ruff clean;
device + tool-executor suites green.
* fix: log instead of silently passing on failed-dispatch cleanup
Addresses the code-quality lint on the best-effort hash delete in
dispatch_invocation's failure path: replace the bare `except: pass` with a
logger.debug carrying the invocation_id. No behavior change — cleanup stays
best-effort and still returns a failed Invocation.
60 lines
2.3 KiB
Python
60 lines
2.3 KiB
Python
"""Tests for ``DeviceBroker.cleanup_invocation`` queued-command removal."""
|
|
|
|
from __future__ import annotations
|
|
|
|
|
|
def _dispatch_offline(broker, invocation_id="inv_stale", device_id="dev_x"):
|
|
# No session draining -> the envelope stays queued on the device's list.
|
|
envelope = {"invocation_id": invocation_id, "action": "run_command"}
|
|
return broker.dispatch_invocation(device_id, "user_x", envelope)
|
|
|
|
|
|
def test_cleanup_removes_queued_command_so_it_doesnt_replay(broker_env):
|
|
# A timed-out invocation cleaned up while still queued must NOT later be
|
|
# delivered to (and run by) a freshly connected session.
|
|
broker, fake = broker_env
|
|
inv = _dispatch_offline(broker)
|
|
assert fake.llen("dev:cmd:dev_x") == 1
|
|
|
|
broker.cleanup_invocation(inv.invocation_id)
|
|
|
|
assert fake.llen("dev:cmd:dev_x") == 0
|
|
assert broker.get_invocation(inv.invocation_id) is None
|
|
|
|
# A session that connects now finds nothing queued.
|
|
sess = broker.register_session("dev_x", "user_x")
|
|
assert broker.next_command(sess, timeout=0.05) is None
|
|
|
|
|
|
def test_cleanup_deletes_invocation_and_output(broker_env):
|
|
# The metadata hash and output stream are removed on cleanup.
|
|
broker, fake = broker_env
|
|
_dispatch_offline(broker, invocation_id="inv_live", device_id="dev_live")
|
|
broker.submit_output_chunk(
|
|
"inv_live", {"stream": "control", "exit_code": 0, "duration_ms": 1}
|
|
)
|
|
assert fake.exists("dev:inv:inv_live") == 1
|
|
assert fake.exists("dev:out:inv_live") == 1
|
|
|
|
broker.cleanup_invocation("inv_live")
|
|
|
|
assert fake.exists("dev:inv:inv_live") == 0
|
|
assert fake.exists("dev:out:inv_live") == 0
|
|
assert broker.get_invocation("inv_live") is None
|
|
|
|
|
|
def test_cleanup_keeps_other_queued_commands(broker_env):
|
|
# Cleaning one invocation must leave a sibling queued command intact.
|
|
broker, fake = broker_env
|
|
inv1 = _dispatch_offline(broker, invocation_id="inv_1", device_id="dev_y")
|
|
_dispatch_offline(broker, invocation_id="inv_2", device_id="dev_y")
|
|
assert fake.llen("dev:cmd:dev_y") == 2
|
|
|
|
broker.cleanup_invocation(inv1.invocation_id)
|
|
assert fake.llen("dev:cmd:dev_y") == 1
|
|
|
|
sess = broker.register_session("dev_y", "user_y")
|
|
envelope = broker.next_command(sess, timeout=0.05)
|
|
assert envelope is not None
|
|
assert envelope["invocation_id"] == "inv_2"
|