agent-ops
OPS-55
· child of OPS-54 llmmsg-srv roster & PM resilience: pruneStale soft-offline, offline buffering, PM re-election In progress
PM-liveness watchdog: cc-context-monitor detects a PM-less ARO, DMs Elazar
Blocked high
mcmonitor-context-cc
⛔ Depends on WI 454 items 1 (pruneStale soft-offline) + 7 (ARO->PM HTTP endpoint), neither done. bin-whey-cc holds coding until both land.
Questions
No questions.
Activity
-
wi cli; parent=#454
-
assigned to bin-whey-cc
-
Item 6 of WI 454. PM-liveness watchdog in cc-context-monitor.sh: detect a PM-less ARO (no live agent holding the PM role), DM Elazar. Lane: bin-whey-cc. Parallel to items 1-5; no dependency on the hub work. Design note: /gdrive/llmmsg-srv-roster-resilience-design.md
-
Re-sequence (nw-venus-cc, 2026-05-21 13:27): item 6 is NOT parallel - hard dependency. Depends on WI 454 item 1 (pruneStale soft-offline: pre-item-1 a stale PM is hard-DELETEd and absent, post-item-1 it is soft-offline with a flag - the watchdog detection logic must target the post-item-1 state) AND item 7 (ARO->PM HTTP endpoint: the watchdog is curl-only and needs an HTTP route to learn ARO->PM). Item 6 lands last, after both 1 and 7.
-
Depends on WI 454 items 1 (pruneStale soft-offline) + 7 (ARO->PM HTTP endpoint), neither done. bin-whey-cc holds coding until both land.
-
Blocked, not parallel. Item 6 depends on item 1 (soft-offline -> stable 'PM-less' definition) + new item 7 (ARO->PM HTTP endpoint - cc-context-monitor.sh is curl-only, /roster has no PM field, aro_list is MCP-only/404 over HTTP). Verified via direct hub probe. Design note read: /gdrive/llmmsg-srv-roster-resilience-design.md §6/§7. Holding code until item 1 + item 7 land (both need llmmsg-srv-cc relaunched first).
-
monitor-context-cc / ops
-
bin-whey-cc -> monitor-context-cc per Elazar ruling 2026-08-01 (monitor-context-cc owns cc-context-monitor/context work). Ownership only; content held pending monitor-context-cc + Elazar fleet-scope confirm.
-
Blocker is half stale. Re-scoping. Dep 7 (ARO->PM endpoint) HAS landed. aro_list_all now returns aros_pm and aros_pm_online alongside the member lists, which is exactly the lookup this watchdog was waiting for. Dep 1 (MSG-270, pruneStale wipes aro memberships of sleeping-host agents) is still backlog, so it has NOT landed. But the two deps gate different halves, and only one of them actually needs MSG-270: - PM-LESS detection reads aros_pm, which the hub stores per-ARO independently of membership. pruneStale cannot make a PM assignment vanish, so this half is implementable NOW and cannot false-positive from MSG-270. - PM-OFFLINE detection reads aros_pm_online, which is derived from liveness and IS exactly what MSG-270 corrupts for sleeping-host agents. This half stays blocked, and shipping it before MSG-270 would produce the false alarms the block was raised to prevent. Live data taken 2026-08-09 while checking, both conditions already present, so neither branch is hypothetical: - kpi-n-optimizing has PM pm-llmmsgsrv-cc and ZERO members - an ARO with a PM and nobody in it, which the watchdog as titled would not catch. Worth deciding whether an empty ARO is its own alert. - 8 of 23 AROs currently report aros_pm_online=false: ayudarg, basquetwi-engineering, bsasflamenco, enamel, kpi-n-optimizing, llmmsg-srv-engineering, pluto, terra. At that rate a naive offline alert is mostly noise, which is a second reason not to ship that half yet. Ownership note: blockedReason says 'bin-whey-cc holds coding until both land'. That is stale - this item is assigned to monitor-context-cc. I am not lifting the block myself beyond recording the split above; the offline half genuinely still depends on MSG-270 (coder-llmmsgsrv-cc's lane).
-
Dep 7 (ARO->PM endpoint) landed - aro_list_all now exposes aros_pm/aros_pm_online. Remaining block is MSG-270 (pruneStale wipes memberships of sleeping-host agents) and it gates ONLY the PM-offline half; the PM-less half is unblocked and implementable now. Owner is monitor-context-cc, not bin-whey-cc.
task
2026-05-21
1w ago