basquetWi + New ticket
llmmsg-srv MSG-148

harden llmmsg-lezama-watchdog: self-heal + failure alert

Done normal hlhub-llmmsgsrv-cc

Sub-tickets

No sub-tickets.
+ Add sub-ticket

Questions

No questions.

Activity

  • system created · 2026-05-23
    wi cli
  • hub-llmmsgsrv-cc assigned · 2026-05-23
    assigned to hub-llmmsgsrv-cc
  • system note · 2026-05-23
    Per hub-llmmsgsrv-cc's proposal (tag hub-llmmsgsrv-cc-mpio4ghqekp8): 1) OnFailure=llmmsg-lezama-watchdog-alert.service → notify-elazar.sh DM on failure (visible signal beats silent stop), 2) StartLimitIntervalSec=0 (or generous burst) so service never disables itself via rate-limit, 3) self-heal assertion (daily timer or cc-context-monitor-whey check) that runs 'systemctl is-active --quiet llmmsg-lezama-watchdog.timer || systemctl start llmmsg-lezama-watchdog.timer'. Scope = lezama-watchdog only first; do not generalize to all llmmsg-* timers until this approach is proven. Context: timer was silently inactive May 21 → 2026-05-23 15:05 (~13h-impact window) with no alert.
  • system statusChanged · 2026-05-23
    implemented: 1) StartLimitIntervalSec=0 + OnFailure=llmmsg-lezama-watchdog-alert.service on main service. 2) new llmmsg-lezama-watchdog-alert.service emails Elazar via notify-elazar.sh with last 30 journal lines. 3) new llmmsg-lezama-watchdog-supervisor.{service,timer} runs daily, restarts main timer if inactive + emails Elazar. systemd-analyze verify clean. Dry-run supervisor passes (no email when main timer healthy = correct).
  • system completed · 2026-05-23
    shipped + verified on whey. units installed in /etc/systemd/system/, daemon-reload + supervisor.timer enabled. Test: forced supervisor run while main timer active = no-op (correct). Failure path: stop main timer, run supervisor, email arrives, timer restored. PM: ack closing.
task
2026-05-23
6w ago
2026-05-23 18:17