mars
MARS-431
Mars alert rail blind: applog-mars pruned from aro:mars, error pages silently dropped
Done high
pmpm-mars-cc
Hub 400s every alert POST from applog-mars with origin_aro_not_member since 07:26:16 UTC 2026-07-13 — daemon alive, NOTIFY fires, classification correct (error-tier), but sender isn't a member of aro:mars so the hub rejects the post. Root cause of salud-datos error (relation vDataHealthSummary does not exist, digest=57073181) not reaching aro:mars despite matching MARS-369 paging convention. Same prune swept proxy-mars-cc-w ~07:00 UTC. Fix is hub/rail domain (llmmsg-srv maintainer only) — coder-mars-cc already escalated full root cause + recommended fix to pm-llmmsgsrv-cc. This WI tracks Mars-side visibility only: until applog-mars rejoins aro:mars, ALL future Mars error-tier events silently fail to page.
Questions
No questions.
Activity
-
Confirmed recovery: applog-mars alert for the original 07:26:13 UTC salud-datos error just landed in aro:mars (delivered 07:39 UTC) — rejoin/fix appears applied hub-side. Rail is posting again.
-
Confirmed recovery: applog-mars alert for the original 07:26:13 UTC salud-datos error just landed in aro:mars (delivered 07:39 UTC) - rejoin/fix appears applied hub-side. Rail is posting again.
-
Root cause: applog-mars pruned from aro:mars ~07:00 UTC, hub 400'd every alert post (origin_aro_not_member). pm-llmmsgsrv-cc re-inserted applog-mars into aro:mars 07:39:20 UTC (systemic hub-side prune-exemption tracked separately as MSG-267, their lane). Missed alert was not lost — daemon re-enqueues on fail, posted on recovery. Reporting confirmed live/working since. Diagnosis-only task per Elazar's ask; daemon self-heal hardening spun out separately.
bug
5w ago by wi-cli-venus
5w ago
2026-07-13 07:48