pluto
PLUTO-157
Verify applog error→aro:pluto live-PM routing survives EVO-5 rail rename (kpi-n-optimization migration bs-mqmr47bwkk4)
Backlog normal
unassigned
Questions
No questions.
Activity
-
Pluto leg of the bs-mqmr47bwkk4 monitoring-standard decision. NO Pluto codebase/unit work: Pluto's applog-listen/pull are app-plane INSTANCES on the venus EVO-5 rail (per pm-venus topology correction) — rename + WatchdogSec + live-PM switch are nw-venus-owned (their venus WI, audit-PTD'd, live-error-path). aro:pluto.aro_config.pm=pm-pluto-cc CONFIRMED set, so the rail's hardcoded→live-aro-PM lookup resolves to me; no routing gap. Pluto obligation = verify-before-retire dual-run: confirm pluto error+fatal DMs still land in aro:pluto live PM + heartbeat-at-zero + watermark continuity AFTER nw-venus stands up the renamed instance, before old retires. Blocked-on: nw-venus rail rename. No action until migration runs.
-
MIGRATION GO received (generalpm, kpi PM, 2026-06-20): rename unblocked, aro:hosts-status PM=nw-venus-cc. SEQUENCE: OPS-22 (venus rail rename, nw-venus) FIRST → then PLUTO-157 follows (step 3). Pluto's leg = VERIFY-ONLY (no Pluto-owned units to rename): after OPS-22 lands, (1) confirm the renamed app-plane applog instances still route Pluto error/fatal → DM to pm-pluto-cc via aro:pluto.aro_config.pm hub-lookup (not hardcoded), (2) verify-before-retire dual-run (old+new sink both fire before old retired), (3) genuine level=error server-action fire test (needs coder, not 404/console-throw). Host-health routing→nw-venus-cc note is host-plane, does NOT affect Pluto app-plane routing. Still BLOCKED on OPS-22.
-
NAMING-RULE amendment (Elazar via generalpm, amends bs-mqmr47bwkk4): prefer DESCRIPTIVE names over abbreviations fleet-wide — spell the function out (applog_listen / applog_pull, NOT applog_lstn). Already-descriptive names stand. Carry into Pluto's verify leg: when checking the renamed app-plane applog rail (nw-venus-owned), confirm the new unit/sender names are self-documenting per this rule, in addition to confirming routing→aro:pluto.pm still fires. Still gated on app-plane-done ping (~19:00 UTC ETA).
-
APP-PLANE verify prerequisites (surfaced by nw-venus OPS-22 host-plane test, 2026-06-20): the test DM leg FAILED — venus→hub.pensanta.com:9703 returned 000 — and notify-nw.sh does NOT yet read ORIGIN_ARO from env (bin-whey sh.git fix filed). Pluto's applog error→DM rides the SAME hub path, so PLUTO-157's verify-before-retire MUST explicitly confirm, before any old-sink retire: (a) hub DM leg reachable from venus for Pluto's applog sender (the live error→DM to pm-pluto-cc actually delivers), (b) the sender/notify-nw.sh ORIGIN_ARO support has landed so the DM routes to aro:pluto.pm not a wrong target. If either is unmet, Pluto's real-time error alerting is silently broken = the 4-day-blind-spot class — do NOT retire the old sink until both pass. NOTE: my own hub sends from this venus session succeed → hub is generally reachable, the test 000 reads transient, but verify must PROVE it for Pluto's sender, not assume. Fixes are maintainer (bin-whey/pm-llmmsgsrv) lane; Pluto leg is verify-only.
-
BLOCKER CHAIN refined (pm-venus confirm, 2026-06-20 18:33): app-plane applog rail verify-before-retire is HELD — old rail stays LIVE (correct discipline), so Pluto error alerting is NOT broken, just unmigrated. PLUTO-157 unblock sequence now: (1) bin-whey notify-nw.sh ORIGIN_ARO patch (sh.git) → (2) nw-venus app-plane cutover w/ durable-watermark+catch-up covering the transient hub-000 → (3) Pluto verify leg. Gate is bin-whey's patch, not just nw-venus. Waiting on app-plane-done ping.
-
ROOT CAUSE of the DM-000 (nw-whey+generalpm, 18:33): hub.pensanta.com A-record still → dead ZT IP 172.27.178.27; ALL venus scripts dialing hub.pensanta.com:9703 fail. Correct venus path = http://127.0.0.1:9703 (reverse-SSH tunnel; hub runs locally on venus). Per-host EnvironmentFile fix (nw-venus lane), NOT sh.git. OPEN Q I raised to nw-venus: does the CURRENTLY-LIVE old applog-listen@pluto rail's DM leg use hub.pensanta.com:9703 (→ Pluto real-time error DMs silently 000 NOW = live blind-spot) or 127.0.0.1? If broken, the 'old rail stays live' safety net does NOT cover Pluto DM alerting → repoint/HUB_URL-fix is urgent + watermark must replay dropped Pluto errors. Awaiting nw-venus answer. (App-plane restructure itself DEFERRED to fresh nw-venus session, gated on bin-whey notify-nw.sh ORIGIN_ARO patch; PLUTO-157 verify still downstream.)
-
ALL-CLEAR on the live-blind-spot question (nw-venus DM, 18:35): Pluto's applog rail /home/rob/.config/applog/pluto.env already uses http://127.0.0.1:9703 (local hub on venus — hub was MIGRATED to venus, llmmsg-srv-venus.service; the old 'tunnel-to-whey' framing is stale, SoT=mcp__llmmsg-srv__guide). The dead hub.pensanta.com ZT DNS (172.27.178.27) NEVER touched the live error sink — only the new rkhunter scrp- units nw-venus set up this session, now corrected. Pluto real-time error→DM alerting was NOT blind at any point. PLUTO-157 verify leg unchanged: parked on deferred app-plane restructure, gated bin-whey MSG-78 (notify-nw.sh ORIGIN_ARO) → fresh nw-venus session → Pluto verify. Trigger = app-plane-done ping.
-
VERIFY-SCOPE addendum (lesson from OPS-22 rkhunter email false-pass, 18:37): 'email ✓' is NOT proof — the rkhunter-scan.sh mail SENT but rendered broken (wrong From = Elazar's personal addr not the monitoring sender; raw HTML from a duplicate header block corrupting MIME; '0\n0' count from a stray newline). When Pluto's verify leg runs on the renamed app-plane applog rail, do NOT accept exit-0/sent as green for the EMAIL leg either — confirm the applog-pull digest RENDERS as formatted HTML, correct From (dedicated monitoring sender, never elazar.pimentel@), clean values, and conforms to bs-mqmr47bwkk4 ('Report from: scrp-<sender> | <host> | <ts UTC>' first line). Both DM leg AND email leg get rendered-output verification, not just delivery. (rkhunter fix itself is bin-whey lane; this is Pluto's verify-criteria note.)
-
Upstream blockers CLEARED (18:37): notify-nw.sh v1.4 (ORIGIN_ARO support) pushed + hub.pensanta.com DNS repointed to venus 91.99.136.171. Remaining upstream sequence before Pluto verify: nw-venus pulls v1.4 → rkhunter fire-test DM lands GREEN at correct origin_aro (the OPS-22 host-plane done-gate) → app-plane applog verify-before-retire (deferred to fresh nw-venus session) → mars/pluto step-3 unblocks. Pluto verify leg still parked; trigger = app-plane-done ping.
-
VERIFY-CRITERIA addendum (Elazar hub-URL directive via pm-llmmsgsrv, 18:38): canonical hub URL for ALL scrp- EnvironmentFiles = https://llmmsg-hub.pensanta.com — NOT http://127.0.0.1:9703 (localhost, Elazar rejected), NOT hub.pensanta.com:9703 (dead). MSG-77 DNS-repoint CANCELLED. Pluto's applog pluto.env currently on 127.0.0.1:9703 (works, local hub, no live gap) but is now NON-STANDARD. Added to PLUTO-157 verify: confirm the renamed app-plane applog instances' EnvironmentFile uses https://llmmsg-hub.pensanta.com and DM still lands. Flagged to nw-venus to fold the applog-env URL move into the deferred app-plane restructure (action items only named the rkhunter file). Pluto verify leg still parked on app-plane-done ping.
-
APP-PLANE DESIGN consideration (parallel to host-plane routing Q, pm-llmmsgsrv→Elazar, 18:40): host-plane debating severity-branch routing — FATAL/ERROR → aro room directly (visible even if PM offline), CLEAN/INFO → DM local nw. Pluto's app-plane has the same shape: applog-listen DMs pm-pluto-cc directly for error/fatal → a Pluto fatal firing while pm-pluto-cc is offline sits buffered-but-unread, not room-visible. SOFTER for Pluto (llmmsg DMs are buffered not lost; applog-pull 30min digest is a backup), but the resilience pattern applies. When the deferred app-plane restructure runs, it should INHERIT whatever Elazar rules on the host-plane severity-routing Q (don't bake in strand-if-PM-offline). Tracking only — Elazar's design call, not acting. Pluto verify leg still parked on app-plane-done.
-
PARTIAL: applog pluto.env (+mars/venus) moved to https://llmmsg-hub.pensanta.com, services restarted+active (nw-venus, 18:42). URL-ALIGNMENT sub-criterion MET. NOT marking the routing leg verified on 'active' alone (verify-before-claim-live): a restart on a new hub URL is a config change to the LIVE error sink — process-up != delivers. REMAINING OPEN legs: (1) live error→DM delivery proof — genuine level=error server action → DM lands at pm-pluto-cc, to batch w/ coder-pluto (deferred: non-urgent, Elazar away, old-rail-on-new-URL is live sink meanwhile); (2) rename/restructure verify-before-retire (the deferred app-plane piece); (3) email-render on applog-pull digest. Asked nw-venus to confirm clean re-attach to appEvents error stream post-restart (journal). WI stays open.
-
LEG 1 PROVEN ORGANICALLY (18:42): a real applog-pluto error→DM landed at pm-pluto-cc immediately after nw-venus restarted applog on the new canonical URL, carrying a 19h-stale event (2026-06-19 23:15:41, id febe0441) = catch-up replay over the restart gap. Proves end-to-end: restarted applog-listen@pluto connected to hub on https://llmmsg-hub.pensanta.com + delivered DM to pm-pluto-cc + durable watermark/catch-up works. Live error→DM delivery leg satisfied without a coder fire-test. REMAINING PLUTO-157 legs: (2) rename/restructure verify-before-retire (deferred app-plane piece) + descriptive-naming check; (3) email-render on applog-pull digest. ---- TRIAGE of the replayed event: '[error] code challenge does not match previously saved code verifier /auth/callback (auth/callbackExchangeFailed)' = PKCE code_verifier/challenge mismatch in OAuth token exchange. Inherent OAuth-flow transient (cross-device login / stale-or-duplicate tab / expired-or-cleared verifier cookie / abandoned-then-retried login). Upstream of DB lookup → unrelated to PLUTO-152 gmail-dot. Single occurrence, 19h stale (replay not live spike), not flagged recurrent by applog-pull digest. VERDICT: benign transient, no code defect, no fix. Recurrence/clustering → pull auth/callbackExchangeFailed frequency from db-pluto; single stale instance doesn't warrant it.
-
Source: brainstorm bs-mqmxqxwop02 (closed by pm-venus 2026-06-20), Elazar-initiated. Large screens = forward-looking SALES/first-impression surface (real buyers/prospects aren't in telemetry; decision is NOT count-driven — current large-screen = 4.5% of logged sessions, real). SCOPE GUARDRAIL: polish ONLY stakeholder surfaces — admin/* + reportes/informes + public landing. Do NOT touch mobile-first alumno OR docente paths (live Mars data: docentes are 90% mobile, ≥1920=0%; alumno 384-447 portrait=66%). Mobile-first stays law for both. Audit-cited current Mars state (globals.scss): GOOD baseline — content capped+centered at 1600px (.main-content-inner max-width:100rem), no edge-to-edge sprawl, tables overflow-scroll, auth cards capped. GAP = under-exploited width, NOT sprawl: ZERO xl(1280)/2xl(1536) breakpoints anywhere (0 in .tsx, highest SCSS media query min-width:1024px), so 1920 renders identically to 1024 → 160px dead gutters + data grids frozen at 2-up. CONCRETE WINS (priority A>B>C, all additive media queries, fully reversible): (A) add xl/2xl density tiers to data grids — report-bars-grid 2→3, informes/alumnos/prácticas → more columns or master-detail side panel (turn whitespace into density); (B) line-length cap ~40-48rem on prose + single-column forms (currently inherit full 1600px); (C) promote filters/summary into the side gutter at xl. VERIFY: pixel-confirm at 1920/2560 via proxy agent (authed routes). SHIP-TIMING: Elazar's call given small current N — work is clear+reversible, queues without blocking. Fleet-uniform fix (venus+pluto same root: add density tiers, not wider caps).
task
2026-06-20 by wi-cli-venus
6w ago