Skip to content

feat(reboot): one Reboot operation — sync, land, newest image, ONE roll/restart, verify; plus a rate-limited wedge watchdog (stacked on #6124) - #6132

Merged
rbuergi merged 7 commits into
mainfrom
feat/instance-reboot
Oct 5, 2026
Merged

rbuergi merged 7 commits into
mainfrom
feat/instance-reboot

Conversation

@rbuergi

@rbuergi rbuergi commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Instance Reboot — one step back to a known-good, newest state

Maintainer directive after the control-instance outage: "can you please build some form of reboot feature." Policy instance-reboot (register row added as proposed). Manual: Doc/Architecture/InstanceReboot (new).

🚨 Stacked on #6124 (module reload + packages-auto-update). This branch carries #6124's two commits plus ONE commit of its own (f1d4a02ee4). Its module step CALLS the reload's resolve-and-land instead of forking it. Merge #6124 first. If this merges first, #6124's content lands with it.

What a reboot does

A request is written to Admin/_Reboot/{id} (InstanceRebootRequest, written only by InstanceReboot.Request, as System). An executor on the request's own hub then runs these steps in order. Each step is recorded as Ok, Skipped or Failed with its reason.

  1. Sync — every GitSynced module source is imported at its branch HEAD (ModuleSourceUpdate.UpdateAll). This uses the same UpdateToLatestFromGitHub call as git_hub_sync op=update. The seal is asked as for a person's update; modules are judged per manifest hash, and an incompatible floor or unmet dependency is declined by name.
  2. Modules — the newest compatible version of every installed module is landed through ModuleReloadExecutor.ResolveAndLand, extracted from feat(modules): reload + two-phase uninstall requests, policy packages-auto-update — independent of seal/sync #6124's executor and now shared. A floor decline is named; it is not red.
  3. Image — SelfUpdateHostedService.SelectImage (new IInstanceRebootActivation) picks the newest image that the instance's own update policy admits and the availability gate clears. It picks only from tags the registry actually lists. A wedged reboot leaves the combo (dependent-suite) verdict out and says so on the step.
  4. Restart — restartRequestedAt is stamped first, then ONE roll (migration first) or restart happens. The roll floor is not honoured. If no roll or restart can be scheduled, the reboot ends Failed.
  5. Verify — every process that booted after the stamp runs every registered IInstanceRebootCheck: health:nodetype_bake (red only for critical namespaces, default Hosting), health:pending_module_activation, health:content-types, and smoke:* checks. A process that measured no smoke check is RED. The thread-start smoke check ships with the AI module in the Plugins PR.

A red step before the restart does not stop the reboot. It does make the request Failed, naming every red step.

Who can start one

  • MeshOperations.RebootInstance(reason, wedged) — global admin only. The caller is the signature and is recorded as requestedBy. The MCP tool reboot_instance (Plugins) wraps it.

  • Control lane ControlLaneOperation.Reboot / RebootOperation. The plan is fixed (PlanFor), so the control side binds its digest without a dry run, with the requester as approver. The lane run ends when the request is filed.

  • RebootWatchdog (every process). The predicate is explicit and pure (RebootWatchdogRules). It fires when either holds:

    • ≥3 thread starts fail with a load/binding fault (MissingMember/TypeLoad/BadImageFormat/assembly FileNotFound/FileLoad) within 15 min, reported through hub.ReportWedgeFault;
    • OR a critical NodeType is in compile Error on every pass for ≥20 min.

    It is rate-limited: no firing while a reboot is open, at most one self-reboot per 6 h (also remembered in-process), and never twice for evidence a reboot already failed to clear within 24 h. Every firing, and every refused firing, logs at Critical.

Coordinator's lessons — all four now in this PR (second commit 1bb1551fa1)

  • Pullable image. Step 3 chooses only from the registry's own tag listing, so an unminted tag is never chosen.
  • No dependence on a wedged signer. The in-process path (the instance's own MCP reboot_instance, or the watchdog) signs nothing.
    • A self-patching instance restarts through its own service account.
    • 🚨 A control instance never hands its own restart to itself. If the hand-over route is its own inbox (Route.Local) and it cannot self-patch, the Restart step is Refused by name and nothing is handed over. Self-patch is REQUIRED for the control instance.
  • Keep serving. Before a self-patched roll or restart, the step reads spec.strategy/spec.replicas (new IDeploymentUpdater.ReadRolloutStrategyAsync, a default member, so no implementer is obliged; the AKS updater implements it in Plugins#2899).
    • It rolls only if maxUnavailable resolves to 0 and maxSurge to ≥ 1 (RolloutStrategyReading.NonDisruptiveRefusal).
    • Recreate, the 25%/25% default on 4 replicas, a zero surge, an unresolvable value and an unreadable strategy are each Refused by name, and nothing is rolled.
  • Singletons resumed. New check health:singletons-resumed waits up to 35 min, on each node's own stream, for the PR babysitter (Hosting/Babysitter.lastRunAt) and the PR review sweep (Hosting/Triage/Status.lastPrSweepAt) to stamp a pass NEWER than the restart.
    • It is red if one did not, naming its stale last pass.
    • An absent node is reported not measured, by name.
    • Checks now run side by side, and each may state its own budget. IInstanceRebootCheck.Run receives the request.

Tests (Memex.Portal.Shared.Test → InstanceReboot* + RolloutStrategyRuleTest, 19)

  • Acceptance, through the real registry client, landing, reconciler, executor, agent and SelfUpdateHostedService (only its k8s updater is recorded): M@1.1.0 is landed, and 1.2.0 is published → Sync Skipped (named), Modules lands 1.2.0, Image Skipped (named), exactly ONE restart, despite being inside the roll floor. A process booted after the restart, loading 1.2.0, passes the thread-start stand-in → Done, with requestedBy recorded.
  • Negative: the booted process still binds 1.1.0. The stand-in throws the incident's MissingMethodException → Failed, naming Verify: … smoke:thread-start.
  • A floor above the running platform is declined by name; the reboot still restarts, and N keeps serving.
  • A failing step is red by name: an updater that cannot restart → Restart Failed (Unavailable…), the request Failed, asked exactly once, and earlier steps are not painted red.
  • Watchdog (real mesh):
    • a healthy instance files nothing;
    • 5 non-binding failures are not a wedge (negative);
    • 3 MissingMethodException thread-start faults → ONE watchdog reboot (Trigger Watchdog, Wedged, evidence recorded); the next pass is Suppressed "rate-limited", and still exactly 1 request.
  • Rules: the fault classifier, the threshold and window, the compile-leg continuity and its reset on an unreadable pass, and rate-limit/open/same-evidence. Each has its control.
  • Lane: fixed plan digest = PlanFor; executing it files one Person reboot for the lane requester; a wrong target is refused.
  • Disruptive rollout: with a 25%/25% strategy on 4 replicas, the Restart step is Refused … maxUnavailable resolves to 1 of 4, and 0 restarts are issued. The acceptance scenario, with a non-disruptive strategy and 1 restart, is the control.
  • Control self-hand-over: a real self-updater with CanPatch=false on the local-inbox route is Refused … IS the control instance … nothing was handed over, with 0 restarts.
  • Singletons: a pass 49 min before the restart is red ("did NOT resume"), and the absent review sweep is reported not measured. A fresh pass that lands while the check waits turns it green.
  • Rollout rule, pure: 8 readings, each with its control.
  • Negative-control runs (reverted afterwards):
    • Making the restart honour the roll floor: 3/3 acceptance-class tests went red.
    • Disabling both guards (the rollout rule and the self-hand-over refusal): the rule test, the disruptive test and the control test all went red (3/3). Without its guard, the control instance attempted the hand-over to its own inbox.

Regression

  • Full Memex.Portal.Shared.Test: 2317 passed, 1 skipped, 0 failed (2318 total), run before the last two edits.
  • After the second commit: InstanceReboot* + RolloutStrategy* + ModuleReload* + SelfUpdate* 170/170; MeshWeaver.Documentation.Test 670/670; MeshWeaver.Graph.Test ControlLane 20/20.
  • The last two edits were: IInstanceRebootCheck moved to MeshWeaver.Graph so the AI module can implement it, and the lane test was added.
  • Release -warnaserror clean: Memex.Portal.Shared.Test (which builds Graph, GitSync, PluginCatalog, Mesh.Operations, Portal.Shared), MeshWeaver.Graph.Test, MeshWeaver.Hosting.Test, MeshWeaver.Documentation.Test.

NOT established

  • No real cross-process run: no Kubernetes roll, no multi-replica Orleans cluster, no real GitSync import (the test mesh has no sync source).
  • The thread-start check in these tests is a stand-in. The real one starts a thread, in the AI module (Plugins PR).
  • "Newest promoted image" means what the instance's update policy admits. On Continuous that is the newest admitted continuous build, not only clean promotion tags.
  • A roll handed to the control lane waits on that lane's own handling. On a wedged control instance that cannot self-patch, the hand-over goes to itself.
  • Two replicas' watchdogs could both fire within the listing's lag. The in-process memory closes this per process only.
  • The watchdog is ON by default (Reboot:WatchdogEnabled).
  • The rollout check covers the SELF-PATCH path only. An ordinary instance that hands its roll to the control plane's Roll is not checked there.

Recycle after deploy

None needed for core. New node type InstanceReboot, and the agent and watchdog arm on every mesh hub at boot.

Pairs-with: none — this adds public surface and removes none.
Implementers: IDeploymentUpdater.ReadRolloutStrategyAsync — given a DEFAULT implementation (answers null, which the reboot treats as a refusal), so no implementer is obliged; the AKS updater implements it in Plugins#2899. IInstanceRebootActivation and IInstanceRebootCheck are new.
Mirror-sync: none — this core change adds no catalog key.

🤖 Generated with Claude Code

…ll/restart, verify; plus the instance's own wedge watchdog

An instance can now be brought to a known-good, newest state in one step (policy instance-reboot,
Doc/Architecture/InstanceReboot). A durable request at Admin/_Reboot/{id} runs, in order, each step
reported with a named reason when skipped or failed:
1. Sync every GitSynced module source at its branch HEAD (ModuleSourceUpdate.UpdateAll — the same
   UpdateToLatestFromGitHub import as git_hub_sync op=update, per-module manifest-hash judgement,
   incompatible floor declined by name).
2. Land the newest compatible version of every installed module (ModuleReloadExecutor.ResolveAndLand,
   extracted from the module reload — not forked).
3. Pick the newest image the update policy admits and the availability gate clears
   (SelfUpdateHostedService.SelectImage; a wedged reboot does not wait for the combo verdict, and says so).
4. ONE roll (migration first) or restart, stamped first, roll floor not honoured.
5. Every process booted after the stamp runs IInstanceRebootCheck: critical-namespace bake,
   pending module activation, content types, and a smoke:thread-start check (no smoke measured = RED).

Surfaces: MeshOperations.RebootInstance (global admin; the caller is the signature), the control-lane
Reboot operation (fixed one-step plan, digest bound without a dry run), and RebootWatchdog — fires only
on repeated load/binding faults on thread start or a critical NodeType continuously in compile Error,
rate-limited (one per 6 h, never twice for the same evidence), alarmed at Critical.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@rbuergi
rbuergi enabled auto-merge October 5, 2026 08:37
@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results

    22 files      22 suites   45m 42s ⏱️
10 771 tests 10 578 ✅ 193 💤 0 ❌
10 784 runs  10 591 ✅ 193 💤 0 ❌

Results for commit 6377fe0.

♻️ This comment has been updated with latest results.

@systemorph-com systemorph-com Bot added the thread:pr-systemorph-meshweaver-6132-f0592c https://memex.systemorph.com/Hosting/Triage/_Thread/pr-systemorph-meshweaver-6132 label Oct 5, 2026
…f-hand-over

- Restart step refuses, by name, a self-patched roll whose portal Deployment rollout could drop serving
  pods: RollingUpdate with maxUnavailable resolving to 0 and maxSurge to >= 1 is required; Recreate,
  the 25%/25% default, a zero surge, an unresolvable or unreadable strategy are refused. New
  IDeploymentUpdater.ReadRolloutStrategyAsync (default: null => refused) + RolloutStrategyReading.
- A CONTROL instance (hand-over route = its own inbox) that cannot self-patch is refused instead of
  handing its own restart to itself; self-patch through its own service account is required.
- health:singletons-resumed: the PR babysitter (Hosting/Babysitter.lastRunAt) and PR review sweep
  (Hosting/Triage/Status.lastPrSweepAt) must stamp a pass newer than the restart within 35 min, else red.
- IInstanceRebootCheck.Run now receives the request; checks may state their own budget and run side by side.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@rbuergi
rbuergi disabled auto-merge October 5, 2026 09:02
@rbuergi
rbuergi enabled auto-merge October 5, 2026 09:03
@rbuergi
rbuergi disabled auto-merge October 5, 2026 10:56
@rbuergi
rbuergi enabled auto-merge October 5, 2026 10:57
@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 0)

  3 files    3 suites   3m 42s ⏱️
566 tests 375 ✅ 191 💤 0 ❌
570 runs  379 ✅ 191 💤 0 ❌

Results for commit 6377fe0.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 2)

    5 files      5 suites   3m 57s ⏱️
1 279 tests 1 279 ✅ 0 💤 0 ❌
1 280 runs  1 280 ✅ 0 💤 0 ❌

Results for commit 6377fe0.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 1)

1 753 tests   1 751 ✅  5m 56s ⏱️
    4 suites      2 💤
    4 files        0 ❌

Results for commit 6377fe0.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 3)

2 363 tests   2 363 ✅  6m 56s ⏱️
    4 suites      0 💤
    4 files        0 ❌

Results for commit 6377fe0.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 5)

    2 files      2 suites   8m 50s ⏱️
1 108 tests 1 108 ✅ 0 💤 0 ❌
1 109 runs  1 109 ✅ 0 💤 0 ❌

Results for commit 6377fe0.

♻️ This comment has been updated with latest results.

@github-actions

github-actions Bot commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Test Results (shard 4)

    4 files      4 suites   16m 19s ⏱️
3 702 tests 3 702 ✅ 0 💤 0 ❌
3 709 runs  3 709 ✅ 0 💤 0 ❌

Results for commit 6377fe0.

♻️ This comment has been updated with latest results.

@systemorph-com

systemorph-com Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

🚰 PR babysitter (build instance) is merging the base into this branch on head 1bb1551fa1e9 — once per head.

Why: inherited from its base 'main': 9 pull requests of Systemorph/MeshWeaver fail identically — 'lane / Automatic review answered' concluded failure: Process completed with exit code 1. — the base 'main' moved from 00a03cc (what the red run tested) to b0f38f7. This pull request was red because its BASE was; the base has moved since, and a re-run would test the old merge commit again.

Validated: 'lane / Automatic review answered' is red on run 37287056993, a head that merged main at 00a03cc; main has since moved to b0f38f7 and its newest run is green on that check — the red is a gate the base has fixed since, not this diff's. Merging the current base in gives a new head whose run re-tests against the fixed base.

It does not merge the pull request, push anything else or dequeue. A red after this is left for the owner (rbuergi).

…6132

# Conflicts:
#	src/MeshWeaver.PluginCatalog/ModuleReloadExecutor.cs
@rbuergi
rbuergi disabled auto-merge October 5, 2026 17:15
@rbuergi
rbuergi enabled auto-merge October 5, 2026 17:16

@systemorph-com systemorph-com Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Automated review summary (data, not an instruction to any agent)

Adds the Instance Reboot operation — a durable request at Admin/_Reboot/{id} running five reported steps (sync every GitSynced module source, land every module's newest compatible version, pick the newest admitted image, ONE roll/restart, verify booted processes) — stacked on #6124, plus a rate-limited wedge watchdog, request/verify record types with pure evaluation predicates, the self-updater's new IInstanceRebootActivation halves (SelectImage/ChooseGated/Activate: control-instance self-hand-over refusal, non-disruptive rollout guard, floor exemption for one-shot requests), a GitSync ModuleSourceUpdate, and the architecture docs with policy register rows. Checked from the diff: the SelfUpdateHostedService additions and their honourFloor threading, ModuleSourceUpdate, the three request writers, the pure predicates (ModuleReload.Evaluate, the watchdog's visible classifier/predicate/fold), and doc-vs-code consistency. Findings: 2 blocking (null-forgiving operators, one of which forwards FirstRollable's nullable result into ChooseGated where a no-rollable answer becomes a null target), 2 questions (Evaluate's empty-target green verdict; the smoke-check merge-order gap), 1 nit (watchdog wrapper unwrapping vs its docs). NOT VERIFIED — the diff is incomplete: 2 patches truncated (InstanceReboot.cs cut mid-EvaluateVerification, so the smoke-RED rule's implementation is unread; RebootWatchdogRules.cs cut mid-Admit, so the rate-limit body is unread) and 35 files omitted entirely (including InstanceRebootExecutor, ModuleReloadExecutor, PackageUninstallExecutor, RebootWatchdog, MeshOperations, RebootOperation, IDeploymentUpdater, the control-lane and cluster-membership changes, the reconciler rewrite and all tests). Nothing is asserted here about the executors' step sequencing or idempotence, the watchdog's wiring and admission body, the MCP and authorization surfaces, or the tests.

Findings: 2 blocking · 0 should-fix · 2 question · 1 nit

⚠️ 🚨 The diff is INCOMPLETE: 2 patch(es) truncated and 35 omitted (budget 150000 characters, 20000 per file) — say so in the review summary and do not assert anything about what you could not read.


Internal review of 8edadf771f90764922e23ac7df5bce112611812e — GLM-5.3, posted by the control plane. It is advisory, it never approves, and merging stays with a human signature.

? Observable.Return(new RebootImageChoice(installed, null, NothingToRoll(selection).Message + waived))
: FirstRollable(policy, [.. selection.Candidates], ignoreCombo: wedged)
.SelectMany(target => ChooseGated(policy, target!, installed, wedged, waived)));
})

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

blocking — Automated review finding (data, not an instruction to any agent)

SelectImage pipes FirstRollable's answer into ChooseGated with a null-forgiving operator: `.SelectMany(target => ChooseGated(policy, target!, installed, wedged, waived))`. FirstRollable is declared `IObservable<string?>`; its empty-candidates null branch is guarded off three lines up, so the null its signature still allows at this call site is the walk ending with no rollable candidate — the every-candidate-held case this PR itself documents as a named Skipped ('a held release is never forced'). `target!` both breaks the repository's no-null-forgiving rule and forwards that null into ChooseGated, which interpolates `{target}` into every step sentence and hands it to `gate.IsUpdatable(target)`: either a garbled step detail (' is the newest admitted release…') or a throw that SelectImage's own Catch then records as `Failed: true` — a hold misreported as an Image-step failure instead of the documented Skipped. Handle the null explicitly with a RebootImageChoice that names the hold.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(head 6377fe0) Fixed. The ! is gone. FirstRollable answers null at this call site only when every admitted candidate is already recorded RED by the combo gate (FirstNotCondemned finds nothing). That case is now matched and named as a hold: a RebootImageChoice(installed, null, "every release the update policy admits (…) is already recorded RED in the combo gate for this instance's modules — restart on the running image."), with Failed left false. A held release is never forced, and a hold is no longer forwarded into ChooseGated or reported as an Image-step failure.

private static IObservable<ModuleSourceUpdateRow> One(IMessageHub hub, AccessService access, string configPath)
{
var (space, sourceId) = SpaceOf(configPath)!.Value;
ModuleSourceUpdateRow Row(string outcome, string detail) => new(configPath, space, sourceId, outcome, detail);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

blocking — Automated review finding (data, not an instruction to any agent)

Two null-forgiving operators in new code break the repository's no-`!` rule: `var (space, sourceId) = SpaceOf(configPath)!.Value;` here, and `log!.Messages` at line 147. Both silence nullable warnings the compiler cannot discharge — SpaceOf's non-nullness is established only by UpdateAll's `Where(p => SpaceOf(p) is not null)` in another method, and log's only by the `Where(log => log is not null && …)` two operators upstream. The fix is mechanical (pattern-match instead, e.g. `SpaceOf(configPath) is { } s`), and would fail loudly at compile time if a future caller ever reached One() without the filter.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(head 6377fe0) Fixed, both of them.

  • SpaceOf(configPath)!.Value: the parse is now matched once in UpdateAll (SpaceOf(path) is { } parsed ? One(…, parsed.Space, parsed.SourceId) : Observable.Empty<…>()), and One takes space and sourceId as parameters. A future caller that skips the filter now fails at compile time. The separate Where(p => SpaceOf(p) is not null) stays, so the dedupe and order run over valid paths only.
  • log!.Messages: the stream is narrowed with .OfType<ActivityLog>() before .Where(log => log.Status.IsTerminal()), so log is non-null by type.

IReadOnlyCollection<string>? aliveMembers = null)
{
var targets = request.Items
.Where(i => i.Failure is null && !string.IsNullOrWhiteSpace(i.TargetVersion))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

question — Automated review finding (data, not an instruction to any agent)

Evaluate builds `targets` only from items that have no failure and a target version. When every item failed or was declined, `targets` is empty and, with at least one counted report, the predicate answers a GREEN verdict — mismatches is empty over an empty target set — whose detail reads 'N replica(s) report  loaded' with an empty module list. Can the executor reach Evaluate in that state, or does it red the request from the items' own failures before consulting the predicate? ModuleReloadExecutor.cs is omitted from this diff, so I cannot tell whether an all-declined reload could be marked Done here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(head 6377fe0) It can't be marked Done in that state, and that is decided before Evaluate. ModuleReloadExecutor.Decide builds needing from the same filter (no failure, has a target version, and the target differs from the running version). When every item failed or was declined, needing is empty and the request is finished right there: Activation = NotNeeded, Status = failure is null ? Done : Failed, CompletedAt stamped, with the items' own failures joined into Failure. So an all-failed or all-declined reload ends Failed and never reaches AwaitingRestart/Activating, the only states that call Evaluate. A mixed request that does reach Evaluate keeps the failures it already recorded, because the final write is Status = r.Failure is null ? Done : Failed. A green verdict over the healthy subset therefore cannot turn a partial failure into Done. No code change.

| `packages-auto-update` | Every installed package updates by itself as soon as a newer COMPATIBLE version is published to the registry: the only inputs are "a version was published" and "its declared platform floor ≤ the running platform". It depends on NOTHING else — not a platform image build, deploy or roll, not a publication seal or a seal for the running framework identity, not a green build of the platform or of the package repository beyond the package's own publish, not the platform's CD or fleet arming, not a per-identity prebuilt match, and not the partition's sync source (the module lane's sync-owned hold is retired). A fresh install is `Auto`; a record SEEDED with the old reminder-only default is migrated to `Auto` at boot; the only opt-outs kept are a per-package pin (`None`) or a policy a global administrator chose on the catalog card, each named in the log. An incompatible floor is declined by name and the running version keeps serving. What lands is activated through the module reload path (live, else exactly one automatic restart, no approval). Manual: [Module Reload](../ModuleReload) → "Auto-update". | in force | 2026-10-05 | maintainer |
| `package-uninstall-request` | A platform admin uninstalls a package through ONE durable request (`Admin/_PackageUninstall/{id}`), in two phases. Phase 1, no confirmation: refuses by name a package not installed, a partition shared with another installed package or holding user data; retires the module (live, else exactly one automatic restart), closes the package's hubs, removes the install record, blocks every unattended re-install, and records exactly what phase 2 would destroy. Phase 2, the irreversible drop of the partition storage and its registry record through the platform's governed partition teardown, runs only after the REQUESTER repeats the partition name, recorded with who and when; without it the package stays uninstalled with its data retained. **Owed:** the MCP `uninstall_package` tool (MeshWeaver.Plugins), and a package-page entry. Manual: [Package Uninstall](../PackageUninstall). | proposed | — | maintainer |
| `module-reload-request` | Any authorised caller (a platform admin, an agent through MCP, the platform's own watchers) can ask an instance to reload one module, or all, through ONE durable request (`Admin/_ModuleReload/{id}`). The instance resolves the newest compatible published version — its declared platform floor against the running platform, never a seal or a green platform build — lands it through the existing landing path, activates it live when the module can be swapped and otherwise by exactly one automatic restart through the self-update restart path, with no governed activity, no approval and no confirmation, and reports on the request node the version found, what landed, how it activated and what every replica loaded, or RED with the reason. **Owed:** the MCP `reload_module` tool and the fleet-target intake's switch from `RefreshModules` (MeshWeaver.Plugins), and the live loader's swap registering `IModuleLiveActivation` (MeshWeaver#6121, slice 2). Manual: [Module Reload](../ModuleReload). | proposed | — | maintainer |
| `instance-reboot` | ONE operation brings an instance to a known-good, NEWEST state: a durable request (`Admin/_Reboot/{id}`) runs, in order, Sync (every GitSynced module source at its branch HEAD, the `git_hub_sync op=update` import), Modules (the newest COMPATIBLE published version of every installed module, the module-reload landing), Image (the newest image the instance's update policy admits and its availability gate clears), ONE roll/restart that activates all of it, and Verify (every process booted after it: the critical-namespace bake, pending module activation, content types, and a thread-start smoke check — a process that measured no smoke check is RED). Every step is reported on the request with a named reason when skipped or failed. A person's call (MCP `reboot_instance`, the Reboot button, a platform admin) IS the signature — no second approver, the requester recorded; the instance's own watchdog files it without approval ONLY on the explicit wedge predicate (repeated load/binding faults on thread start, or a critical NodeType continuously in compile Error), at most once per interval, never twice for evidence a reboot already failed to clear, and every firing or refused firing is a Critical alarm. **Owed:** the MeshWeaver.Plugins half (the MCP tool, the deployment page's button and the `Reboot` InstanceAction kind, the AI module's thread-start smoke check and wedge reports). Manual: [Instance Reboot](../InstanceReboot). | proposed | — | maintainer |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

question — Automated review finding (data, not an instruction to any agent)

The `instance-reboot` register row marks the MeshWeaver.Plugins half as Owed — including 'the AI module's thread-start smoke check' — while the reboot's Verify step makes a counted process that measured NO smoke check RED (per InstanceReboot.md and the request's own documented rule). Between this merge and that owed half being rolled out, every reboot on a host without the AI module would end Failed by construction: verification that cannot measure fails closed. Is that gap intended to order the merges (the Plugins half first, or together), and if so should the register row or the manual say so explicitly, so an operator reading a Failed reboot knows the missing instrument is the cause and not a broken instance?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(head 6377fe0) Yes, that ordering is intended, and it is now written down in both places. It fails closed on purpose: a reboot that cannot measure whether a thread starts must not report green, because that is the incident's own failure shape. Added:

  • Register row (instance-reboot): until the smoke check is rolled onto an instance, every reboot there ends Failed at Verify by design. The verdict names the cause ("no smoke check was measured (none registered on this host)"). Merge order: this core half first, then the MeshWeaver.Plugins half, and a reboot is trusted green only once both are live.
  • InstanceReboot.md → "What is NOT established": the same point for an operator reading a Failed reboot. The missing instrument is the cause, not a broken instance, and the rest of the verdict says which readings passed.

case FileLoadException { FileName: { } file } when IsAssembly(file):
return Line(exception);
}
exception = exception is TargetInvocationException or TypeInitializationException or InvalidOperationException

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit — Automated review finding (data, not an instruction to any agent)

LoadOrBindingFault unwraps `InvalidOperationException` on its way to InnerException, but neither its own XML doc ('Unwraps aggregate, target-invocation and type-initialization wrappers') nor the matching sentence in InstanceReboot.md lists it. The classifier is documented as explicit and conservative, so description and code should agree: name the wrapper in both docs, or drop it from the unwrap set.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(head 6377fe0) Fixed by documenting the wrapper. Dropping it would lose real cases, because dependency injection activation and Rx operators surface a binding fault as an InvalidOperationException's inner exception. The XML doc of LoadOrBindingFault and the matching sentence in InstanceReboot.md now name InvalidOperationException beside the aggregate, target-invocation and type-initialization wrappers. Both also state the conservative half: a wrapper is only looked THROUGH and never classified itself, so an InvalidOperationException with no load or binding fault inside is not a wedge signal. The code is unchanged: a non-matching inner exception still answers null.

@rbuergi
rbuergi disabled auto-merge October 5, 2026 17:59
rbuergi added a commit that referenced this pull request Oct 5, 2026
…aits

test(bake): wait for the held compile stamp, never sample HeldCount (red on #6132 shard 5)
@rbuergi
rbuergi enabled auto-merge October 5, 2026 18:00
rbuergi and others added 2 commits October 5, 2026 20:00
…no null-forgiving in the source update, the smoke-check merge order and the IOE unwrap are written down

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@systemorph-com

systemorph-com Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

🩹 PR babysitter (build instance) is re-running the failed jobs of run(s) 37364268896, 37364303894, 37364285499, 37364268963, 37364268738, 37364268731, 37364263816, 37364303491, 37364287058 on head 6377fe07ea18 — once per head, never again for this commit.

Why: no runner acquired the job ('Consolidate test results' concluded cancelled: The job was not acquired by Runner of type hosted even after multiple attempts); 16 more infra check(s).

Validated: 'Consolidate test results' concluded cancelled because 'The job was not acquired by Runner of type hosted even after multiple attempts', with 16 more checks of the same signature, and both readable log tails answer 404 — the runners were lost mid-job. Nothing in the evidence names a failing test, compile error or any of the changed files (reboot/watchdog sources). Re-running the failed jobs of these nine runs re-takes the measurement.

It does not merge, push or dequeue. A red after this re-run is left for the owner (rbuergi).

@rbuergi
rbuergi merged commit 1d1e202 into main Oct 5, 2026
99 of 121 checks passed
@rbuergi
rbuergi deleted the feat/instance-reboot branch October 10, 2026 13:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

thread:pr-systemorph-meshweaver-6132-f0592c https://memex.systemorph.com/Hosting/Triage/_Thread/pr-systemorph-meshweaver-6132

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant