Stop and escalate
Pause when the failure affects a broad ring, servicing health is unclear, data loss is possible, or the recovery step is not approved.
A repeatable first-response path for endpoint engineers. Capture the failure, separate policy from client state, collect evidence, and choose a safe next action before reaching for repair scripts.
Check an item only when you have observed it or recorded a clear exception. Preserve the original error and timestamp before changing state.
Pause when the failure affects a broad ring, servicing health is unclear, data loss is possible, or the recovery step is not approved.
Retain the original code, timestamps, comparison device, policy state, logs, action taken, and verification result.
This runbook organizes investigation; it does not promise that every update error has one universal fix.