Choose the correct support surface
Support works when each kind of question lands on the surface built for it:
- Documentation and Guides — architecture and rollout orientation (you are here). Start with getting started or the pilot rollout checklist.
- System Status — qualitative public health notices for marketing edges and the identity handoff. We publish no uptime percentages there; a public number we cannot keep truthful in real time does not get published.
- Contact — institutional questions that need a human desk.
- The authenticated console — live deployment issues for credentialed operators. In-product support routes carry the context a public form never has: your brand, your deployment, your grants.
The split is deliberate. Public status pages orient; they cannot replace incident operations, and a support flow that starts from an anonymous form starts from zero context.
The incident record
When something goes wrong, the record is the product of the response. Capture, as it happens:
- Affected workflow and scope — which brands, locations, and operator queues are touched.
- First known signal — what told you, and when. Not the polished timeline reconstructed after recovery.
- Owner — the named operator running the response.
- Operating decisions — what was decided, by whom, on which evidence, through which approval.
- Mitigation and exception state — what was degraded deliberately, what failed on its own.
- Recovery evidence — how the restored workflow was verified, not merely restarted.
Preserve what operators knew at each step. A timeline rewritten after the fact is a narrative; the audit trail & evidence model exists so the record is written at the moment of action instead.
Safe communication
- State observed impact without guessing at cause.
- Separate confirmed facts from active investigation.
- Avoid publishing customer, tenant, credential, or record-level detail.
- Do not invent response-time or availability commitments.
- Close the loop only after the affected workflow is verified, not merely after a process restarts.
Degrade safely, on purpose
Readiness includes deciding — before an incident — how each workflow behaves when its dependencies are uncertain. A roster that cannot reach its qualification data does not guess at qualifications; it holds publication and pages the scheduler. A payout flow that cannot verify its inputs pauses with a visible exception instead of posting on a guess. Safe degradation is a designed state with a named owner, not an improvised outage: operators can see what is degraded, why, and what unblocks it. The same fail-closed posture the platform applies to data isolation and ledger entries applies to availability.
After the incident
Every incident ends in a review whose output is bounded follow-up work, not blame: what signaled first, which decisions helped, which controls were missing, and what changes make the next occurrence cheaper. Findings land as workflow or policy changes with owners — and the review itself is part of the record, linked to the incident timeline it examines.
The authenticated incident record is authoritative. Public updates are a bounded communication surface derived from confirmed operational state. The essay incident readiness without status theater covers the operating posture behind these rules.
