Sheet BEA-11 — ITSM Manual
Major Incident Operations
Declaring a major incident, running roles and the comms drumbeat, standing down, and closing with a mandatory post-incident review.
Scope
Major incidents are won or lost on coordination. This runbook keeps the operating rhythm, roles, and stand-down model while removing private platform detail.
This is the operator runbook for Major Incident Management (MIM): declaring a major incident, running it with a commander and a comms drumbeat, standing it down when service is restored, and closing it with a mandatory post-incident review (PIR) that spawns the root-cause problem. MIM is a process layer on the ordinary incident ticket — you drive it from the incident’s detail pane on the agent desk (the relevant workflow) and from the command palette (⌘K). The policy that governs auto-declaration and comms cadence lives in the Administration console (the relevant workflow).
Beacon ships MIM ready to use: a fresh tenant has a “Default major incident policy” (auto-declare candidate at priority 1, 30-minute comms cadence), so the process works on day one.
Seeing what is running right now
The service desk carries a live major-incidents board at the top of the queue, and it appears the moment anything is declared — you do not go and look for it. One row per active major: number, priority, title, affected service, how long it has been running, who is running it, and when the next stakeholder update is due.
Two things it makes impossible to miss:
-
Overdue comms. A row whose update is late is tinted red, badged with how late, and sorted to the top, and the board header counts them. If two majors are running and one has gone quiet, that is the first thing on the screen.
-
An incident nobody is leading. If no incident commander has been assigned the row says so in as many words rather than showing an empty space.
Open runbook on any row takes you straight into that incident’s major-incident panel — the runbook, not just the ticket. The board collapses if you need the screen, and it shows at most ten rows, telling you how many more there are rather than truncating in silence.
Requesters see a version of the same honesty in the Help Center banner: whether somebody is leading the response, when the next update is due, and — said out loud — how late an update already is. They are never shown who is running it; that is internal staffing.
When to declare
Declare a major incident when the impact is broad or business-critical and needs a
coordinated response rather than a single agent working a queue. Priority-1
incidents are auto-flagged as candidates by policy; the decision to actually
declare stays with a human (needs incidents.major.declare).
Declaring early also protects the requester experience: while a major incident is active, the requester Help Center shows a banner (“we’re already on it”) to suppress duplicate tickets, and new requester-raised tickets that keyword-match the major incident are auto-linked to it.
Declare + assign roles
On the incident, open the major-incident panel and Declare major (or run “Declare major incident — INC-…” from ⌘K). Assign the two roles — each through the person picker: type a name or email and pick the person from your Portal directory. No more hand-typing an identity string: a picked person is a verified directory identity Beacon can notify and report on. Per role you can instead click Assign me (your own signed-in identity).
-
Incident commander — owns the response and the decision-making. Required, and the dialog opens with you already selected — declaring yourself commander is one click. Beacon will not accept a declaration without a commander: the commander is who gets notified that they are running it, so a nameless declaration notified nobody while looking exactly like a real one. If your Portal directory is synced, an unverifiable person gets you a clear error naming the field rather than a silent dead end.
-
Communications lead — owns stakeholder updates. Optional; the cadence reminders fall back to the commander.
Optionally pick the affected service (a business-service CI from the registry). This names the service in the runbook panel and on the requester banner — org-scoped: requesters only see the service name when it’s a tenant-wide CI or they belong to the CI’s organization. Re-declare to change or clear it later.
Both roles are notified on assignment. Re-declaring updates the roles without resetting the declaration clock.
Hand a role over, or leave it empty
Use Update roles in the major-incident panel — the same permission as declaring. The form shows whoever holds each role now; handing over is one pick in the person picker.
-
Change a holder: pick the new person from the directory (or Assign me). The whole identity is replaced atomically — the new holder never inherits any fragment of the old one. A role you don’t touch isn’t sent at all, so nothing can be wiped by accident.
-
Vacate a role: click Vacate role. The role becomes empty and reminders stop going to whoever held it — this is how you remove a comms lead who has left, rather than having them absorb every overdue-comms reminder for the life of the incident.
-
A live major must keep an incident commander. The commander role has no Vacate button while the response is live; hand the role over instead, or stand the incident down first. Once it has stood down, the commander can be cleared like any other role.
Only the people whose role actually changed are notified. The incoming holder is told they now hold it; the outgoing holder is told they’ve been relieved and who has it now (or that it is vacant). Correcting a display name notifies nobody, and re-submitting the same roles unchanged does nothing at all — no notification, no timeline entry, no audit row. Every real change is recorded on the incident timeline and in the audit log with the before and after values.
Check that paging will actually fire
The panel shows a Paging: ready / not ready badge with the reason — for example “Page-on-declare is enabled but no alert channel is deliverable (1 disabled by the delivery circuit breaker after repeated failures) — page the on-call rota manually.” You do not need the alerting-admin permission to see it; the incident-manager role includes the read-only paging check. If it says not ready, page your rota by hand and tell your administrator: a declaration does not (and never did) guarantee a page, and Beacon now says so before you declare instead of only on the timeline afterwards.
If there is no war room
The panel states why rather than showing an empty space: either no policy names a Teams team (nothing to create), or one does but the Teams connection is not bound / creation failed — in which case the incident timeline carries the exact reason. Coordinate on your usual bridge and carry on; nothing about the declaration depends on it.
Run the comms drumbeat
The communications lead posts status updates on a cadence (the shortest cadence across your active MIM policies — 30 minutes by default). Post an update from the major-incident panel, or from ⌘K. Each update:
- is a public note on the incident, so requesters and watchers see it;
- notifies every watcher on the incident;
- emails every stakeholder-list recipient whose list receives updates (2b);
- resets the cadence clock.
The dialog tells you who is about to hear from you — “This goes to 3 watcher(s) on this incident and 9 stakeholder-list recipient(s)” — before you write a word. It used to be titled “Post stakeholder update” while reaching no stakeholder list at all, which meant the one word on the screen named the one audience that was not on the distribution.
If the cadence lapses without an update, Beacon reminds the comms lead (falling back to the commander) and marks the incident overdue on the timeline — the drumbeat is never silently dropped. Posting a real update clears the reminder.
If an update doesn’t reach someone, the incident says so. When the mail
relay gives up on a comms message after its retries, the incident’s timeline
records incident.comms_delivery_failed with the address and the relay’s own
error, and the panel shows an “N undelivered” badge beside the cadence. It
appears only once delivery has genuinely been abandoned — a retry still in
flight is not a failure — and it stays visible after stand-down, because a
message that failed during the response still failed.
a. War room (Teams channel, one click on declare)
If your major-incident policy names a War-room team (the Teams team id in
the relevant workflow → Major incidents → policy) and a Teams connection is bound, the
first declaration creates a dedicated Teams channel named
MI <number> <title> and puts a Join war room link on the major-incident
panel and the timeline.
Prerequisites (both outside Beacon): the Beacon bot’s Azure AD app must hold
the Channel.Create Microsoft Graph application permission with admin
consent in your AAD tenant, and the team id must be a team the app can create
channels in. If either is missing, declaration still succeeds and the timeline
records exactly why no war room was created — fix the consent/team id and
re-declare is not needed for existing majors (create the channel by hand) but
the next major will get one automatically.
b. Stakeholder lists (briefed without becoming watchers)
the relevant workflow → Major incidents → Stakeholder lists (needs
incidents.major.declare). Each list holds plain email recipients (execs,
account managers — they don’t need Beacon accounts) and two filters: an
optional customer organization (blank = fires for every customer’s majors —
lists are additive, unlike policies) and an optional priority threshold
(“P1 only”). On declaration every matched recipient gets the
incident.stakeholder_declared email; on stand-down they get
incident.stakeholder_stood_down; and in between they get each status update
(incident.stakeholder_update). Edit all three texts under the relevant workflow →
Notification templates.
“Receives comms updates” is on by default, and you can switch it off per list. A list that has it off still gets the declare and stand-down notices — use that for an audience who needs only “it started / it’s over”. Before this existed, every list was effectively in that mode: the cadence, the overdue reminder and the post-update dialog all ran for an audience that was never on the distribution.
The audience is fixed when the incident is declared. If you add or remove recipients on a list while an incident is running, the change applies to the next incident — the one in flight keeps the audience it was declared with. That is deliberate: otherwise editing a list would retroactively change who the record says should have been briefed on updates already sent, and “who did we tell?” would stop being answerable.
Stakeholders are still not watchers: they never appear in the ticket’s watcher panel and they receive nothing else about the ticket. Use a watcher for someone following the ticket itself; use a stakeholder list for people who need the briefing by email.
c. Attach child incidents
Reports of the same outage can be attached under the major (major-incident
panel → Child incidents; needs tickets.link.manage). Attached children
show with their status, and — the payoff — are resolved automatically when
the major stands down, each with a public closing note and a requester
notification rendered from the incident.child_resolved template. A child
that can’t legally resolve (e.g. an open gating task) is skipped with an
honest timeline note instead. Children you resolve yourself before stand-down
are left untouched, and a child reopened after stand-down stays open.
Stand down
When service is restored, stand down the major incident (major-incident panel;
needs incidents.major.declare). This clears the major flag and stamps the
stand-down time. The requester banner and auto-linking stop. Within a minute
the scheduler resolves any attached child incidents (see 2c) and every
stakeholder list that was briefed gets the stood-down notice (see 2b). You can
move the incident to resolved at this point — but note that final
closure is still gated on the PIR (next step).
Stand down asks first, and tells you what it is about to do. It used to be a single click that emailed executives and auto-resolved every attached ticket with no confirmation at all. Now it opens a confirmation that names the actual numbers — “auto-resolve 4 attached child incident(s) that are still open”, “email the stand-down notice to 9 stakeholder-list recipient(s)” — because “are you sure?” about an unstated impact is not a safeguard. Only children that are still open are counted: one already resolved is not something stand-down is going to touch.
Resolving stands it down for you (and who gets told)
You do not have to stand down before resolving. If you resolve or close an
incident that is still flagged major, Beacon derives the stand-down from the
resolution rather than refusing the transition — saying the incident is
resolved is already saying the response is over. This matters in practice
because an agent who can transition tickets but does not hold
incidents.major.declare can still settle the incident: they are not trapped,
and the incident does not silently stay “live” on the board.
A derived stand-down does everything a manual one does — clears the flag, stamps
the stand-down time, sends stakeholders the stood-down notice, resolves the
alert pages, arms the child-resolution sweep — and is labelled, not
disguised. The timeline reads “Major incident stood down automatically — the
incident was resolved”, and the audit row carries derived_from with the actor
being whoever resolved.
Because the resolver is often not you, the incident commander and
communications lead each get a notification when a stand-down is derived
this way — “Major incident INC-… stood down automatically” — so you find out
that the response you were running has ended, and that the comms cadence you
were keeping no longer applies. Whoever caused it is not notified (you do not
need telling about your own action), and pressing Stand down yourself sends
nothing, because you already know. The wording is templatable per tenant like
any other notification (the relevant workflow, event
incident.major_stood_down).
One case has no stand-down to derive: an incident flagged major with no
major-incident record behind it (an import, or a flag set through a path that
never opened one). Settling it just clears the flag, and the timeline and audit
say exactly that — incident.major_flag_cleared — rather than claiming a
stand-down that never happened. Stakeholders are not notified in that case,
because there was no briefed audience to close out.
Close with a PIR (mandatory) + auto-spawned RCA
A major incident cannot be closed until its post-incident review is filed —
the state machine blocks incident → closed for any incident that was ever
declared major and has no PIR, and there is no back door (the generic transition
endpoint enforces it too). Standing the incident down ends the response; it does
not retract the declaration or waive the review. A major incident may sit at
resolved before the PIR; only the final closed step is gated.
Because resolving a still-live major stands it down automatically, the normal sequence leaves you on a stood-down incident with the review still owed. The major-incident panel stays on the ticket in that state — the roles, war room, affected service, attached child incidents and the PIR/RCA block are all still there; only the live-response actions (Post update, Stand down) drop away and the update cadence reads update cadence ended. The ticket header carries a PIR outstanding badge for as long as the review is owed, so you find out before you try to close, not from a refusal afterwards.
Submit the PIR from the major-incident panel with:
- a summary (required) of what happened and the impact,
- an optional timeline of events, and
- lessons learned.
On the first PIR submission Beacon automatically:
- spawns a root-cause Problem ticket seeded from your PIR summary,
- links it to the incident both ways, and
- notifies the incident commander.
Re-submitting the PIR updates it in place and does not create a second problem. From there, the RCA follows the problem-management workflow (see needed — raise a linked change. The investigation is never dropped on the floor.
Tuning the policy
In the relevant workflow (needs incidents.major.declare) you set, per
policy:
-
Auto-declare priority — which incidents become MIM candidates (or manual only), and
-
Comms cadence (minutes) — how often the commander/comms lead must post an update. With multiple active policies, the shortest cadence wins.
-
Page alert channels on declare — opt-in: the first declaration pages every active PagerDuty/OpsGenie channel (configured under the relevant workflow) with a
criticalpage. Re-declaring (role updates) never double-pages; stand-down or ticket resolution clears the page. See
Deactivate a policy to disable auto-declaration and cadence reminders for it; a tenant with no active policy has cadence reminders off.
The flow at a glance
Quick reference — who can do what
| Action | Permission |
|---|---|
| View MIM state / active banner | tickets.view |
| Declare / stand down / post comms / submit PIR | incidents.major.declare |
| Manage the MIM policy | incidents.major.declare |