Risk incidents, containment, and learning
Respond to active harm and preserve an accurate account of events.
The loss graph bends upward at 2 a.m. The first job is to stop the bleeding without breaking the rest of the payment system. The second is to understand what happened well enough that the next team does not repeat it.
Declare the incident from impact
An incident can involve active fraud, a broken sanctions feed, duplicate payouts, missing notices, data exposure, or a liquidity shortfall. Define severity from customer harm, financial exposure, legal duties, service impact, and uncertainty.
Use a clear declaration process with an incident lead and decision roles. Preserve the initial evidence and assumptions. Early estimates will change; version them rather than treating the first number as permanent truth. A small confirmed issue with large unknown exposure may need more attention than its current loss count suggests.
An incident can exist before the service is fully unavailable. If a release approves payments without a required check, every endpoint may respond quickly while the platform creates new exposure. Declare incidents from customer, financial, security, and control impact as well as infrastructure health. A shared severity definition helps teams recognize this class of failure early.
The initial scope will often be uncertain. Record the known start time, suspected affected population, systems involved, and the evidence supporting each statement. Update these facts as the investigation improves. A confident estimate without a reproducible population query can send recovery teams toward the wrong accounts and make later reconciliation harder.
Inside the mechanism. Declare an incident from the affected capability and consequence. A small population can justify urgent response if the effect is severe or continuing. Establish known customers, actions, amounts, control failures, and uncertainties. Use an incident owner and a shared source of current facts. Global error percentages can hide concentrated harm and should not be the only declaration trigger.
A concrete example. A low technical error rate can conceal concentrated customer harm or unreconciled money. Declaration should use the affected service and consequence, not only infrastructure alerts. The daily source population is 22,000 items, but 440 are outside the completed monitoring run. The included population creates 453 hits and 371 unique cases. With 82 cases already open and capacity for 430, the queue closes at 23. Coverage, duplicate work, and staffing are separate causes; reducing one number does not prove that the overall control improved.
When the assumption fails. A small group of duplicate debits remains below the global error-rate threshold. Identify affected customers, actions, amounts, and required controls and appoint an incident owner. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.
A low technical error rate can conceal concentrated customer harm or unreconciled money. Declaration should use the affected service and consequence, not only infrastructure alerts.
- DetectIdentify the abnormal outcome
- AssessEstimate harm exposure and uncertainty
- DeclareAssign command and response roles
- Confirmed loss
- Evidence supports the amount
- Potential exposure
- Additional impact remains possible
Incident scope
Illustrative data; not a real customer record or a prescribed policy.
- Confirmed duplicates4
Known affected payments
- Unreviewed batch2000
Potential broader scope
- Severityuncertainty included
Not based only on four cases
Known loss may be only part of the event
Include uncertainty in severity and scope. Known loss may be only part of the event.
- Failure mode 1avoid
- Wait for perfect facts before organizing. Harm can continue.
- Failure mode 2avoid
- Use the first estimate forever. Evidence develops.
- Failure mode 3avoid
- Assign several competing incident leads. Decision authority becomes unclear.
Contain the harmful capability
Containment should target the path causing harm: pause a payout route, restrict a compromised credential, stop a defective rule, or hold work under the approved contingency. Avoid a broad shutdown when a narrower action can safely stop the problem.
Record the authority, time, scope, and expected effect of each action. Confirm that it worked using direct evidence. A kill switch that changes a dashboard flag but leaves workers active is not containment. Consider customer consequences and legal constraints while preserving the records needed for investigation.
Inside the mechanism. Contain the harmful capability at the narrowest effective boundary. Preserve read access and evidence needed to classify unknown outcomes where possible. Record who authorized the action, its scope, and the conditions for reversal. A broad shutdown can create new recovery work if it destroys visibility into actions already in flight. Containment and financial correction are separate workstreams.
A concrete example. Containment should stop the action causing harm while preserving evidence and essential service where possible. The scope depends on the failure mechanism. The case identifies 2,728 eligible records from a source population of 3,100. The required workflow completes for 2,646, but 40 completed records miss the illustrative internal target. Another 82 remain incomplete. Communication evidence covers 2,620 generated notices. Scope, completion, timeliness, and delivery are four separate properties of the customer outcome.
When the assumption fails. A broad shutdown obscures unknown payment outcomes and prevents customer support from seeing them. Disable the affected capability, retain read access, and record the authority and rollback conditions. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.
Containment should stop the action causing harm while preserving evidence and essential service where possible. The scope depends on the failure mechanism.
- TargetIdentify the harmful path
- ActApply an authorized bounded restriction
- ConfirmObserve that the harmful effect stopped
- Command issued
- Containment was requested
- Containment effective
- Evidence shows the path is stopped
Payout containment
Illustrative data; not a real customer record or a prescribed policy.
- Switchdisabled
Control command
- Worker releasesstill occurring
Action ineffective
- Responsestop actual release path
Verify the downstream effect
A control flag may not stop running workers
Verify containment at the effect boundary. A control flag may not stop running workers.
- Failure mode 1avoid
- Assume the command succeeded. The action needs evidence.
- Failure mode 2avoid
- Delete records to stop processing. That destroys the investigation trail.
- Failure mode 3avoid
- Ignore customer obligations during a pause. They still require handling.
Maintain a decision and evidence log
An incident log records observations, decisions, owners, timestamps, and evidence references. Keep facts separate from hypotheses. Use a common time basis and preserve source timestamps when they differ.
Record why a decision was reasonable with the information available then. This reduces hindsight distortion during review. Limit sensitive content to the appropriate access group and use references in broad coordination channels. The log should support handoff when responders change shifts and should make unresolved questions visible.
Inside the mechanism. Keep a chronological log with observed facts, hypotheses, decisions, owners, and next evidence. Correct disproven hypotheses explicitly rather than rewriting history. Material decisions need the information available at the time, not only the explanation formed afterward. Use protected references for sensitive artifacts. The log supports coordination during the incident and a reliable review after it.
A concrete example. An incident log separates observed facts, working hypotheses, decisions, and open tasks. That distinction matters when new evidence disproves the first explanation. The case identifies 650 eligible records from a source population of 670. The required workflow completes for 630, but 9 completed records miss the illustrative internal target. Another 20 remain incomplete. Communication evidence covers 624 generated notices. Scope, completion, timeliness, and delivery are four separate properties of the customer outcome.
When the assumption fails. An early hypothesis is copied into customer-impact reporting as a confirmed cause. Timestamp evidence and decisions, name owners, and preserve corrections without rewriting history. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.
An incident log separates observed facts, working hypotheses, decisions, and open tasks. That distinction matters when new evidence disproves the first explanation.
- ObserveRecord facts with sources
- DecideCapture rationale authority and time
- HandoffShow unresolved work and next owners
- Hypothesis
- Possible explanation under investigation
- Established fact
- Supported observation or confirmed result
Incident log
Illustrative data; not a real customer record or a prescribed policy.
- 02:10duplicate releases observed
Fact
- 02:15retry defect suspected
Hypothesis
- 02:20route paused and verified
Action with evidence
The record must show what was known when
Version facts hypotheses and decisions. The record must show what was known when.
- Failure mode 1avoid
- Rewrite early notes as if the cause was obvious. That loses the real decision context.
- Failure mode 2avoid
- Paste secrets into broad channels. Coordination does not remove access limits.
- Failure mode 3avoid
- Handoff without open questions. The next team can miss critical work.
Recover with a reconciled population
Recovery should restore safe service and resolve the affected obligations. Identify all potentially affected records with a reproducible query, confirm impact, and assign the appropriate remediation. Keep uncertain records in scope until resolved.
Use idempotent correction jobs and independent reconciliation. A refund or reversal batch can itself create duplicate value if retried poorly. Resume service under defined guardrails and monitor the original failure mode. The incident is not complete merely because the graph returns to normal while historical customers remain affected.
Recovery is complete when the affected population has a defined and verified outcome. Some payments may need replay, some may already have succeeded, and some may require manual resolution or customer communication. Use stable identifiers to assign each item to a resolution state and verify totals against independent records. Do not replay an entire time range solely because the service was degraded during it. An outage window is a useful investigation boundary, but it does not prove that every action within the window failed.
Inside the mechanism. Recovery requires a complete affected population with known success, known failure, and unknown outcomes distinguished. Join stable business keys to external records and ledger effects before replay. An action that succeeded externally during a timeout may need reconciliation, not another execution. Verify corrected balances, required controls, customer communications, and remaining exceptions before declaring financial recovery.
A concrete example. Recovery must classify every affected financial action and resolve unknown outcomes before replay. Restoring throughput does not repair duplicate or missing postings. The batch begins with $1,360,000 of instructions and $1,305,600.00 of captured value. At the observation cutoff, $39,168.00 remains pending. After the stated refunds, fees, and restrictions, $1,056,230.40 is available for payout. The unresolved instruction count is 2; an unknown external result is handled separately from a known decline.
When the assumption fails. A bulk retry repeats actions that succeeded externally before the outage. Join original action identifiers to external and ledger evidence and repair only the supported exceptions. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.
Recovery must classify every affected financial action and resolve unknown outcomes before replay. Restoring throughput does not repair duplicate or missing postings.
- PopulationIdentify and confirm affected records
- RemediateExecute controlled corrections
- ResumeVerify safe service and remaining obligations
- Service recovery
- New activity works again
- Impact resolution
- Affected historical records have supported outcomes
Recovery population
Illustrative data; not a real customer record or a prescribed policy.
- Confirmed affected180
Reviewed scope
- Corrected176
Completed actions
- Unresolved4
Incident work remains
Restored service does not resolve past impact
Reconcile remediation against the affected population. Restored service does not resolve past impact.
- Failure mode 1avoid
- Run unkeyed correction scripts repeatedly. They can create new duplicates.
- Failure mode 2avoid
- Drop uncertain records from the count. The scope becomes falsely small.
- Failure mode 3avoid
- Close at the first normal chart. Historical obligations may remain.
Learn from the system that allowed the event
A useful incident review explains the trigger, contributing conditions, detection, response, and customer impact. Avoid stopping at a person made a mistake. Examine why the system permitted the error and why existing controls did not catch it sooner.
Choose corrective actions that change the failure path and can be verified. Assign owners, dates, and acceptance evidence. Track completion and retest important controls. A long list of reminders can be less useful than one enforced invariant. The review should improve the system’s ability to prevent, detect, or contain the next similar event.
Inside the mechanism. A useful review explains the conditions that allowed harm, the missing detection, and the limits of the response. Avoid ending with a reminder to be careful. Assign changes to the relevant contract, control, capacity, or operating process and define evidence of effectiveness. Re-test the failure mechanism and reconcile any remaining customer remediation. Closure should mean the agreed behavior is demonstrated.
A concrete example. A useful review explains how the system permitted and failed to detect harm. Remediation needs an owner and a check that the changed behavior addresses the mechanism. The case identifies 989 eligible records from a source population of 1,150. The required workflow completes for 959, but 14 completed records miss the illustrative internal target. Another 30 remain incomplete. Communication evidence covers 949 generated notices. Scope, completion, timeliness, and delivery are four separate properties of the customer outcome.
When the assumption fails. The review ends with a reminder to be careful and no test of the failed control. Change the relevant contract or control and verify the same failure no longer escapes detection. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.
A useful review explains how the system permitted and failed to detect harm. Remediation needs an owner and a check that the changed behavior addresses the mechanism.
- ExplainTrace trigger and contributing conditions
- RepairChoose enforceable changes
- VerifyRetest the relevant failure path
- Individual action
- What a person did
- System condition
- Why that action could cause or escape harm
Post-incident action
Illustrative data; not a real customer record or a prescribed policy.
- Causeduplicate posting allowed
Missing invariant
- Fixunique operation constraint
Enforced boundary
- Evidenceretry fault test passes
Verifiable acceptance
Lessons should change behavior in the system
Choose repairs with observable acceptance evidence. Lessons should change behavior in the system.
- Failure mode 1avoid
- End at human error. Contributing design conditions remain.
- Failure mode 2avoid
- Assign reminders without owners. The work may not happen.
- Failure mode 3avoid
- Close after code review alone. The failure path still needs verification.
Chapter connections
This chapter builds on Operational resilience and failure design. Continue with A complete risk system: the Lantern case to follow the next part of the system. Use the glossary for terminology and risk mathematics for formulas and worked calculations.