Incident Management
Incident management is the operational practice of restoring the agreed service state as quickly as possible after a disruption. Problem management is deliberately kept separate: it removes the underlying cause permanently — a restart can resolve an incident without touching the problem. The flow: logging, categorisation, prioritisation by impact and urgency, diagnosis, escalation, resolution, closure. Common metrics are time to acknowledge, time to restore and first-level resolution rate.
Incident Management in practice
Two special cases break the standard flow. A major incident affects business-critical services, is led by a named person and runs on its own communication line to the business and to management; technical resolution is only half the job, reliable information the other half. A security incident follows its own rules, because evidence preservation and reporting duties are added.
Under Article 33 GDPR, a personal data breach must be reported to the supervisory authority within 72 hours of becoming aware of it. For significant security incidents, NIS2 — transposed in Germany in Section 32 BSIG — requires a three-stage notification: an early warning within 24 hours of awareness, a notification with an initial assessment within 72 hours, and a final report no later than one month after that 72-hour notification.
These clocks start on awareness, not on resolution — meeting them requires decision paths, access to the reporting portal and notification templates written down in advance. Escalation runs in two directions: functionally toward deeper expertise, hierarchically toward decision authority, for instance to approve taking a system offline. After closure, every major incident deserves a review that records the timeline, the cause and the effectiveness of countermeasures, and hands open items to problem management or change enablement. Without that feedback loop the same incident recurs and the metrics document activity rather than improvement.