Skip to content
Kamiti Labs
The Playbook

Volume 5 · Chapter 5Enterprise Incident Management, Problem Management, Change Management & ITIL Operations Framework

How Production Support Actually Works

The ITIL-aligned procedures engineers follow every day — incident classification and priority, major incident management and war rooms, problem management and the KEDB, change management and the CAB, release control, rollback, disaster recovery, and the runbook library behind it all.

01

Executive Summary

The ITIL service lifecycle

Enterprise operations are not simply about fixing issues. Every production event follows a controlled lifecycle — detect, classify, respond, restore, analyze, improve. This structured approach minimizes business impact while ensuring every incident contributes to continuous improvement.

Operational lifecycle

  1. 1

    Monitoring

  2. 2

    Alert generated

  3. 3

    Incident created

  4. 4

    Investigation

  5. 5

    Resolution

  6. 6

    Validation

  7. 7

    Closure

  8. 8

    RCA

  9. 9

    Knowledge base

  10. 10

    Automation

02

Business Objective

Incident management framework

Restore normal business operations as quickly as possible while minimizing disruption.

Notice the wording. The objective is not to find the root cause immediately. The objective is to restore service. Root cause analysis comes later.

Incident sources

  • New Relic alerts
  • Customer complaints
  • Help desk tickets
  • Synthetic monitoring failures
  • Infrastructure monitoring
  • Security monitoring
  • Business transaction failures
  • Cloud provider alerts
  • Database alerts
  • API failures

Incident lifecycle

  1. 1

    Alert

  2. 2

    Ticket creation

  3. 3

    Incident classification

  4. 4

    Priority assignment

  5. 5

    Assignment

  6. 6

    Investigation

  7. 7

    Workaround

  8. 8

    Service restoration

  9. 9

    Validation

  10. 10

    Closure

03

Standards & Best Practices

Incident classification

Every incident is categorized consistently, regardless of who logs it or which shift picks it up.

Infrastructure

  • CPU
  • Memory
  • Storage
  • Network
  • Server
  • Kubernetes node

Application

  • Java
  • .NET
  • Node.js
  • API
  • Authentication
  • Microservices

Database

  • Oracle
  • SQL Server
  • PostgreSQL
  • MongoDB
  • Redis

Cloud

  • AWS
  • Azure
  • GCP
  • Load balancer
  • Storage
  • IAM

Security

  • Certificate
  • Firewall
  • Access
  • MFA
  • Identity
  • VPN

Business

  • Orders
  • Payments
  • Inventory
  • Dealer portal
  • Customer login
04

Common Challenges

Priority matrix

Priority is a function of business impact and urgency — not how loudly someone is asking for a fix.

ImpactUrgencyPriority
HighHighP1
HighMediumP2
MediumMediumP3
LowLowP4

P1 — Critical

  • ERP down
  • Manufacturing stops
  • Complete production outage
  • Payment gateway down
  • Kubernetes cluster failure

Immediate actions

  • Major Incident declared
  • War Room activated
  • SDM informed
  • Customer notified
  • Executive updates initiated

P2

  • Slow application
  • Database replication issue
  • Kubernetes node unavailable (service still operational)

P3

  • Report generation failure
  • Individual module unavailable
  • Non-critical API issue

P4

  • Cosmetic issue
  • Documentation correction
  • Minor enhancement request
05

Implementation Guidance

Major Incident Management (MIM)

Major incidents require a dedicated operational process — not an ad hoc scramble.

Major incident flow

  1. 1

    P1 detected

  2. 2

    Major Incident declared

  3. 3

    War Room started

  4. 4

    Technical teams engaged

  5. 5

    Business communication

  6. 6

    Recovery

  7. 7

    Validation

  8. 8

    Closure

  9. 9

    RCA meeting

War Room participants

Customer

  • CIO
  • Infrastructure Manager
  • Application Owner
  • Business Owner

Kamiti Labs

  • Service Delivery Manager
  • SRE Lead
  • Technical Account Manager
  • Cloud Engineer
  • DBA
  • Application Engineer
06

Kamiti Recommendations

Root Cause Analysis (RCA)

An RCA should not seek blame. It should identify what happened, why it happened, why it wasn’t detected earlier, how it was resolved, and how recurrence can be prevented.

RCA template

  • Incident summary
  • Timeline
  • Business impact
  • Technical findings
  • Root cause
  • Corrective actions
  • Preventive actions
  • Lessons learned
  • Automation opportunities
07

Standards & Best Practices

Problem management

Some incidents happen repeatedly. Those become problems — tracked and closed separately from the individual incidents they keep spawning.

Workflow

  1. 1

    Recurring incident

  2. 2

    Problem record

  3. 3

    Investigation

  4. 4

    Known error

  5. 5

    Permanent fix

  6. 6

    Verification

  7. 7

    Closure

Objectives

  • Remove recurring failures
  • Improve reliability
  • Reduce incident volume
  • Improve customer satisfaction
08

Deliverables

Known Error Database (KEDB)

A searchable repository of recurring issues means an L1 engineer at 2am doesn’t have to rediscover a fix that’s already known.

ErrorCauseWorkaroundPermanent fix
JVM heap fullMemory leakRestart serviceCode fix
Disk fullLog growthCleanupLog rotation
Redis timeoutNetworkRestartConnection pool tuning
09

Common Challenges

Change management

Not every production modification carries the same level of risk. Changes are categorized and approved accordingly.

Standard change

Low risk, pre-approved, routine, repeatable.

  • Restart service
  • Certificate renewal
  • Scheduled log cleanup

Normal change

Requires review and approval.

  • Version upgrade
  • Database schema update
  • Kubernetes deployment
  • Firewall modification

Emergency change

Implemented to restore service during a major incident.

  • Hotfix deployment
  • Immediate rollback
  • Security patch for active vulnerability

Reviewed retrospectively after implementation.

Change lifecycle

  1. 1

    RFC raised

  2. 2

    Risk assessment

  3. 3

    CAB review

  4. 4

    Approval

  5. 5

    Implementation

  6. 6

    Validation

  7. 7

    Closure

10

Standards & Best Practices

Change Advisory Board (CAB)

The CAB evaluates

  • Business impact
  • Technical risk
  • Rollback readiness
  • Implementation schedule
  • Customer communication
  • Resource availability

Typical participants

  • Service Delivery Manager
  • Technical Lead
  • Application Owner
  • Infrastructure Lead
  • Customer representative
11

Implementation Guidance

Release management

Production releases follow a controlled process across three phases.

Pre-release

  • Code freeze
  • Testing complete
  • Security review
  • Backup verification
  • Rollback prepared

Deployment

  • Deployment window
  • Smoke testing
  • Health checks
  • Monitoring validation

Post-release

  • Business validation
  • Performance observation
  • Incident monitoring
  • Formal release closure
12

Kamiti Recommendations

Rollback strategy

Every deployment includes a tested rollback plan — a rollback plan that hasn’t been tested isn’t a plan.

  • Previous application version available
  • Database rollback plan
  • Configuration backup
  • Infrastructure snapshot (where applicable)
  • Rollback owner assigned
  • Business approval process documented
13

Architecture

Disaster Recovery (DR)

Disaster Recovery planning defines the objectives and procedures that get a business back online after a catastrophic failure.

  • Recovery Time Objective (RTO)
  • Recovery Point Objective (RPO)
  • Backup strategy
  • Secondary site readiness
  • Failover procedures
  • Communication plan
  • Recovery validation
  • DR testing schedule

DR activation workflow

  1. 1

    Disaster declared

  2. 2

    DR team activated

  3. 3

    Customer approval

  4. 4

    Failover

  5. 5

    Validation

  6. 6

    Business confirmation

  7. 7

    Operations continue

14

Deliverables

Knowledge management

Every operational activity should enrich the organization’s knowledge, not just close a ticket. Articles are version-controlled and reviewed periodically to keep them accurate.

Knowledge articles should include

  • Symptoms
  • Root cause
  • Resolution steps
  • Screenshots
  • Logs
  • Commands
  • Validation steps
  • Related incidents
  • Preventive recommendations
15

Deliverables

Enterprise runbook library

Every critical service has a runbook — organized by domain, so the right procedure is always one search away.

Infrastructure

  • CPU utilization
  • Memory exhaustion
  • Disk full
  • Server down

Kubernetes

  • Pod CrashLoopBackOff
  • Node not ready
  • ImagePullBackOff
  • Persistent volume issues

Application

  • Java heap
  • API timeout
  • Authentication failure
  • Payment failure

Database

  • Deadlock
  • Replication delay
  • Slow query
  • Backup failure

Cloud

  • Auto scaling failure
  • IAM access issue
  • Load balancer health
  • DNS failure

Each runbook defines

  • Trigger
  • Impact
  • Diagnostic steps
  • Resolution
  • Validation
  • Escalation criteria

Deliverables

What Volume 5 produces

At the conclusion of this volume, Kamiti Labs delivers the complete set of operational documents that govern day-to-day production support.

  • Incident Management SOP
  • Major Incident Playbook
  • RCA template
  • Problem Management process
  • KEDB structure
  • Change Management SOP
  • CAB charter
  • Release checklist
  • Rollback checklist
  • DR runbook
  • Knowledge base standards
  • Enterprise runbook library

Consultant Tips

Consultant’s note

Confidence comes from consistency, not zero incidents

Mature managed service providers are distinguished not by the absence of incidents, but by the consistency and professionalism with which they respond. Standardized incident handling, disciplined change management, well-documented runbooks, and a culture of continuous learning build customer confidence and reduce operational risk.

These practices are especially important in FMCG environments, where production, logistics, and customer-facing systems often operate under tight business timelines and high availability expectations.