Volume 5 · Chapter 5 — Enterprise Incident Management, Problem Management, Change Management & ITIL Operations Framework
How Production Support Actually Works
The ITIL-aligned procedures engineers follow every day — incident classification and priority, major incident management and war rooms, problem management and the KEDB, change management and the CAB, release control, rollback, disaster recovery, and the runbook library behind it all.
Executive Summary
The ITIL service lifecycle
Enterprise operations are not simply about fixing issues. Every production event follows a controlled lifecycle — detect, classify, respond, restore, analyze, improve. This structured approach minimizes business impact while ensuring every incident contributes to continuous improvement.
Operational lifecycle
- 1
Monitoring
- 2
Alert generated
- 3
Incident created
- 4
Investigation
- 5
Resolution
- 6
Validation
- 7
Closure
- 8
RCA
- 9
Knowledge base
- 10
Automation
Business Objective
Incident management framework
Restore normal business operations as quickly as possible while minimizing disruption.
Notice the wording. The objective is not to find the root cause immediately. The objective is to restore service. Root cause analysis comes later.
Incident sources
- New Relic alerts
- Customer complaints
- Help desk tickets
- Synthetic monitoring failures
- Infrastructure monitoring
- Security monitoring
- Business transaction failures
- Cloud provider alerts
- Database alerts
- API failures
Incident lifecycle
- 1
Alert
- 2
Ticket creation
- 3
Incident classification
- 4
Priority assignment
- 5
Assignment
- 6
Investigation
- 7
Workaround
- 8
Service restoration
- 9
Validation
- 10
Closure
Standards & Best Practices
Incident classification
Every incident is categorized consistently, regardless of who logs it or which shift picks it up.
Infrastructure
- CPU
- Memory
- Storage
- Network
- Server
- Kubernetes node
Application
- Java
- .NET
- Node.js
- API
- Authentication
- Microservices
Database
- Oracle
- SQL Server
- PostgreSQL
- MongoDB
- Redis
Cloud
- AWS
- Azure
- GCP
- Load balancer
- Storage
- IAM
Security
- Certificate
- Firewall
- Access
- MFA
- Identity
- VPN
Business
- Orders
- Payments
- Inventory
- Dealer portal
- Customer login
Common Challenges
Priority matrix
Priority is a function of business impact and urgency — not how loudly someone is asking for a fix.
| Impact | Urgency | Priority |
|---|---|---|
| High | High | P1 |
| High | Medium | P2 |
| Medium | Medium | P3 |
| Low | Low | P4 |
P1 — Critical
- ERP down
- Manufacturing stops
- Complete production outage
- Payment gateway down
- Kubernetes cluster failure
Immediate actions
- Major Incident declared
- War Room activated
- SDM informed
- Customer notified
- Executive updates initiated
P2
- Slow application
- Database replication issue
- Kubernetes node unavailable (service still operational)
P3
- Report generation failure
- Individual module unavailable
- Non-critical API issue
P4
- Cosmetic issue
- Documentation correction
- Minor enhancement request
Implementation Guidance
Major Incident Management (MIM)
Major incidents require a dedicated operational process — not an ad hoc scramble.
Major incident flow
- 1
P1 detected
- 2
Major Incident declared
- 3
War Room started
- 4
Technical teams engaged
- 5
Business communication
- 6
Recovery
- 7
Validation
- 8
Closure
- 9
RCA meeting
War Room participants
Customer
- CIO
- Infrastructure Manager
- Application Owner
- Business Owner
Kamiti Labs
- Service Delivery Manager
- SRE Lead
- Technical Account Manager
- Cloud Engineer
- DBA
- Application Engineer
Kamiti Recommendations
Root Cause Analysis (RCA)
An RCA should not seek blame. It should identify what happened, why it happened, why it wasn’t detected earlier, how it was resolved, and how recurrence can be prevented.
RCA template
- Incident summary
- Timeline
- Business impact
- Technical findings
- Root cause
- Corrective actions
- Preventive actions
- Lessons learned
- Automation opportunities
Standards & Best Practices
Problem management
Some incidents happen repeatedly. Those become problems — tracked and closed separately from the individual incidents they keep spawning.
Workflow
- 1
Recurring incident
- 2
Problem record
- 3
Investigation
- 4
Known error
- 5
Permanent fix
- 6
Verification
- 7
Closure
Objectives
- Remove recurring failures
- Improve reliability
- Reduce incident volume
- Improve customer satisfaction
Deliverables
Known Error Database (KEDB)
A searchable repository of recurring issues means an L1 engineer at 2am doesn’t have to rediscover a fix that’s already known.
| Error | Cause | Workaround | Permanent fix |
|---|---|---|---|
| JVM heap full | Memory leak | Restart service | Code fix |
| Disk full | Log growth | Cleanup | Log rotation |
| Redis timeout | Network | Restart | Connection pool tuning |
Common Challenges
Change management
Not every production modification carries the same level of risk. Changes are categorized and approved accordingly.
Standard change
Low risk, pre-approved, routine, repeatable.
- Restart service
- Certificate renewal
- Scheduled log cleanup
Normal change
Requires review and approval.
- Version upgrade
- Database schema update
- Kubernetes deployment
- Firewall modification
Emergency change
Implemented to restore service during a major incident.
- Hotfix deployment
- Immediate rollback
- Security patch for active vulnerability
Reviewed retrospectively after implementation.
Change lifecycle
- 1
RFC raised
- 2
Risk assessment
- 3
CAB review
- 4
Approval
- 5
Implementation
- 6
Validation
- 7
Closure
Standards & Best Practices
Change Advisory Board (CAB)
The CAB evaluates
- Business impact
- Technical risk
- Rollback readiness
- Implementation schedule
- Customer communication
- Resource availability
Typical participants
- Service Delivery Manager
- Technical Lead
- Application Owner
- Infrastructure Lead
- Customer representative
Implementation Guidance
Release management
Production releases follow a controlled process across three phases.
Pre-release
- Code freeze
- Testing complete
- Security review
- Backup verification
- Rollback prepared
Deployment
- Deployment window
- Smoke testing
- Health checks
- Monitoring validation
Post-release
- Business validation
- Performance observation
- Incident monitoring
- Formal release closure
Kamiti Recommendations
Rollback strategy
Every deployment includes a tested rollback plan — a rollback plan that hasn’t been tested isn’t a plan.
- Previous application version available
- Database rollback plan
- Configuration backup
- Infrastructure snapshot (where applicable)
- Rollback owner assigned
- Business approval process documented
Architecture
Disaster Recovery (DR)
Disaster Recovery planning defines the objectives and procedures that get a business back online after a catastrophic failure.
- Recovery Time Objective (RTO)
- Recovery Point Objective (RPO)
- Backup strategy
- Secondary site readiness
- Failover procedures
- Communication plan
- Recovery validation
- DR testing schedule
DR activation workflow
- 1
Disaster declared
- 2
DR team activated
- 3
Customer approval
- 4
Failover
- 5
Validation
- 6
Business confirmation
- 7
Operations continue
Deliverables
Knowledge management
Every operational activity should enrich the organization’s knowledge, not just close a ticket. Articles are version-controlled and reviewed periodically to keep them accurate.
Knowledge articles should include
- Symptoms
- Root cause
- Resolution steps
- Screenshots
- Logs
- Commands
- Validation steps
- Related incidents
- Preventive recommendations
Deliverables
Enterprise runbook library
Every critical service has a runbook — organized by domain, so the right procedure is always one search away.
Infrastructure
- CPU utilization
- Memory exhaustion
- Disk full
- Server down
Kubernetes
- Pod CrashLoopBackOff
- Node not ready
- ImagePullBackOff
- Persistent volume issues
Application
- Java heap
- API timeout
- Authentication failure
- Payment failure
Database
- Deadlock
- Replication delay
- Slow query
- Backup failure
Cloud
- Auto scaling failure
- IAM access issue
- Load balancer health
- DNS failure
Each runbook defines
- Trigger
- Impact
- Diagnostic steps
- Resolution
- Validation
- Escalation criteria
Deliverables
What Volume 5 produces
At the conclusion of this volume, Kamiti Labs delivers the complete set of operational documents that govern day-to-day production support.
- Incident Management SOP
- Major Incident Playbook
- RCA template
- Problem Management process
- KEDB structure
- Change Management SOP
- CAB charter
- Release checklist
- Rollback checklist
- DR runbook
- Knowledge base standards
- Enterprise runbook library
Consultant Tips
Consultant’s note
Confidence comes from consistency, not zero incidents
Mature managed service providers are distinguished not by the absence of incidents, but by the consistency and professionalism with which they respond. Standardized incident handling, disciplined change management, well-documented runbooks, and a culture of continuous learning build customer confidence and reduce operational risk.
These practices are especially important in FMCG environments, where production, logistics, and customer-facing systems often operate under tight business timelines and high availability expectations.