Volume 8 · Chapter 8 — Enterprise Operations Runbooks, Standard Operating Procedures (SOPs) & Production Playbooks
What To Do At 2 AM
The day-to-day operational manual — SOP structure, shift operations, and runbooks across Linux, Windows, Kubernetes, databases, applications, cloud, network, and security, plus the maintenance procedures and checklists that keep it all auditable.
Executive Summary
Standard Operating Procedures (SOPs)
Operational excellence depends on consistency. Every engineer, regardless of experience level, follows the same approved procedures. A Standard Operating Procedure ensures that operational activities are repeatable, auditable, measurable, safe, compliant, and predictable.
Every SOP defines
- Objective
- Scope
- Prerequisites
- Required permissions
- Procedure
- Validation
- Rollback (if applicable)
- Documentation updates
SOP structure template
Purpose · Scope · Responsibilities · Prerequisites
Procedure · Validation · Rollback · Escalation
Related Documents · Revision History
Business Objective
Shift operations
Every engineer runs the same start-of-shift checklist, regardless of which customer environment they’re covering.
Start-of-shift checklist
Monitoring platform
- Verify New Relic platform accessibility
- Confirm dashboard availability
- Check alert pipeline
- Validate notification channels
Infrastructure
- Review overnight alerts
- Review unresolved incidents
- Verify backup completion
- Check storage utilization
- Validate certificate expirations
Applications
- Review failed deployments
- Review application availability
- Review synthetic monitoring
- Review transaction failures
Customer communication
- Review previous shift handover
- Review maintenance windows
- Review change calendar
Shift handover checklist
- Open incidents
- Pending changes
- Major risks
- Active monitoring exceptions
- Customer escalations
- Known issues
- Planned maintenance
- Action items
Implementation Guidance
Daily operations playbook
Every business day follows the same defined rhythm.
Morning
- Health verification
- Dashboard review
- Backup verification
- Capacity review
Mid-day
- Incident review
- SLA review
- Customer requests
- Monitoring validation
Evening
- Shift summary
- Pending incident review
- Knowledge updates
- Handover preparation
Standards & Best Practices
Linux operations runbooks
Four runbooks cover the most common Linux production issues an on-call engineer will hit.
Runbook 1 — High CPU utilization
Symptoms
- CPU > 90%
- Slow applications
- High load average
Investigation
- Review CPU metrics in New Relic
- Identify affected processes
- Review recent deployments
- Check scheduled jobs
- Review logs
Resolution
- Restart affected process (if approved)
- Optimize workload
- Scale resources if required
- Escalate if persistent
Validation
- CPU returns to baseline
- Application performance restored
- Alerts cleared
Runbook 2 — Memory utilization
Check
- Free memory
- Cache
- Swap
- OOM events
Possible actions
- Restart application (approved scenarios)
- Investigate memory leaks
- Tune JVM or runtime
- Scale infrastructure where justified
Runbook 3 — Disk full
Investigation
- Identify filesystem
- Review log growth
- Check temporary files
- Review backup retention
Resolution
- Archive logs
- Clean temporary files
- Extend storage if approved
Runbook 4 — Server down
Steps
- Verify monitoring
- Ping server
- Check cloud console or hypervisor
- Review recent changes
- Review console logs
- Engage infrastructure team if required
Standards & Best Practices
Windows Server operations
The same runbook discipline applies to the Windows estate.
- High CPU
- Memory pressure
- Service failures
- IIS issues
- Windows Update validation
- Event Viewer analysis
- Disk utilization
- Active Directory checks
Standards & Best Practices
Kubernetes cluster runbooks
Pod CrashLoopBackOff
Symptoms
- Restart count increasing
- Application unavailable
Procedure
- Review pod logs
- Review deployment changes
- Verify ConfigMaps and Secrets
- Check resource limits
- Verify image versions
- Confirm dependencies
- Redeploy if approved
- Escalate if unresolved
ImagePullBackOff
Check
- Image registry
- Authentication
- Repository permissions
- Image tag
- Network connectivity
Node not ready
Investigate
- Kubelet
- Network
- Disk pressure
- Memory pressure
- Control plane connectivity
PVC issues
Check
- StorageClass
- Persistent Volume status
- CSI driver health
- Capacity
Kamiti Recommendations
Container platform maintenance
Routine maintenance keeps the platform itself from becoming the next incident.
- Node upgrades
- Certificate renewal
- Cluster backup
- Namespace review
- Resource optimization
- Image cleanup
- Security patch validation
Standards & Best Practices
Database operations runbooks
Every supported database engine has its own runbook set, tuned to how that engine actually fails.
Oracle
- Listener down
- Tablespace full
- Archive log full
- ASM health
- Session blocking
- Slow queries
- Backup validation
PostgreSQL
- Replication lag
- Vacuum issues
- Connection exhaustion
- Index fragmentation
- Slow queries
- Backup recovery validation
SQL Server
- TempDB full
- Blocking sessions
- Deadlocks
- Backup validation
- High CPU
- Storage growth
MySQL, MongoDB & Redis
- Replication
- Memory utilization
- Connection limits
- Query optimization
- Failover
- Backup verification
Standards & Best Practices
Application operations runbooks
Java
- JVM heap
- Garbage collection
- Thread dumps
- Heap dumps
- Connection pool exhaustion
- API latency
- External dependency failure
.NET
- CLR memory
- IIS recycling
- API failures
- Authentication errors
- Database connectivity
Node.js & Python
- Event loop analysis
- Async bottlenecks
- Dependency failures
- Worker process health
- Runtime memory usage
Architecture
Cloud operations runbooks
AWS
- EC2 recovery
- Auto Scaling
- ELB health
- EBS volume issues
- IAM validation
- Route53 checks
- CloudWatch integration
Azure
- VM recovery
- Load balancer
- Azure Monitor
- AKS validation
- Storage accounts
- Azure SQL
Google Cloud
- Compute Engine
- GKE
- Cloud SQL
- Load balancer
- Cloud Storage
- IAM
Common Challenges
Network troubleshooting
Network playbooks cover the failure modes that sit underneath every other layer.
- DNS failure
- Firewall rules
- SSL certificate expiry
- VPN failure
- Load balancer health
- Network latency
- Packet loss
Each playbook includes
- Symptoms
- Verification
- Isolation steps
- Resolution
- Escalation criteria
Kamiti Recommendations
Security incident playbooks
Security incidents follow a distinct playbook structure, separate from standard incident management.
- Malware detection
- Ransomware indicators
- Suspicious login activity
- Privileged account misuse
- Certificate compromise
- Credential leakage
- API abuse
- DDoS response
Each playbook defines
- Initial containment
- Evidence preservation
- Notification requirements
- Recovery steps
- Lessons learned
Implementation Guidance
Planned maintenance
Before maintenance
- CAB approval
- Customer notification
- Backup verification
- Rollback validation
- Monitoring readiness
During maintenance
- Execute approved tasks
- Monitor systems
- Record deviations
- Communicate status
After maintenance
- Validate functionality
- Confirm monitoring
- Update documentation
- Obtain customer confirmation
KPI & SLA Examples
Operational checklists
The same checklist discipline repeats at every cadence, from daily to quarterly.
Daily
- Dashboard review
- Backup validation
- Certificate review
- Critical service validation
- Open incident review
- Capacity review
Weekly
- Patch compliance review
- Alert tuning review
- Capacity trend review
- Knowledge article updates
- Automation opportunities
Monthly
- SLA reporting
- RCA review
- Capacity forecast
- Security review
- DR readiness
- CSI review
Quarterly
- Disaster Recovery exercise
- Capacity assessment
- Architecture review
- Tool optimization
- Customer governance meeting
Deliverables
Knowledge base management
Knowledge articles are peer-reviewed and linked to relevant incidents or problems — not written once and forgotten.
Each article includes
- Issue summary
- Environment
- Symptoms
- Root cause
- Resolution
- Validation
- References
- Version history
Deliverables
What Volume 8 produces
Standard Operating Procedures
- 100+ SOP documents
- Shift operations guide
- Maintenance procedures
- Operational standards
Runbooks
- Linux runbooks
- Windows runbooks
- Kubernetes runbooks
- Cloud runbooks
- Database runbooks
- Application runbooks
- Network runbooks
- Security playbooks
Checklists
- Daily operations
- Weekly maintenance
- Monthly governance
- Quarterly DR validation
- Annual platform health review
Operational templates
- Shift handover report
- Maintenance checklist
- Incident action log
- RCA template
- Knowledge article template
- Operational audit checklist
Consultant Tips
Consultant’s note
Runbooks are living documents
The value of a runbook lies in its clarity and consistency. A well-designed runbook enables engineers of varying experience levels to respond confidently during high-pressure situations, reducing mean time to resolution and minimizing operational risk.
As systems, platforms, and customer requirements evolve, runbooks should be treated as living documents — regularly reviewed, tested, and improved through operational experience.