Skip to content
Kamiti Labs
The Playbook

Volume 8 · Chapter 8Enterprise Operations Runbooks, Standard Operating Procedures (SOPs) & Production Playbooks

What To Do At 2 AM

The day-to-day operational manual — SOP structure, shift operations, and runbooks across Linux, Windows, Kubernetes, databases, applications, cloud, network, and security, plus the maintenance procedures and checklists that keep it all auditable.

01

Executive Summary

Standard Operating Procedures (SOPs)

Operational excellence depends on consistency. Every engineer, regardless of experience level, follows the same approved procedures. A Standard Operating Procedure ensures that operational activities are repeatable, auditable, measurable, safe, compliant, and predictable.

Every SOP defines

  • Objective
  • Scope
  • Prerequisites
  • Required permissions
  • Procedure
  • Validation
  • Rollback (if applicable)
  • Documentation updates

SOP structure template

Document ID · Version · Author · Reviewer · Approval Date

Purpose · Scope · Responsibilities · Prerequisites

Procedure · Validation · Rollback · Escalation

Related Documents · Revision History
02

Business Objective

Shift operations

Every engineer runs the same start-of-shift checklist, regardless of which customer environment they’re covering.

Start-of-shift checklist

Monitoring platform

  • Verify New Relic platform accessibility
  • Confirm dashboard availability
  • Check alert pipeline
  • Validate notification channels

Infrastructure

  • Review overnight alerts
  • Review unresolved incidents
  • Verify backup completion
  • Check storage utilization
  • Validate certificate expirations

Applications

  • Review failed deployments
  • Review application availability
  • Review synthetic monitoring
  • Review transaction failures

Customer communication

  • Review previous shift handover
  • Review maintenance windows
  • Review change calendar

Shift handover checklist

  • Open incidents
  • Pending changes
  • Major risks
  • Active monitoring exceptions
  • Customer escalations
  • Known issues
  • Planned maintenance
  • Action items
03

Implementation Guidance

Daily operations playbook

Every business day follows the same defined rhythm.

Morning

  • Health verification
  • Dashboard review
  • Backup verification
  • Capacity review

Mid-day

  • Incident review
  • SLA review
  • Customer requests
  • Monitoring validation

Evening

  • Shift summary
  • Pending incident review
  • Knowledge updates
  • Handover preparation
04

Standards & Best Practices

Linux operations runbooks

Four runbooks cover the most common Linux production issues an on-call engineer will hit.

Runbook 1 — High CPU utilization

Symptoms

  • CPU > 90%
  • Slow applications
  • High load average

Investigation

  • Review CPU metrics in New Relic
  • Identify affected processes
  • Review recent deployments
  • Check scheduled jobs
  • Review logs

Resolution

  • Restart affected process (if approved)
  • Optimize workload
  • Scale resources if required
  • Escalate if persistent

Validation

  • CPU returns to baseline
  • Application performance restored
  • Alerts cleared

Runbook 2 — Memory utilization

Check

  • Free memory
  • Cache
  • Swap
  • OOM events

Possible actions

  • Restart application (approved scenarios)
  • Investigate memory leaks
  • Tune JVM or runtime
  • Scale infrastructure where justified

Runbook 3 — Disk full

Investigation

  • Identify filesystem
  • Review log growth
  • Check temporary files
  • Review backup retention

Resolution

  • Archive logs
  • Clean temporary files
  • Extend storage if approved

Runbook 4 — Server down

Steps

  • Verify monitoring
  • Ping server
  • Check cloud console or hypervisor
  • Review recent changes
  • Review console logs
  • Engage infrastructure team if required
05

Standards & Best Practices

Windows Server operations

The same runbook discipline applies to the Windows estate.

  • High CPU
  • Memory pressure
  • Service failures
  • IIS issues
  • Windows Update validation
  • Event Viewer analysis
  • Disk utilization
  • Active Directory checks
06

Standards & Best Practices

Kubernetes cluster runbooks

Pod CrashLoopBackOff

Symptoms

  • Restart count increasing
  • Application unavailable

Procedure

  • Review pod logs
  • Review deployment changes
  • Verify ConfigMaps and Secrets
  • Check resource limits
  • Verify image versions
  • Confirm dependencies
  • Redeploy if approved
  • Escalate if unresolved

ImagePullBackOff

Check

  • Image registry
  • Authentication
  • Repository permissions
  • Image tag
  • Network connectivity

Node not ready

Investigate

  • Kubelet
  • Network
  • Disk pressure
  • Memory pressure
  • Control plane connectivity

PVC issues

Check

  • StorageClass
  • Persistent Volume status
  • CSI driver health
  • Capacity
07

Kamiti Recommendations

Container platform maintenance

Routine maintenance keeps the platform itself from becoming the next incident.

  • Node upgrades
  • Certificate renewal
  • Cluster backup
  • Namespace review
  • Resource optimization
  • Image cleanup
  • Security patch validation
08

Standards & Best Practices

Database operations runbooks

Every supported database engine has its own runbook set, tuned to how that engine actually fails.

Oracle

  • Listener down
  • Tablespace full
  • Archive log full
  • ASM health
  • Session blocking
  • Slow queries
  • Backup validation

PostgreSQL

  • Replication lag
  • Vacuum issues
  • Connection exhaustion
  • Index fragmentation
  • Slow queries
  • Backup recovery validation

SQL Server

  • TempDB full
  • Blocking sessions
  • Deadlocks
  • Backup validation
  • High CPU
  • Storage growth

MySQL, MongoDB & Redis

  • Replication
  • Memory utilization
  • Connection limits
  • Query optimization
  • Failover
  • Backup verification
09

Standards & Best Practices

Application operations runbooks

Java

  • JVM heap
  • Garbage collection
  • Thread dumps
  • Heap dumps
  • Connection pool exhaustion
  • API latency
  • External dependency failure

.NET

  • CLR memory
  • IIS recycling
  • API failures
  • Authentication errors
  • Database connectivity

Node.js & Python

  • Event loop analysis
  • Async bottlenecks
  • Dependency failures
  • Worker process health
  • Runtime memory usage
10

Architecture

Cloud operations runbooks

AWS

  • EC2 recovery
  • Auto Scaling
  • ELB health
  • EBS volume issues
  • IAM validation
  • Route53 checks
  • CloudWatch integration

Azure

  • VM recovery
  • Load balancer
  • Azure Monitor
  • AKS validation
  • Storage accounts
  • Azure SQL

Google Cloud

  • Compute Engine
  • GKE
  • Cloud SQL
  • Load balancer
  • Cloud Storage
  • IAM
11

Common Challenges

Network troubleshooting

Network playbooks cover the failure modes that sit underneath every other layer.

  • DNS failure
  • Firewall rules
  • SSL certificate expiry
  • VPN failure
  • Load balancer health
  • Network latency
  • Packet loss

Each playbook includes

  • Symptoms
  • Verification
  • Isolation steps
  • Resolution
  • Escalation criteria
12

Kamiti Recommendations

Security incident playbooks

Security incidents follow a distinct playbook structure, separate from standard incident management.

  • Malware detection
  • Ransomware indicators
  • Suspicious login activity
  • Privileged account misuse
  • Certificate compromise
  • Credential leakage
  • API abuse
  • DDoS response

Each playbook defines

  • Initial containment
  • Evidence preservation
  • Notification requirements
  • Recovery steps
  • Lessons learned
13

Implementation Guidance

Planned maintenance

Before maintenance

  • CAB approval
  • Customer notification
  • Backup verification
  • Rollback validation
  • Monitoring readiness

During maintenance

  • Execute approved tasks
  • Monitor systems
  • Record deviations
  • Communicate status

After maintenance

  • Validate functionality
  • Confirm monitoring
  • Update documentation
  • Obtain customer confirmation
14

KPI & SLA Examples

Operational checklists

The same checklist discipline repeats at every cadence, from daily to quarterly.

Daily

  • Dashboard review
  • Backup validation
  • Certificate review
  • Critical service validation
  • Open incident review
  • Capacity review

Weekly

  • Patch compliance review
  • Alert tuning review
  • Capacity trend review
  • Knowledge article updates
  • Automation opportunities

Monthly

  • SLA reporting
  • RCA review
  • Capacity forecast
  • Security review
  • DR readiness
  • CSI review

Quarterly

  • Disaster Recovery exercise
  • Capacity assessment
  • Architecture review
  • Tool optimization
  • Customer governance meeting
15

Deliverables

Knowledge base management

Knowledge articles are peer-reviewed and linked to relevant incidents or problems — not written once and forgotten.

Each article includes

  • Issue summary
  • Environment
  • Symptoms
  • Root cause
  • Resolution
  • Validation
  • References
  • Version history

Deliverables

What Volume 8 produces

Standard Operating Procedures

  • 100+ SOP documents
  • Shift operations guide
  • Maintenance procedures
  • Operational standards

Runbooks

  • Linux runbooks
  • Windows runbooks
  • Kubernetes runbooks
  • Cloud runbooks
  • Database runbooks
  • Application runbooks
  • Network runbooks
  • Security playbooks

Checklists

  • Daily operations
  • Weekly maintenance
  • Monthly governance
  • Quarterly DR validation
  • Annual platform health review

Operational templates

  • Shift handover report
  • Maintenance checklist
  • Incident action log
  • RCA template
  • Knowledge article template
  • Operational audit checklist

Consultant Tips

Consultant’s note

Runbooks are living documents

The value of a runbook lies in its clarity and consistency. A well-designed runbook enables engineers of varying experience levels to respond confidently during high-pressure situations, reducing mean time to resolution and minimizing operational risk.

As systems, platforms, and customer requirements evolve, runbooks should be treated as living documents — regularly reviewed, tested, and improved through operational experience.