Skip to content
Kamiti Labs
The Playbook

Volume 7 · Chapter 7Enterprise New Relic Implementation, Instrumentation & Observability Engineering

The Implementation Guide

The technical build standard engineers follow end to end — account governance and tagging, agent deployment, per-language instrumentation, OpenTelemetry, Kubernetes and database monitoring, dashboards and alerting, access governance, and the phased production rollout.

01

Executive Summary

Enterprise observability strategy

Modern enterprise environments consist of distributed applications running across on-premises data centers, Kubernetes clusters, public cloud platforms, SaaS services, databases, APIs, and mobile applications. Monitoring these components individually creates operational silos.

The objective of enterprise observability is to provide end-to-end visibility, unified telemetry, business transaction monitoring, root cause identification, AI-assisted incident management, executive reporting, and continuous reliability improvement.

Observability pillars

1. Metrics

  • CPU
  • Memory
  • JVM heap
  • API response time
  • Kubernetes node utilization

2. Logs

  • Application logs
  • System logs
  • Security logs
  • Audit logs
  • Database logs

3. Distributed traces

  • API Gateway
  • Microservices
  • Database
  • Cache
  • Message queue

4. Events

  • Deployments
  • Auto scaling
  • Configuration changes
  • Business transactions
  • Login events

5. Business KPIs

  • Orders
  • Payments
  • Dealer logins
  • Inventory updates
  • Manufacturing jobs

Enterprise observability stack

  1. 1

    Business KPIs

  2. 2

    Applications

  3. 3

    Containers

  4. 4

    Operating systems

  5. 5

    Cloud infrastructure

  6. 6

    Telemetry collection

  7. 7

    New Relic platform

  8. 8

    Dashboards, alerts, AI & automation

02

Architecture

New Relic enterprise architecture

Logical architecture

  1. 1

    Users, applications, servers, databases, containers & cloud

  2. 2

    New Relic agents

  3. 3

    OpenTelemetry collector

  4. 4

    Secure telemetry pipeline

  5. 5

    New Relic cloud

  6. 6

    Dashboards, alerts, AI & reports

Enterprise components

  • New Relic One Platform
  • Infrastructure Agent
  • APM Agent
  • Browser Agent
  • Mobile Agent
  • Kubernetes integration
  • OpenTelemetry Collector
  • Log Forwarder
  • NRQL Engine
  • AI Monitoring
  • Alert Engine
03

Standards & Best Practices

Enterprise account governance

Large organizations separate environments at the account level, and every monitored entity carries standardized metadata.

Account hierarchy

Enterprise Account
  • Production
  • UAT
  • QA
  • Development
  • Sandbox

Tagging standards

TagExample
Business unitManufacturing
EnvironmentProduction
ApplicationDealer Portal
OwnerSales IT
Support groupKamiti SRE
CriticalityP1
Cost centerFMCG-SALES
RegionSingapore

Naming convention

<Environment>-<BusinessUnit>-<Application>-<Component>

PROD-SALES-DEALERPORTAL-API01

Consistent naming improves searchability, reporting, and automation.

04

Implementation Guidance

Infrastructure agent deployment

Linux

  • RHEL
  • Rocky Linux
  • Ubuntu
  • Debian
  • Oracle Linux
  • SUSE

Windows

  • Windows Server 2019
  • Windows Server 2022

Cloud

  • AWS EC2
  • Azure VM
  • Google Compute Engine

Deployment process

  1. 1

    Validate OS compatibility.

  2. 2

    Create API keys.

  3. 3

    Install Infrastructure Agent.

  4. 4

    Configure entity tags.

  5. 5

    Verify connectivity.

  6. 6

    Confirm metrics ingestion.

  7. 7

    Validate dashboards.

  8. 8

    Enable alert policies.

Post-installation checklist

  • Host visible in New Relic
  • CPU metrics available
  • Memory metrics available
  • Disk metrics available
  • Network metrics available
  • Entity correctly tagged
  • Alerts functioning
05

Standards & Best Practices

APM instrumentation standards

Application Performance Monitoring is standardized across every supported technology in the estate.

Java

  • JVM heap
  • Garbage collection
  • Thread pools
  • Response time
  • Exceptions
  • Transactions
  • Database calls
  • External services

.NET

  • CLR performance
  • Memory usage
  • API latency
  • SQL queries
  • Exceptions

Node.js

  • Event loop delay
  • API response
  • MongoDB
  • Redis
  • Async operations

Python

  • Flask
  • Django
  • FastAPI
  • Celery
  • Database queries

Go

  • Goroutines
  • Garbage collection
  • Memory
  • API performance

PHP

  • Laravel
  • Symfony
  • WordPress
  • Drupal
06

Implementation Guidance

OpenTelemetry integration

Kamiti Labs standardizes telemetry collection using OpenTelemetry wherever feasible.

Architecture

  1. 1

    Application

  2. 2

    OTel SDK

  3. 3

    OTel Collector

  4. 4

    New Relic Exporter

  5. 5

    New Relic Platform

Benefits

  • Vendor-neutral instrumentation
  • Consistent telemetry
  • Multi-cloud support
  • Future portability
  • Reduced instrumentation effort
07

Architecture

Kubernetes observability

Cluster monitoring

  • Control plane
  • API server
  • etcd
  • Scheduler
  • Controller manager

Worker nodes

  • CPU
  • Memory
  • Network
  • Storage
  • Kernel health

Pods

  • Restart count
  • CrashLoopBackOff
  • Pending pods
  • Image pull errors
  • Resource consumption

Workloads

  • Deployments
  • StatefulSets
  • DaemonSets
  • Jobs
  • CronJobs

Services

  • Ingress
  • Service mesh
  • DNS
  • Load balancers
08

Standards & Best Practices

Database monitoring

Every production database requires dedicated monitoring, tuned to its engine.

Oracle

  • Sessions
  • Tablespaces
  • Wait events
  • ASM
  • Archive logs

PostgreSQL

  • Connections
  • Locks
  • Replication lag
  • Vacuum health
  • Query latency

SQL Server

  • TempDB
  • Blocking
  • Deadlocks
  • Buffer cache

MySQL

  • InnoDB
  • Replication
  • Query performance

MongoDB

  • Replica sets
  • Query execution
  • Index efficiency

Redis

  • Memory usage
  • Hit ratio
  • Evictions
  • Replication
09

Architecture

Messaging & streaming observability

Modern enterprises depend on asynchronous communication — every broker and queue is monitored, not just the services that call them.

Apache Kafka

  • Broker health
  • Consumer lag
  • Topic throughput
  • Partition balance

RabbitMQ

  • Queue depth
  • Consumer status
  • Message rate

ActiveMQ / IBM MQ

  • Queue health
  • Dead letter queues
  • Channel status
10

KPI & SLA Examples

Browser & mobile monitoring

Browser monitoring

  • Page load time
  • Largest Contentful Paint (LCP)
  • First Input Delay (FID)
  • Cumulative Layout Shift (CLS)
  • JavaScript errors
  • Session duration

Mobile monitoring

  • App launch time
  • API performance
  • Crash rate
  • Network latency
  • Device distribution
  • OS versions
11

Implementation Guidance

Synthetic monitoring

Synthetic monitoring validates application availability before users experience issues. Tests execute from multiple geographic locations where applicable.

  • Homepage availability
  • Login journey
  • Dealer portal
  • Payment flow
  • API endpoints
  • Search functionality
12

Deliverables

Dashboard engineering standards

Dashboards are role-specific — the same underlying telemetry, presented differently for each audience.

Executive dashboard

  • SLA compliance
  • Availability
  • Revenue impact
  • Major incidents

Operations dashboard

  • Infrastructure health
  • Active alerts
  • Service status

Application dashboard

  • Transactions
  • Errors
  • Latency
  • JVM metrics

DBA dashboard

  • Slow queries
  • Blocking
  • Storage
  • Replication

Kubernetes dashboard

  • Cluster health
  • Pod health
  • Resource usage
  • Autoscaling
13

Common Challenges

Alert engineering

Alerts are actionable, prioritized, noise-free, and business-aware — anything less erodes trust in the platform.

Every alert includes

  • Severity
  • Trigger
  • Threshold
  • Notification targets
  • Runbook link
  • Escalation path

Review alert effectiveness regularly to reduce false positives.

14

Standards & Best Practices

Security & access governance

Access is granted on least-privilege principles, reviewed periodically.

RolePermissions
AdministratorPlatform administration
Observability EngineerConfigure monitoring
Operations EngineerView dashboards, acknowledge alerts
Application OwnerView application metrics
ExecutiveRead-only business dashboards
AuditorReporting & audit access

Enable SSO and MFA where supported, and review access periodically.

15

Kamiti Recommendations

Production rollout methodology

  1. 1

    Phase 1 — Pilot applications

    5–10 services, dashboard validation, alert tuning.

  2. 2

    Phase 2 — Business-critical applications

    ERP, WMS, Dealer Portal, APIs.

  3. 3

    Phase 3 — Enterprise expansion

    Remaining applications, databases, infrastructure, Kubernetes, mobile, browser, business KPIs.

  4. 4

    Phase 4 — Operational optimization

    AI correlation, automation, capacity planning, reliability engineering.

16

Deliverables

Validation & acceptance

Before handover, every layer of the platform is validated against a fixed checklist.

Infrastructure

  • All hosts reporting
  • Metrics complete
  • Alerts tested

Applications

  • APM data flowing
  • Traces visible
  • Transactions captured

Databases

  • Health metrics visible
  • Query monitoring active

Dashboards

  • Executive dashboards approved
  • Operational dashboards validated

Operations

  • Alert routing tested
  • Runbooks linked
  • On-call procedures verified

Deliverables

What Volume 7 produces

Architecture

  • Enterprise observability architecture
  • Telemetry standards
  • Agent deployment guide
  • OpenTelemetry design

Engineering

  • Instrumentation standards
  • Dashboard catalogue
  • Alert catalogue
  • NRQL query library
  • Tagging standards
  • Naming standards

Operations

  • Production rollout plan
  • Validation checklist
  • Acceptance test report
  • Monitoring coverage report
  • Operational handover package

Consultant Tips

Consultant’s note

Quality of telemetry, not quantity of dashboards

A successful observability platform is defined not by the number of dashboards created but by the quality, consistency, and operational value of the telemetry it produces. Standardized instrumentation, governance, tagging, dashboard design, and validation ensure that engineers, operations teams, and business stakeholders all rely on the same trusted operational data.

This consistency is what enables faster troubleshooting, informed decision-making, and long-term service reliability.