Skip to content

Operations & Observability

Overview

Operations and observability standards require production services to expose their behaviour and support operational response and recovery.

Choose a standard below to read it in full.

Standards

Structured Logging

A service's log output is structured, traceable, and free of sensitive data or unnecessary detail.

Read more.

Metrics, Monitoring & Alerting

A service's metrics remain accurate, visible, and bounded, with automated and actionable alerts.

Read more.

Distributed Tracing

Trace context propagates end-to-end through accurately structured spans and deliberate sampling.

Read more.

Telemetry Instrumentation

Telemetry remains consistent across services, safe to introduce, and delivered alongside the change it observes.

Read more.

Observability Platform Integration

Telemetry is centralised, portable, and available for resilient operations.

Read more.

Backup & Disaster Recovery

Recovery is designed and validated against defined objectives.

Read more.

Runbooks

A runbook documents and validates the steps to resolve known failures and perform high-risk procedures.

Read more.