Site reliability engineering

Observability and SLO engineering

A method for turning user journeys into measurable service indicators, budgets and incident evidence without confusing a wall of telemetry with reliability.

15 minute technical paperApplied engineering paper
Technical note 1

Start with the user action

A method for turning user journeys into measurable service indicators, budgets and incident evidence without confusing a wall of telemetry with reliability.

A service-level objective should describe a result users care about. The API process being up is not enough if authentication fails, a queue stalls or a payment provider rejects every request. Begin with a small set of user journeys: submit an application, confirm a transaction, receive a shipment update, load an operator queue. For each journey, define success, failure, valid traffic and the time boundary.

The denominator matters. Including malformed requests can make a healthy service look unreliable. Excluding every inconvenient dependency failure can make a broken journey look healthy. Define which requests the service accepts responsibility for, how planned maintenance is treated and whether a result that arrives after the user gives up counts as success. Write examples. Two teams reading the definition should classify the same event the same way.

Name the user journey and its responsible service boundary.
Define valid events, successful events and the measurement window.
Keep exclusions narrow, explicit and reviewable.
Link every objective to an owner and a response policy.