§ 08 · Domain mastery

Observability & SRE

Signal that survives 3 AM. Postmortems that end in code changes, not blame.

SLOs and error budgets your product team helps set, dashboards a CFO can read, on-call rotations with clear ownership, and blameless postmortems that actually change the codebase. Cost and reliability treated as the same conversation, not siloed OKRs.

Talk to a partnerAll services
§ I

Pillars

01

SLOs tied to user journeys

Availability of endpoints is not availability of your product. SLOs written against what the user is trying to do, with error budgets that inform release cadence.

02

OpenTelemetry as the standard

One vendor-neutral instrumentation across languages. When you change backends, you're not rewriting instrumentation for three months.

03

On-call that respects sleep

Rotations sized for the incident load, runbooks that pass a real drill, and a paging discipline that treats false positives as bugs.

04

Blameless postmortems, actioned

The template is the easy part. The habit that turns action items into shipped code is what separates practice from theater.

05

Cost + reliability as one lens

The cheapest way to make a system reliable is often to remove things. Cost signals in the reliability conversation, and vice versa.

06

Chaos and drills, not hope

Game days for the top failure modes, DR drills that actually flip traffic, load tests before Black Friday — not after.

§ II

Domain vocabulary

The words that separate insiders from readers.

Error budgetThe allowed unreliability in an SLO. Teams that treat it as a resource ship faster than those that don't.
SLI vs. SLOThe metric vs. the target. Confusing them makes 'we hit our SLO' meaningless.
Burn rate alertAlerting on trajectory to breach, not on breach itself. The difference between paging early and paging useless.
MTTRMean time to recover. Widely quoted, easily gamed. Insiders pair it with mean time between incidents.
Postmortem cultureNot the doc — the follow-through. Teams that ship the action items outperform teams that write beautiful docs.
Golden signalLatency, traffic, errors, saturation. Naming which one you're alerting on is a mark of clarity.
OTel collectorThe vendor-neutral pipeline for telemetry. Where you win back leverage from any single observability vendor.
Blast radiusShared with cloud vocab; the SRE version is about traffic routing and cell isolation.
§ III

Open problems we help with

§ IV

What we ship

  1. 01SLOs tied to user journeys with error budgets.
  2. 02OpenTelemetry instrumentation across your services.
  3. 03Runbooks tested in game days.
  4. 04On-call rotations sized to the real load.
  5. 05A postmortem template plus the habit to close its action items.
  6. 06A cost + reliability dashboard the whole org reads.
NEXT STEP
How does this look in your case?
Talk to a partner
Season