Services

Site Reliability Engineering

Site Reliability Engineering treats reliability as measurable, not a matter of luck. I use service level objectives and error budgets instead of vague uptime targets, building observability in rather than bolting it on after an incident.

Dashboards and alerts reflect what users experience, not just what is easy to measure. I instrument systems well enough to debug at short notice.

Problems this usually starts from

These issues usually prompt a conversation:

  • Alerts fire constantly and everyone ignores them.
  • Dashboards exist, but nobody trusts them during an incident.
  • On-call means guessing the problem instead of being told.
  • Nobody agrees on what “reliable enough” means, turning every outage into a debate.
  • Logs, metrics, and traces live in separate places that do not connect.
  • Observability was bolted on late, missing exactly what is needed mid-incident.
  • Postmortems happen, but the same incidents recur regardless.

What an engagement looks like

I start by understanding what is measured today and where the gaps are between the dashboards and the user experience. From there, I agree on SLOs and error budgets that reflect reality, tighten alerting so it is signal rather than noise, and build or consolidate the observability stack (Prometheus, Grafana, Loki, Mimir). An incident at 3am should not start with guessing which dashboard to open.

Scope is either a fixed piece of work (an observability audit, a tool migration) or an ongoing placement. The deliverables stay useful after the engagement ends: reality-reflecting dashboards, actionable alerts, and runbooks documented for the next person on call.

Typical work includes SLOs tied to what users notice instead of easy infrastructure metrics, Grafana dashboards built for incidents rather than general overviews, and Loki/Mimir pipelines that make finding logs a five-second answer.

Scope can start small and grow once the working relationship fits. The engagement and pricing page details how I agree on and price work, including fixed-scope observability audits. The fractional DevOps cost in the UK page compares costs against hiring.

Questions

Remote or on-site?

Either — depends on the engagement. Comfortable with a regular on-site day where required.

Long contracts only, or shorter work too?

Both. That includes urgent, one-off issues — a service that’s gone down, traffic causing everything to slow to a crawl, an incident that needs someone on it now — not just fixed-scope audits or longer placements. No running contract required to get started on something like that.

What if Prometheus/Grafana or a similar stack is already in place?

Common, and usually the starting point rather than something to replace — the work is more often tightening what’s there than ripping it out.

Do you get involved in on-call itself?

Joining a rotation fits a longer placement better than fractional work, where by default nobody covers the days in between. More often the work is making on-call less miserable: better alerts, better runbooks. If on-call is part of what you need, say so at the start and I’ll tell you what I can take on.

Can I get fractional SRE rather than hiring someone?

Yes, for the build-and-fix side: SLOs, alerting, dashboards and runbooks, bought a piece at a time or as a block of days each month. On-call cover isn’t included by default. How it’s scoped and priced.

Do you work with a specific cloud provider?

Most commonly AWS, but the observability stack — Prometheus, Grafana, Loki, Mimir — isn’t tied to one provider and carries over to whatever’s actually in use.

Get in touch

Also

  • DevOps Engineering — CI/CD, automation, and GitOps — making releases routine instead of risky.
  • Platform Engineering — The internal platform other engineers build on, standardised and self-service.
  • Cloud Infrastructure — Architecture and provisioning on AWS, defined as code and version-controlled.