Site Reliability Engineering
Site Reliability Engineering treats reliability as measurable, not a matter of luck. I use service level objectives and error budgets instead of vague uptime targets, building observability in rather than bolting it on after an incident.
Dashboards and alerts reflect what users experience, not just what is easy to measure. I instrument systems well enough to debug at short notice.
Problems this usually starts from
These issues usually prompt a conversation:
- Alerts fire constantly and everyone ignores them.
- Dashboards exist, but nobody trusts them during an incident.
- On-call means guessing the problem instead of being told.
- Nobody agrees on what “reliable enough” means, turning every outage into a debate.
- Logs, metrics, and traces live in separate places that do not connect.
- Observability was bolted on late, missing exactly what is needed mid-incident.
- Postmortems happen, but the same incidents recur regardless.
What an engagement looks like
I start by understanding what is measured today and where the gaps are between the dashboards and the user experience. From there, I agree on SLOs and error budgets that reflect reality, tighten alerting so it is signal rather than noise, and build or consolidate the observability stack (Prometheus, Grafana, Loki, Mimir). An incident at 3am should not start with guessing which dashboard to open.
Scope is either a fixed piece of work (an observability audit, a tool migration) or an ongoing placement. The deliverables stay useful after the engagement ends: reality-reflecting dashboards, actionable alerts, and runbooks documented for the next person on call.
Typical work includes SLOs tied to what users notice instead of easy infrastructure metrics, Grafana dashboards built for incidents rather than general overviews, and Loki/Mimir pipelines that make finding logs a five-second answer.
Scope can start small and grow once the working relationship fits. The engagement and pricing page details how I agree on and price work, including fixed-scope observability audits. The fractional DevOps cost in the UK page compares costs against hiring.
Questions
Remote or on-site?
Long contracts only, or shorter work too?
What if Prometheus/Grafana or a similar stack is already in place?
Do you get involved in on-call itself?
Can I get fractional SRE rather than hiring someone?
Do you work with a specific cloud provider?
Also
- DevOps Engineering — CI/CD, automation, and GitOps — making releases routine instead of risky.
- Platform Engineering — The internal platform other engineers build on, standardised and self-service.
- Cloud Infrastructure — Architecture and provisioning on AWS, defined as code and version-controlled.