Insights
Fractional SRE in the UK: What Fits in a Few Days a Month, and What Doesn't
“Fractional SRE” means buying site reliability engineering in pieces instead of hiring a full-time SRE. What the term leaves out is that SRE is two different jobs. One is building the measurement: service level objectives, the alerts that fire when you’re burning through them, and the observability stack underneath. The other is running the rota: being the person who gets paged at 2am.
The first job fits a few days a month well. The second doesn’t, by its nature: it needs someone reachable on the days that weren’t bought. Most of what I’d want to know before buying fractional SRE comes down to which of the two you actually need.
The build side fits in days
An SLO is a target for something users notice. Your monitoring measures it continuously, not a person. “99.9% of API requests succeed over 30 days” is an SLO. Once defined and wired up, it keeps measuring on the days nobody works on it. That’s what makes this half of SRE suit part-time work: the output is configuration that keeps running on its own.
Pieces of work that fit in a day or a few days:
- Defining the first SLOs for one service or user journey: which requests count, what counts as a failure, and what target the team can live with.
- Alerting on the SLO instead of on CPU, memory and disk, so the page arrives when users are affected.
- An alert clean-up. The rule I use: every alert either leads to an action or gets deleted. Alerts that fire daily and get ignored are worse than no alert, because they teach everyone to ignore the pager.
- Runbooks linked from the alerts, so whoever is paged starts with what to check first.
- Dashboards built for the incident, showing the SLO and its burn rate on top, with the detail underneath.
- Consolidating collection, for example moving from Promtail to Grafana Alloy, so logs and metrics arrive through one pipeline with the same labels.
None of these needs someone present every day. Each one ends with something in Git that keeps working after the work is finished.
Alerting on the error budget, in practice
A 99.9% target over 30 days leaves an error budget of 0.1%, which is 43.2 minutes of total failure a month. Alerting on how fast you’re spending that budget is the part of SRE I’d start with, because it replaces a pile of threshold alerts with a small number that mean something.
The approach below is the multi-window, multi-burn-rate alert from the
Google SRE Workbook. It pages when the budget is
burning fast enough to matter, and uses a short window alongside the long one so the alert clears
soon after the problem stops. These are Prometheus rules for an HTTP service whose request counter
carries a status-code label; change the metric and label names to match yours. The rules pass
promtool check rules, and the alerts pass promtool test rules, on Prometheus 3.15.0.
groups:
- name: slo-api-availability
rules:
# Error ratio over each window the alerts need.
- record: slo:sli_error:ratio_rate5m
expr: |
sum(rate(http_requests_total{job="api", code=~"5.."}[5m]))
/
sum(rate(http_requests_total{job="api"}[5m]))
- record: slo:sli_error:ratio_rate30m
expr: |
sum(rate(http_requests_total{job="api", code=~"5.."}[30m]))
/
sum(rate(http_requests_total{job="api"}[30m]))
- record: slo:sli_error:ratio_rate1h
expr: |
sum(rate(http_requests_total{job="api", code=~"5.."}[1h]))
/
sum(rate(http_requests_total{job="api"}[1h]))
- record: slo:sli_error:ratio_rate6h
expr: |
sum(rate(http_requests_total{job="api", code=~"5.."}[6h]))
/
sum(rate(http_requests_total{job="api"}[6h]))
# 99.9% SLO: the error budget is 0.001.
# Fast burn: 2% of a 30-day budget gone in an hour.
- alert: APIErrorBudgetFastBurn
expr: |
slo:sli_error:ratio_rate1h > (14.4 * 0.001)
and
slo:sli_error:ratio_rate5m > (14.4 * 0.001)
labels:
severity: page
annotations:
summary: "API is spending its error budget about 14 times faster than it can afford"
runbook_url: "https://example.internal/runbooks/api-error-budget"
# Slow burn: 5% of the budget gone in six hours.
- alert: APIErrorBudgetSlowBurn
expr: |
slo:sli_error:ratio_rate6h > (6 * 0.001)
and
slo:sli_error:ratio_rate30m > (6 * 0.001)
labels:
severity: page
annotations:
summary: "API is spending its error budget 6 times faster than it can afford"
runbook_url: "https://example.internal/runbooks/api-error-budget"A burn rate of 14.4 over an hour spends 2% of a 30-day budget; a rate of 6 over six hours spends 5%. The workbook’s third tier, a burn rate of 1 over three days, suits a ticket, not a page.
Two things are worth deciding before copying this. First, whether 5xx responses are the right definition of failure: a fast 404 on a missing page probably isn’t one, and a 200 that took eight seconds probably is. Second, who the page goes to. The rules don’t care whether that’s a full-time engineer or a rota of your developers, and in a fractional arrangement it will be one of those, not the person who wrote the rules.
Where this goes wrong on small services
I’ve hit this one myself. Burn-rate alerting assumes enough traffic that a percentage means something. On a quiet service it doesn’t. At 240 requests an hour, four failures is 1.7% of the hour, which is over the 1.44% fast-burn threshold. If one of the twenty or so requests in the last five minutes also failed, both windows are over the line and it pages, for four failed requests that no user would have noticed as an outage.
There are a few ways round it, and they trade off differently:
Require a minimum number of requests before the alert can fire. The simplest fix, and the one I’d reach for first:
expr: | ( slo:sli_error:ratio_rate1h > (14.4 * 0.001) and slo:sli_error:ratio_rate5m > (14.4 * 0.001) ) and sum(increase(http_requests_total{job="api"}[1h])) > 1000At 240 requests an hour this stays quiet. At 6,000 an hour with 2% failing, it still pages. Pick the threshold from your own traffic, not from this example.
Use longer windows for low-volume services, so a few failures are a smaller share. The cost is slower detection.
Send synthetic traffic at a steady rate, so there is always a baseline. It works, but now the SLO is partly measuring your probe rather than your users.
None of these is free. Which one fits depends on how quiet the service is and how much a slow alert would cost you.
The rota doesn’t fit
On-call is presence. Someone has to be reachable, with access and context, on whichever day something breaks. A few days a month, by definition, doesn’t cover the other days. Buying fractional SRE and expecting it to include the pager gets you either a gap in cover or an arrangement that has quietly become a part-time job with a standby clause.
What part-time work can do is make your own on-call less miserable. Fewer, better alerts. A runbook behind each one. Dashboards that answer “is it us?” in the first minute. Those change what it’s like to be the developer on the rota, even though the rota itself stays yours.
If you need someone else holding the pager around the clock, that’s a managed service with more than one engineer behind it, a hire, or a longer placement where joining the rota is part of the role. Some UK providers sell fractional infrastructure with 24/7 cover included; that’s a different product from buying days, and the difference belongs in the contract. I cover the same question for DevOps work generally in what fractional DevOps costs in the UK.
What it costs against hiring
ITJobsWatch put the UK median advertised salary for a permanent site reliability engineer at £80,000, from 135 adverts in the six months to 26 September 2026. Add employer’s National Insurance at 15% above £5,000 and the minimum employer pension, 3% of earnings between £6,240 and £50,270, and that’s £92,571 a year before equipment, benefits or a recruitment fee.
Contract rates for the same role are lower per day than direct rates: the ITJobsWatch UK median for an SRE contract was £530 a day, from 79 rates, over the same period. UK providers who publish direct day rates sit between about £1,000 and £1,200; they’re listed with sources in the cost article, which also explains why the two differ and has a calculator for your own salary figure, days and rate.
Four days a month at £1,100, the middle of that published range rather than any one provider’s price, is £52,800 a year. That buys the build side described above. It doesn’t buy anyone on call, and it doesn’t buy someone who owns the systems between visits.
When fractional SRE is the right shape
It fits when a team has developers, runs its own services, and has nobody whose job is reliability: the alerts are noisy, nobody has agreed what “reliable enough” means, and each outage turns into an argument about whether it counts. The work there is mostly definition and configuration, and it finishes.
It doesn’t fit when the real need is someone watching production all day, or when the systems are changing so fast that SLOs would need rewriting every week. The first needs a hire or a managed service. The second needs the architecture to settle before measuring it is worth the effort.
I cover how I work on SLOs, alerting and the Prometheus, Grafana, Loki and Mimir stack on the site reliability engineering page. I detail pricing and engagement options on the engagement and pricing page.