Support service
Hear about incidents before your customers do
When customers are the first to report an outage and nobody measures how long recovery takes, reliability is luck. We connect metrics, logs and traces, define SLOs and set up an incident practice.
- Duration
- 3 to 8 weeks
- Model
- Build and hand over
- Output
- SLOs, clean alerts and incident runbooks
Signs reliability depends on luck
Observability gaps stay invisible until an incident. Then they decide how long the incident lasts.
- Customers or the support team notice outages before your monitoring does.
- Time to recover from an incident is not measured, so nobody knows whether it is getting better.
- The on-call engineer receives so many alerts that the important ones get lost.
- Logs exist, but finding the requests behind one customer’s complaint takes hours.
- A request crosses several services, and nobody can see where the time was spent.
- Incidents end with a fix but no written review, so the same failure comes back.
- Nobody can say what reliable enough means for the product, so every outage turns into an argument.
Reliability you can measure
Monitoring usually grows by accident: a dashboard for every incident, an alert for every metric, logs in several formats. The result is a lot of data and very few answers when something breaks.
We start from the user. Service level objectives define what good looks like for each important journey, such as login, checkout or an API call, and error budgets turn that into a rule for when to slow down and fix. Alerts are rebuilt around those objectives, so an alert means a user is affected.
Underneath, metrics, structured logs and distributed traces are connected, so an engineer can go from an alert to the failing request in minutes. Incident roles, runbooks and blameless reviews complete the loop, so each incident makes the next one shorter.
Scope
Included
- Metrics, structured logging and distributed tracing based on OpenTelemetry
- Service level indicators and objectives for key user journeys
- Error budget policy
- Alert review and rebuild around SLOs
- Dashboards for services and user journeys
- Incident roles, severity levels and communication templates
- Runbooks for the most common failures
- Blameless post-incident review practice
Not included
- Running your on-call rotation
- Buying and administering a commercial monitoring suite
- Application performance tuning
- Round-the-clock operations as an outsourced service
How the work runs
- Week 1Map what exists. Current monitoring, alerts, logs and incident history are reviewed, and the key user journeys are listed together with the people who own them.
- Weeks 2 to 4Instrument and define. Tracing and structured logs are added where they are missing, and SLOs are defined and agreed for each key journey.
- Weeks 4 to 6Clean up the alerts. Alerts are rebuilt around the SLOs and noisy ones are removed. Dashboards show journeys, not just servers.
- Final weeksPractise incidents. Incident roles, runbooks and the review practice are introduced and rehearsed with your team in a drill.
What you receive
Everything runs on your infrastructure and is owned by your team.
Who this service is not for
- Teams looking for someone to carry their pager. We build the practice; your team runs it.
- Systems that have no production users yet. SLOs need real traffic to be meaningful.
- Organisations that want dashboards for their own sake. Every chart we add answers a question someone asks during incidents.
Frequently asked questions
Do we need a commercial monitoring product?
Not necessarily. Open standards and open-source tools cover most needs, and the managed services of your cloud provider fill the gaps. If you already use a commercial product, we build on it.
What is an SLO in practice?
A target such as 99.9 percent of checkout requests succeeding within one second over 28 days. It tells the team when reliability needs attention and when it is safe to ship faster.
How long until alerts become useful?
Usually within the first weeks, once the noisiest alerts are removed and the first SLO-based alerts are in place.
Do you run our incident reviews?
We run the first ones with your team and hand over the format. Reviews work best when they are run by the people who operate the system.
How is this different from Performance and Scalability?
This service builds the visibility and the practices that keep the system healthy. Performance and Scalability finds and removes specific performance limits.
Start with a short technical call
Thirty minutes. You describe your last serious incident, and we tell you what would have shortened it.
