What is MTTR (Mean Time to Resolution)?

What is MTTR (Mean Time to Resolution)?

MTTR (Mean Time to Resolution) is the average time from when an incident is first detected to when it is fully resolved and the system returns to normal operation. It is a core KPI in site reliability engineering (SRE): teams use it to measure and improve the efficiency of their incident response.

A lower MTTR means less downtime, fewer missed service level agreements (SLAs), and a better experience for your users. The higher the MTTR, the longer outages, more missed service level agreements (SLAs), and a worse experience for your users. That makes MTTR a metric that engineering managers report to the business and one of the main outcomes that teams look for when they evaluate observability and AIOps platforms.


What does the R in MTTR stand for?

MTTR has four common readings: mean time to repair, mean time to recovery, mean time to resolution, and mean time to respond. Each one starts or stops the clock at a different point, so when you compare MTTR figures from other teams or vendors, check which definition they use. Mean time to resolution stops the clock only when the incident is fully resolved and the system is back to normal operation.

The acronym also appears outside incident management, with a meaning of its own in each field:

  • Security and vulnerability remediation: MTTR usually means mean time to respond or mean time to remediate: how long it takes to contain a threat or to patch a vulnerability once it has been found.
  • Equipment maintenance: MTTR means mean time to repair: the average time to fix a failed machine or component and return it to service. Maintenance teams read it next to MTBF to judge how available their equipment is.
  • Service level agreements: An MTTR target can be part of an SLA. The provider commits to restoring service within a set time for each priority level, and a missed target can lead to penalties or service credits.

How to calculate MTTR

To calculate MTTR for a given period, add up the resolution times of incidents resolved during that period, then divide by the number of those incidents.

MTTR = Total resolution time across all incidents ÷ Number of incidents resolved

Suppose your team resolved three incidents in a month. The first took 45 minutes, the second took 2 hours, and the third took 75 minutes.

MTTR = (45 + 120 + 75) / 3 = 80 minutes

The result is only comparable from month to month if every incident is timed the same way. Before you start tracking MTTR, agree on three things:

  • When the clock starts: Use the moment the incident is detected, which is usually the timestamp of the first alert. Some teams start the clock when the incident begins, which adds detection time to the total.
  • When the clock stops: Stop when the service is back to normal operation, not when a fix is first deployed.
  • Which incidents count: Track MTTR for each severity level. One long-running, low-priority ticket can hide a fast response to critical outages.

Limitations of MTTR

An average also hides outliers. In the example above, a single two-hour incident accounts for half of the total, so review the median and the longest incidents next to the mean.

MTTR also says nothing about how severe each incident was or how many incidents you prevented, so a team can lower it by closing many small incidents quickly while major outages still take hours. When a period has only a handful of incidents, one slow incident moves the average a long way, and month-to-month changes are often noise. It doesn't show whether a fix was held either: for security fixes, some teams track mean time to safe remediation (MTTSR), which counts a remediation only once it is verified and the exposure doesn't come back. Read MTTR together with reliability and maintainability measures such as MTBF, MTTF, and availability, not on its own.

MTTR benchmarks

There is no universal target for MTTR, because the result depends on the system, the severity of the incident, and how you define "resolved." A useful public reference is DORA, the research program behind the State of DevOps reports, which measures a closely related metric: the time it takes to recover from a failed deployment. In its 2024 Accelerate State of DevOps Report, elite teams recover in less than one hour, high performers recover in less than one day, and the lowest-performing teams take between a week and a month.

For a critical, customer-facing service, an MTTR of under one hour puts you alongside the top-performing teams. For everything else, the most useful benchmark is your own trend over time.


MTTR, MTTD, MTTA, MTBF, and MTTF: the five mean-time metrics

MTTR is one of five mean-time metrics that SRE teams track together. Three of them (MTTD, MTTA, and MTTR) measure stages of a single incident, and the other two (MTBF and MTTF) measure how long a system runs before it fails.

MetricFull nameWhat it measuresSRE goal
MTTDMean Time to DetectHow long an incident goes unnoticed after it beginsDetect problems before your users do.
MTTAMean Time to AcknowledgeHow long an alert waits before a responder picks it upGet the right responder on it quickly.
MTTRMean Time to ResolutionHow long it takes to fully resolve an incident once it is detectedReturn to normal operation quickly.
MTBFMean Time Between FailuresHow long the system runs between one incident and the nextFail less often.
MTTFMean Time to FailureHow long a component runs before it fails, for parts that are replaced rather than repairedReplace components before they fail.

 

MTTA is part of MTTR: the time an alert waits for a responder is counted inside the resolution time, so noisy alerts and unclear on-call routing raise both. MTTD and MTBF affect MTTR from outside it:

  • MTTD (Mean Time to Detect)—the average time between when an incident begins and when the team becomes aware of it. Shorter MTTD reduces MTTR. An incident that is caught early has affected fewer services and produced fewer alerts, so there is less to investigate and less to repair. If your MTTR clock starts when the incident begins, every minute of detection time also counts toward the total.
  • MTBF (Mean Time Between Failures)—the average time between incidents. Higher MTBF means fewer incidents, which complements lower MTTR. MTTR describes how quickly you recover, and MTBF describes how often you have to. Together they determine the availability of a service: Availability = MTBF / (MTBF + MTTR).

What causes high MTTR?

An incident moves through three stages: noticing the problem, finding its cause, and getting the right people to fix it. MTTR covers the last two, and a problem that is noticed late is harder to investigate and fix. High MTTR usually comes from one of three root causes:

  • Slow detection: Alerts fire too late, or they fire so often that responders develop alert fatigue and start to ignore them. Static thresholds miss anomalies that never cross the line, and gaps in infrastructure monitoring mean that users report some failures before an alert does.
  • Slow investigation: There is no unified view of metrics, logs, and traces. Engineers pivot between tools, re-enter the same time range in each one, and correlate the results by hand.
  • Poor coordination: The incident workflow and service ownership are unclear. Time goes to finding out who owns the failing service, waiting on handoffs, and repeating work that another responder has already done.

Each cause gets worse as systems grow. In a distributed application, one failing dependency can raise alerts in dozens of services, and the telemetry that explains it is spread across all of them.


How to reduce MTTR

Reducing MTTR means shortening each stage of an incident. To reduce MTTR, look at where your team loses time during an incident. Faster detection and investigation, supported by clear coordination, can help you resolve problems sooner.

Faster detection

A static threshold fires only after a value crosses a line that someone chose in advance, and it can't account for daily or weekly patterns. ML-powered anomaly detection reduces MTTD by flagging unusual behavior before it becomes an outage—catching the signal before a human-configured threshold would fire. A machine learning model builds a baseline of normal behavior for each metric or log stream, including its seasonality, and scores every deviation by severity. That also reduces alert noise: responders see fewer alerts, and the ones they see are worth acting on.

Faster investigation

Investigation is the root cause analysis phase of an incident: finding out why it happened. It goes faster when metrics, logs, and traces are on one platform. You can follow a problem from the metric that raised the alert to the traces of the affected requests and then to the logs of the failing service, without changing tools.

The next step is to automate the correlation. AIOps correlates anomalies across metrics, logs, and traces to surface probable root causes automatically—cutting investigation time from hours to minutes. With AI agents, the first pass of the investigation can run as soon as the alert fires. Elastic's agentic investigation uses AI to automatically analyze incident signals and suggest root causes—reducing the manual investigation burden on on-call engineers.

Proactive prevention

Prevention raises MTBF, and the incidents you prevent never enter the MTTR calculation. Two practices help:

  • Anomaly forecasting: Forecast a metric such as disk usage or queue depth from its history, and act before it reaches a limit.
  • SLO tracking: Define service level objectives (SLOs) and alert on how fast the error budget is burning so that a slow degradation gets attention before it becomes an outage.

Pair these levers with clear on-call ownership, runbooks, and a defined incident workflow, which address the coordination delays. After an incident, a blameless postmortem looks at what to change in the system instead of who is at fault, so the same failure is less likely to happen again.


How Elastic reduces MTTR

Elastic Observability applies all three levers on one platform, from the first anomaly through root cause investigation.

  • Detection: Elastic's ML anomaly detection shortens MTTD, so responders start investigating sooner, while the incident affects fewer services. Preconfigured anomaly detection jobs for logs, APM, and infrastructure work out of the box, with no data science expertise required. The same machine learning engine forecasts trends, so you can act before a limit is reached.
  • Investigation: Elastic Observability is built on Elasticsearch, which stores and queries all of its data. Logs, metrics, and traces are unified in Elasticsearch, which eliminates tool-switching during an investigation. A single ES|QL query can cross all three signals, and OpenTelemetry data is stored natively. AIOps features such as log categorization, log pattern analysis and latency and error correlation in APM point to the fields and services behind a spike.
  • Agentic investigation: When an alert fires, Elastic's AI agents lead the investigation: they correlate signals across services, surface the probable root cause, and recommend next steps, with full transparency so that SREs stay in control. Elastic Workflows can then open the ticket and notify the owning team, which shortens the coordination stage as well.

Because every signal shares one datastore, the context gathered during detection carries through to the investigation and the fix. For the concepts behind these capabilities, read what AIOps is. For a walkthrough of noise reduction, anomaly detection, and root cause analysis in Elastic Observability, see AIOps and MTTR reduction.