Every IT provider advertises fast response, and the numbers quoted are rarely comparable. Some measure from ticket creation to any automated acknowledgement. Some measure to a human. Some measure only during business hours. Some measure an average across all tickets, which is dominated by the trivial ones.
Here is what reasonable actually looks like, and which figures are worth asking for.
Benchmarks by severity
Response times only make sense per severity tier. Reasonable targets for a well-run managed provider:
- Critical, business down — 15 to 30 minutes, with continuous effort until restored
- High, department or function blocked — 1 hour
- Medium, individual blocked with a workaround — 4 hours
- Low, requests and questions — 1 business day
These are response targets, not resolution. A provider committing to 15 minutes on critical incidents is committing to have someone actively working, which is what matters when the business is stopped.
Anything substantially slower than this on critical incidents is worth questioning. Anything dramatically faster across every tier is worth questioning differently — it may indicate an automated acknowledgement being counted as a response.
Why average response time misleads
A provider quoting a single average figure is quoting a number dominated by the easy tickets.
If 80% of tickets are password resets and software questions answered in minutes, the average will look excellent even if every genuine outage takes four hours. The average tells you about ticket mix, not about performance when it counts.
Ask for the median response on critical and high severity tickets specifically, and ask for the 90th percentile alongside it. The 90th percentile is where the bad experiences live, and it is the number that predicts how it will feel to be a client.
What determines the number
- Staffing model. A pooled queue with several engineers responds faster than a named individual who may be on another call, though the named individual usually resolves faster once engaged.
- Coverage hours, and whether the clock runs outside them.
- Triage quality. Providers with good intake processes route correctly the first time; poor triage produces fast acknowledgement followed by long silence.
- Client load per engineer. This is the number nobody advertises and it drives everything. Ask how many clients and endpoints each engineer supports.
- Monitoring maturity. A provider whose tooling detects a failing server before anyone reports it has effectively negative response time on the incidents that matter most.
That last point reframes the metric. The best response time on an outage is the one where the problem was resolved before staff noticed, and no ticket exists to measure.
Questions that produce honest answers
- What is your median response on critical tickets over the last quarter, and what is your 90th percentile?
- Does an automated acknowledgement count as a response in your reporting?
- Does the clock run outside business hours, and how is that defined?
- Who assigns severity, and what happens if we disagree?
- How many endpoints does each engineer support?
- What percentage of incidents did your monitoring detect before a user reported them?
That final question is the most revealing in the list, and the one least often asked. A provider who can answer it has genuine monitoring maturity. A provider who cannot has agents installed and a dashboard nobody opens.
What to expect realistically
For a well-run managed provider, most tickets should get a human response within an hour during business hours, critical incidents within half an hour at any time covered by the agreement, and routine requests within a business day.
More useful than any single figure: does response time hold up during a bad week? Every provider performs well on a quiet Tuesday. The question is what happens when three clients have problems simultaneously, and that is what capacity per engineer actually determines.