In short. 99.9% allows roughly 43 minutes of downtime a month, and that arithmetic is the easy part — every calculator online gives you it. What almost nobody publishes is what the percentage excludes: scheduled maintenance, third-party failures, degradation that is slow rather than down, and the fact that the number is usually measured by the provider being measured.
The arithmetic, since you came for it
Availability is the share of a period during which the service was up. The allowance follows directly from the length of that period, which is why monthly and annual figures differ so sharply.
| Availability | Per month (30 days) | Per year | Common name |
|---|---|---|---|
| 99% | 7 h 12 m | 3 d 15 h | Two nines |
| 99.5% | 3 h 36 m | 1 d 19 h | — |
| 99.9% | 43 m 12 s | 8 h 45 m | Three nines |
| 99.95% | 21 m 36 s | 4 h 22 m | — |
| 99.99% | 4 m 19 s | 52 m 34 s | Four nines |
| 99.999% | 25.9 s | 5 m 15 s | Five nines |
# The whole calculation, if you prefer to derive it
# minutes in a 30-day month
30 * 24 * 60 = 43200
# allowed downtime at 99.9%
43200 * (1 - 0.999) = 43.2 minutes
# annual
365 * 24 * 60 * (1 - 0.999) = 525.6 minutes = 8 h 45 m 36 s
Two practical notes. Each additional nine divides the allowance by ten, so the step from three nines to four is not an incremental improvement — it is a different operational model. And a monthly window is far stricter than an annual one at the same percentage, because a single bad afternoon consumes a month's budget while barely denting a year's.

Five things the percentage does not cover
This is where a published availability figure and a customer's experience separate. None of the following is unusual or dishonest — they are standard terms — but each one makes the effective guarantee smaller than the headline.
1. Scheduled maintenance
Planned windows are almost always excluded from the calculation. A provider can be unavailable for a maintenance window every month and still report 100% availability, because that time never entered the denominator.
What to look for: how much notice is required, whether there is a cap on total maintenance time, and whether the window is restricted to low-traffic hours in your timezone rather than theirs.
2. Anything upstream
Failures attributed to networks, DNS providers, certificate authorities or other third parties outside the provider's control are typically excluded. This matters more than it sounds, because a large share of real outages are exactly that — an upstream provider, a routing problem, an expired certificate at a dependency.
3. Slow is not down
Most agreements define availability as a successful response, not a timely one. A service answering every request in eleven seconds is technically up and practically unusable, and it consumes none of the downtime budget.
This is the single largest gap between the number and the experience. If response time matters to you — and for anything user-facing it does — it needs to be a separate commitment with its own threshold, because the availability figure will not capture it.
4. Partial failure counts as full availability
If checkout is broken while the marketing pages serve perfectly, an availability measure based on the homepage records no downtime. Whether a partial outage counts depends entirely on what is being probed, and that is usually not specified in the agreement at all.
5. The measurement window
Availability measured in one-minute samples and availability measured in five-minute samples produce different numbers from identical reality. A ninety-second outage may not appear at all in five-minute sampling. Neither approach is wrong; they are simply not comparable, and a figure quoted without its sampling interval is not a figure you can check.
Read the exclusions before the percentage. Two providers advertising the same number can differ by an order of magnitude in what that number actually promises, and the difference lives entirely in the paragraphs most people skip.
Who measures it decides what it says
In most agreements below enterprise scale, availability is measured by the provider's own monitoring. That is not necessarily self-serving — they have better instrumentation than you do — but it means the party with the obligation is also the party with the measurement.
Three questions are worth asking, and the answers are usually in the agreement:
| Question | Why it matters |
|---|---|
| Who measures? | Provider, customer, or an agreed third party |
| From where? | Inside their network measures their servers; outside measures what users experience |
| How often? | The sampling interval sets the smallest outage that can be detected |
| What is probed? | A health endpoint can pass while the application fails |
| What counts as failure? | One failed sample, or several consecutive ones |
A health endpoint that returns "ok" whenever the web server is running is a common and important gap. It passes while the database is unreachable, so the availability figure stays perfect through an outage that stopped every real request. Probing something that exercises the actual dependencies is what makes the number mean anything.

What a service credit is actually worth
The remedy for a breached availability commitment is normally a service credit — a percentage of the fee for the affected period, applied to a future invoice. Four characteristics are near-universal and worth understanding before you rely on one:
- It is a credit, not a refund. It reduces a future bill and has no cash value if you leave.
- It is capped, commonly at a fraction of the monthly fee for that service.
- You usually have to claim it, within a limited window, with evidence. Credits are rarely applied automatically.
- It is proportional to what you pay, not to what the outage cost you. An hour of downtime during a launch is worth far more to you than a share of one month's hosting fee.
A service credit is not compensation for your losses and is not designed to be. It is a pricing adjustment that signals the provider takes the commitment seriously. Treating it as insurance is the mistake — if downtime would genuinely hurt, the answer is architecture and a fallback, not a stronger clause.
Verifying a provider's claim yourself
You do not have to accept a published figure on trust, and measuring it independently is straightforward. The only requirement is that your measurement is honest about its own limits.
- Probe from outside the provider's network, from more than one region. That is the path your users take.
- Probe something meaningful — a page that touches the database, not a static health file.
- Record every check, not just the failures, so you can compute availability rather than count incidents.
- Use a sampling interval at least as fine as the smallest outage you care about, and state it whenever you quote a number.
- Distinguish slow from down in your own records, since the agreement will not.
# A single check tells you the state now; availability needs the series.
curl -o /dev/null -s -w 'code=%{http_code} ttfb=%{time_starttransfer}s\n' https://example.com/
# Availability from recorded checks is just arithmetic:
# successful_checks / total_checks * 100
# which is why the interval and the probe target matter more than the formula.
Scheduled monitoring keeps that series from several regions and records response time alongside status, which is what lets you separate "slow" from "down" — the distinction the agreement leaves out. Our own plans differ mainly in check frequency and history depth, and frequency is the parameter that decides how small an outage you can detect at all.
Reading an SLA before you sign
Work through the document in this order. The percentage is the last thing to look at, not the first.
| Clause | What to check | Warning sign |
|---|---|---|
| Definition of downtime | Does slow count? Does partial? | Only "complete unavailability" |
| Exclusions | Maintenance, third parties, force majeure | Uncapped maintenance windows |
| Measurement | Who, from where, how often | Unspecified sampling interval |
| Window | Monthly or annual | Annual — hides a very bad month |
| Remedy | Credit size, cap, claim process | Automatic-sounding but claim-required |
| Claim deadline | How long you have | A window shorter than your billing cycle |
| Termination right | Repeated breach lets you leave | No exit for chronic failure |
The last row is often the most valuable clause in the document and the least discussed. A credit compensates you for one bad month; the right to leave without penalty after repeated breaches is what actually changes a provider's incentives.

If you are the one publishing an SLA
Committing to a number changes what you have to build, and the honest sequence is to measure first and promise second.
- Measure your current availability for at least a full quarter before committing to anything. Promising a figure you have not achieved is a decision to breach.
- Set the target below your measured performance, not at it. The gap is your margin for a bad month.
- Define downtime precisely, including whether slow counts. Ambiguity favours nobody at the point of a dispute.
- State the measurement method — interval, probe target, vantage point. A number without a method is not verifiable and invites argument.
- Publish a status page and update it during incidents. Availability commitments are believed in proportion to how transparently failures are reported.
- Track your error budget, which is simply the allowance the percentage grants you. Spending it deliberately on releases is what the number is for.
The relationship between the target, the budget and release velocity is the substance of error budgets, and the vocabulary that separates an internal objective from a customer commitment is covered in SLI, SLO and SLA explained.
What affects the number in practice
Availability is lost in a small number of recurring ways, and they are worth knowing because most are cheaper to prevent than the nines suggest:
- Certificate expiry. A fully predictable, fully preventable total outage. It remains one of the most common causes.
- Domain and DNS lapses. Same category: a calendar event that becomes an outage.
- Deploys. Usually the largest single contributor, and the one an error budget is designed to price.
- Dependency failures. A payment provider, an identity provider, a third-party API in the response path.
- Capacity under peak. Not a failure until traffic arrives, which is why it is discovered at the worst moment.
Two of those five are date-driven and can be removed entirely by watching the dates. SSL checks and domain expiry lookups take seconds; monitoring them continuously removes an entire class of self-inflicted downtime.
How to check your own availability
Start with the current state — uptime monitoring from several regions gives you the series that availability is computed from, and records response time alongside status so slowness does not hide inside a passing check. Verify the two date-driven risks directly with the SSL checker and WHOIS lookup, and publish what you find on a status page so customers see incidents from you rather than from their own users.
One caution about your own numbers: availability computed from a single vantage point measures the path between that point and your server, not availability in general. Checking from several regions is what separates "our service was down" from "one network could not reach us" — a distinction covered in the layer isolation guide.

Frequently asked questions
How much downtime does 99.9% allow?
About 43 minutes in a 30-day month, or 8 hours 45 minutes over a year. The monthly figure is the demanding one: a single 45-minute incident breaches a monthly commitment while barely registering against an annual one.
Is 99.9% good?
It depends entirely on the exclusions. A 99.9% commitment that counts slow responses as downtime, measures from outside the provider's network at one-minute intervals and caps maintenance windows is a stronger promise than a 99.99% commitment with none of those.
Does scheduled maintenance count as downtime?
Usually not. Planned windows are typically excluded from the calculation, which means a service can be unavailable regularly and still report a perfect figure. Check whether total maintenance time is capped and how much notice you get.
What do I actually get if the SLA is breached?
Normally a service credit against a future invoice, capped at a fraction of that period's fee, and usually only if you claim it within a stated window. It is not compensation for your losses and is not intended to be.
Can I verify a provider's uptime claim?
Yes, by monitoring from outside their network against a page that touches real dependencies, and recording every check. State your sampling interval whenever you quote the result — a figure without one cannot be compared to anything.
Should my status page show the percentage?
Only alongside its method: the period, the sampling interval and what was probed. A bare percentage invites the same confusion this article is about, and it is the confusion that damages trust when customers experience something different.
Checklist
- Read the exclusions before the percentage.
- Check whether the window is monthly or annual — the same number means very different things.
- Establish whether slow responses count as downtime; usually they do not.
- Find out who measures, from where, and how often.
- Probe something that touches real dependencies, not a static health file.
- Treat a service credit as a pricing adjustment, not as insurance.
- Note the claim deadline — credits are rarely automatic.
- Look for a termination right on repeated breach; it is the clause with teeth.
- Measure your own availability for a quarter before publishing a target.
- Quote every availability figure with its sampling interval and vantage point.