99.9% uptime: a commitment, not a guarantee
On this page
The word “guarantee” does a lot of work in web hosting.
You can find 99.9%, 99.99%, and occasionally 99.999% uptime guarantees attached to hosting plans at almost every price point.
Taken literally, the difference is substantial.
99.9% uptime allows roughly 43 minutes of downtime in an average month.
99.99% allows a little over four minutes.
99.999% leaves only about 26 seconds.
Four minutes is not much time to detect a problem, diagnose it, replace failed hardware, reroute traffic, recover a service, or occasionally even finish rebooting a server.
So what exactly is being guaranteed?
An uptime number does not prevent an outage
The first thing worth separating is the engineering objective from the contractual language around it.
A hosting provider can design infrastructure for high availability. It can use redundant power, multiple network connections, RAID or replicated storage, spare hardware, monitoring, automated recovery and carefully planned maintenance.
All of those things reduce the probability and duration of outages.
None of them makes an outage impossible.
This remains true even at a scale far beyond ordinary web hosting.
| Company / platform | Window | Outage duration | Primary root cause & context |
|---|---|---|---|
| AWS (US-EAST-1) | Oct 2025 | ~15 hours | A severe cascading failure triggered by a core DNS resolution fault dropped the monthly uptime of regional Lambda, EC2, and NLB services to ~97.9%. |
| CrowdStrike | Jul 2024 | 24+ hours (recovery took days) | A faulty Falcon sensor content update crashed roughly 8.5 million Microsoft Windows machines globally, reducing monthly service availability drastically. |
| Microsoft Azure | Oct 2025 | 8+ hours | A global cloud disruption originating from an Azure Front Door configuration change knocked out global portals, Teams, and Outlook, violating monthly SLAs. |
| Atlassian | Apr 2022 | ~14 days | A faulty internal script deleted site data across Jira, Confluence, and Opsgenie, dropping monthly platform availability to roughly 53.3%. |
| Meta (Facebook, WhatsApp, Instagram) | Oct 2021 | 6 hours | A global backbone routing error (Border Gateway Protocol / BGP) isolated all Meta data centers from the internet, dropping monthly availability to ~99.1%. |
| Google Cloud (GCP) | Jul 2022 | ~24 hours | An unprecedented UK heatwave caused cooling hardware failures at a London data center, knocking out regional Storage and GKE compute services. |
| PlayStation Network | 2011 | 23 days | An external security breach forced Sony to pull the entire network offline to rebuild infrastructure, reducing its full annual uptime to roughly 93.7%. |
| Fastly | Jun 2021 | ~1 hour | A customer changing a valid configuration triggered a hidden software bug, taking down the entire edge cloud platform and dropping monthly uptime to ~99.8%. |
| Cloudflare | Nov 2025 | 6+ hours | A major software bug combined with internal user permissions errors broke global CDN routing, dropping monthly performance baseline. |
| OVHcloud | 2021 | Multiple weeks (destruction) | A massive, catastrophic fire physically destroyed the SBG2 data center in Strasbourg, knocking affected customers completely offline and wiping out annual SLA metrics. |
| BlackBerry | Oct 2011 | 3–4 days | A core switch failure inside BlackBerry’s private network backed up global infrastructure, cutting messaging and email and dropping monthly availability to ~87%. |
| Dyn (Oracle) | Oct 2016 | ~11 hours | A massive Mirai botnet DDoS attack overwhelmed Managed DNS servers, taking down primary internet pathways and crashing monthly availability. |
These are organizations with engineering resources and redundancy that a normal hosting company could not realistically reproduce.
The point is not that redundancy doesn’t work.
It is that redundancy protects against particular failures. It does not protect automatically against every possible failure.
More redundancy solves some problems and creates others
A second server can protect against the failure of the first server.
A second datacenter can protect against some failures affecting the first datacenter.
A second region can protect against some failures affecting the first region.
But two regions do not help much if the same broken configuration is deployed to both.
Neither does geographic separation automatically protect you from a shared DNS provider, a common authentication system, a software defect present everywhere, compromised administrative credentials, or a control system on which all locations depend.
There is an important distinction here that is similar to the distinction between redundancy and backups.
Redundancy reduces particular failure modes. It does not guarantee availability.
As infrastructure becomes more elaborate, you can remove more single points of failure. But you also introduce coordination, replication, routing and control systems that need to work correctly themselves.
Highly available architecture is therefore less about reaching a point where failure becomes impossible and more about continually reducing the number of ways in which a failure can become an outage.
One bad incident can consume a lot of nines
Uptime percentages can also be slightly deceptive because they compress very different operating histories into one neat number.
A service that is unavailable for 40 minutes once and otherwise works perfectly can have almost the same monthly uptime as a service suffering dozens of short interruptions throughout the month.
The percentage does not tell you which happened.
And the more nines you add, the less room there is for even a single unusual incident.
With a 99.99% monthly target, one ten-minute outage has already exceeded the month’s entire downtime allowance by more than twice.
That does not necessarily mean the provider has poor infrastructure. It may mean something genuinely unusual happened.
It does mean that the 99.99% target was missed.
Months of uneventful operation do not make a serious outage disappear.
You cannot average away Tuesday’s outage by pointing out that Monday worked beautifully.
So what is an “uptime guarantee”?
This is where the terminology becomes important.
No hosting provider can literally guarantee that its infrastructure will never experience more than a certain number of minutes of downtime.
What providers can do is define a service level and explain what happens if they fail to meet it.
A properly written Service Level Agreement can be quite honest about this. It can define what is measured, over what period, what qualifies as downtime, what does not, and what contractual remedy applies when the service level is missed.
In other words, the SLA does not make the outage impossible.
It defines accountability after the outage happens.
That is rather different from the everyday meaning of the word “guarantee.”
And the number alone tells you surprisingly little without the rest of that information.
A nominal 99.99% commitment with extensive exclusions and unclear measurement can tell you less about the service than a clearly explained 99.9% standard.
Before comparing uptime percentages, it is worth asking:
- What exactly is being measured?
- Over what period?
- From where?
- What counts as downtime?
- What is excluded?
- And what does the provider actually commit to doing when performance falls short?
The answers are at least as important as the number of nines.
What uptime are we actually measuring?
Even the apparently simple question “Was the website up?” can have several answers.
- A WordPress plugin can crash and produce an error page while the server itself is operating normally.
- A payment gateway can stop responding while both the website and hosting platform remain available.
- A registrar-level DNS failure can make a domain unreachable even though the hosting server is working perfectly.
- An external CDN can fail between the visitor and the origin server.
- Or the hosting infrastructure itself can become unavailable.
These are all real availability problems for somebody visiting a website, but they are not necessarily failures of the same service.
That is why an uptime commitment needs a defined boundary.
How we approach uptime at ModHost
At ModHost, our published Service Commitments set a target of 99.9% uptime per calendar month, measured at the server level for the hosting service itself.
That corresponds to roughly 43 minutes of downtime per month.
Our actual historical uptime, averaged across our servers over the past ten years, has been higher than that target.
We could therefore put a larger number on the page.
We don’t.
99.9% is the number we are prepared to commit to through normal operations, hardware failures, maintenance, unexpected infrastructure problems and the occasional genuinely bad day. Our Service Commitments deliberately describe it as a target rather than a contractual guarantee.
That does not mean we regard 43 minutes of downtime every month as acceptable.
We do not operate servers with the objective of carefully using up the monthly downtime allowance.
The operational objective is continuous availability.
99.9% is the threshold at which we consider our service to have fallen short of the published commitment.
Being clear about what counts matters too
Our uptime measurement covers periods when the hosting service for a customer’s website is unavailable because of ModHost-managed infrastructure.
It does not count scheduled maintenance announced in advance, problems caused by customer code or applications, account-level resource overages, attacks specifically targeting an individual site, failures of external services such as third-party CDNs or APIs, registrar-level DNS problems, or other circumstances outside the hosting service itself.
Those exclusions are not there to make the uptime statistic look better.
They define what service the statistic is actually measuring.
We also publish how scheduled maintenance is handled, how incidents should be reported, and what happens when a customer experiences sustained or recurring problems. We do not provide automatic compensation every time a target is missed, but we take sustained shortfalls seriously and may address them directly with affected customers through measures such as service credits or plan adjustments.
That is why we call the page Service Commitments.
It says what we aim for, how we operate and what customers can reasonably expect from us.
It does not pretend that infrastructure can be made incapable of failing.
Reliability is what happens before the percentage
Uptime numbers still matter.
We monitor ours. We publish a target. We maintain infrastructure designed to minimize avoidable downtime, and when something does fail, getting the service back to normal becomes the priority.
But the number at the bottom of an uptime calculator is the result of those practices, not the cause of them.
A fourth nine does not add a second power supply.
It does not configure redundant storage.
It does not monitor a server.
It does not make somebody respond to an alert at three in the morning.
And it does not prevent a perfectly reasonable configuration change from occasionally producing a result nobody expected.
A meaningful uptime commitment should therefore tell you something about the standard a provider is prepared to stand behind — not merely how many nines fit comfortably on a pricing page.
Uptime is an engineering outcome, not a marketing number.
Reliability is built by everything you do before you ever need to explain why that number was missed.