Email Outage Statistics 2027: SLAs, Downtime and Credits
What a 99.9% email SLA actually permits, how long the notable provider outages really ran, and the arithmetic on what an SLA credit pays out. Spoiler: not much.
What a 99.9% email SLA actually permits, how long the notable provider outages really ran, and the arithmetic on what an SLA credit pays out. Spoiler: not much.
The email outage statistics all start in the same place: a 99.9% uptime SLA gives your email provider legal permission to be unavailable for 8 hours and 45 minutes a year. If they blow through it, the compensation on a 20-seat Google Workspace account comes to about $14.
Fourteen dollars. For a morning where nobody could send a quote.
I went looking for email outage statistics because a founder asked a simple question — "which provider has the best uptime guarantee?" — and the honest answer isn't a percentage at all. Everybody publishes 99.9%. The interesting number is what happens when they miss it, and almost nobody reads that part of the contract. So I read it. Several of them. On a Sunday, which tells you something about how I spend weekends.
Four kinds of source behind the email outage statistics below, kept separate on purpose.
Published SLA terms — the Google Workspace SLA, the Microsoft Online Services SLA covering Exchange Online, the AWS SLAs for Amazon SES and Amazon WorkMail, and Zoho's published uptime commitment. Contract documents, not marketing pages.
Post-incident reports and status-page histories from the providers, cross-checked against Downdetector's aggregates for start and recovery timing.
Industry cost surveys — the Uptime Institute's Annual Outage Analysis and ITIC's Hourly Cost of Downtime survey, most recent published waves. Both are self-reported and skew toward organisations big enough to have someone whose job is filling in surveys. I quoted one of these in a deck for two years before I read its methodology page. Don't be me.
Arithmetic. The downtime tables below are percentages applied to a 30-day month (43,200 minutes) and a 365-day year (525,600 minutes). Stating the assumption because a 31-day month shifts the figures slightly — I built the first version of this table on 31 days and had to redo it.
One caveat: outage durations are hard to pin down. A single incident can run 90 minutes for one tenant and most of a working day for another in a different region. Every duration below is approximate.
Start here, because the marketing number and the lived experience are further apart than people expect.
| Uptime | Downtime per 30-day month | Downtime per year |
|---|---|---|
| 99% | 7h 12m | 3d 15h 36m |
| 99.5% | 3h 36m | 1d 19h 48m |
| 99.9% | 43m 12s | 8h 45m 36s |
| 99.95% | 21m 36s | 4h 22m 48s |
| 99.99% | 4m 19s | 52m 34s |
| 99.999% | 25.9s | 5m 15s |
The gap between 99.9% and 99.99% is the whole ballgame. One is nine hours a year. The other is under an hour. And every major email provider commits to the first number.
So a five-hour incident on a status page isn't a fluke. It's most of the annual budget, spent in one afternoon — and they're still technically compliant for the month if nothing else breaks. Contractually fine. Operationally a bad day.
| Provider | Published monthly uptime commitment | Credit structure |
|---|---|---|
| Google Workspace | 99.9% | 3 days of service added (99.0–99.9%), 7 days (95.0–99.0%), 15 days (below 95%) |
| Microsoft 365 / Exchange Online | 99.9% | 25% of monthly fee (below 99.9%), 50% (below 99%), 100% (below 95%) |
| Amazon SES | 99.9% | 10% of monthly charge (99.0–99.9%), 25% (95.0–99.0%), 100% (below 95%) |
| Amazon WorkMail | 99.9% | Service credit against the monthly charge, banded |
| Zoho Mail | 99.9% (published commitment — confirm against your own agreement) | No public banded credit table; check what you signed |
Notice what's identical. Everyone lands on 99.9%. The difference isn't the promise, it's the payout — and even there the spread runs from "almost nothing" to "slightly more than almost nothing."
Here's the take I'll defend: an uptime number on a pricing page is a marketing artefact, not an engineering one. It tells you what legal was comfortable signing. Not how the thing behaves at 3am.
Verify these against your own agreement. Enterprise contracts override public pages, and SLA documents get revised without a press release. (One provider's credit table had already changed since the version cached in my notes from earlier this year.)
What genuinely annoys me is that none of these documents defines "downtime" the way you would. Partial degradation, one region, attachments broken but send working — whether that counts is the provider's call, measured with the provider's telemetry.
| Date | Provider | Approx. duration | What broke |
|---|---|---|---|
| 20 Aug 2020 | Google Workspace | ~6 hours | Gmail sending and attachment handling failed; storage subsystem fault |
| 14 Dec 2020 | ~47 minutes | Central authentication quota exhausted — Gmail, Drive and YouTube down together | |
| 25 Jan 2023 | Microsoft 365 | ~5 hours | WAN router configuration change; Exchange Online, Outlook and Teams affected globally |
| 1 Mar 2024 | Microsoft 365 | ~2–3 hours | Outlook and Teams access failures across multiple regions |
| 19 Jul 2024 | Microsoft / Azure | 10+ hours (Central US) | Azure Central US storage incident, coinciding with the CrowdStrike update the same day |
| 25 Nov 2024 | Microsoft 365 | Most of a working day for some tenants | Exchange Online and Teams degraded following a component change |
The December 2020 Google incident is the most instructive on the list, and it's the shortest. Forty-seven minutes. But because it hit the identity layer rather than the mail layer, it took down everything that authenticates against a Google account at once. Duration is a bad proxy for damage.
Microsoft dominates this table, which I want to be careful about. It isn't evidence that Exchange Online is less reliable than Gmail — Microsoft publishes far more detailed incident write-ups, runs a larger enterprise tenant base, and gets covered harder when something wobbles. Selection bias is doing real work here.
Nobody does the arithmetic on this, so let's do it.
Take a five-hour outage — the January 2023 shape. Five hours is 300 minutes out of a 30-day month's 43,200. That's 0.694% unavailability, so the month closes at 99.31% uptime: below the 99.9% commitment, comfortably above 99%. First credit band for both major providers.
Google Workspace, at the entry list price of $84/user/year, is $7.00 per user per month. The 99.0–99.9% band awards 3 days of added service. Three days out of thirty is 10% of a monthly fee, so $0.70 per user. Twenty seats: $14.00.
Microsoft 365, at the entry list price of $72/user/year, is $6.00 per user per month. Below 99.9% triggers a 25% credit, so $1.50 per user. Twenty seats: $30.00.
Both are the cheapest published tier. Pay more per seat and the credit scales with it. Moves the decimal, nothing else.
Now the part that turns a small number into a smaller one: you have to claim it. Neither provider refunds anyone proactively. Google requires the claim within 30 days of the incident. Microsoft requires submission by the end of the month following the month it occurred. Both want you to notice the outage, keep evidence, find the right form, and open a ticket — for fourteen dollars.
I did this once, out of stubbornness. Spent longer finding the right form than the credit was worth, which — yes — was the thing I set out to prove.
Real talk: the service credit isn't compensation. It's a rounding error dressed up as accountability. Once you see it that way, "which provider has the better SLA?" stops being a useful question.
Same lens on us, because it'd be cheap not to. JustEmails is $49/year flat and we don't publish a financially backed SLA. That's a real gap and I'd rather say it plainly than bury it — but note what a credit would be worth if we did. $49/year is about 13 cents a day. Three days of credited service: forty cents. The credit was never the product, at any price point.
Here's my genuinely contrarian take, and the one piece of this that changed how I think about email reliability.
Most email outages don't lose your mail. They delay it.
SMTP is store-and-forward. When a receiving server is unreachable, or returns a 4xx temporary failure, the sender doesn't give up — it queues the message and retries on a backoff schedule. Postfix's default maximal_queue_lifetime is 5 days. Exim's default retry configuration runs about 4 days. Gmail defers and retries for a couple of days before it hard-bounces anything.
So a five-hour inbound outage means five hours of mail arriving late in a burst. Annoying, occasionally expensive, rarely catastrophic. A different thing from what "your email was down" implies — and the gap between those two sentences is where most vendor fear-marketing lives.
Three exceptions, and they're where the real money burns:
DNS is the dangerous one. If your MX records vanish or the domain returns NXDOMAIN, some sending servers treat that as permanent and hard-bounce immediately. No retry, no queue, message gone. A DNS misconfiguration during a migration can do more damage in ten minutes than a mail server offline for six hours — which is why auto-configured DNS matters more than it sounds. What actually breaks during a provider switch.
Outbound transactional email has no retry buffer on the human side. A password reset forty minutes late is a support ticket. An OTP that arrives after the code expires is a failed login and possibly a lost signup.
Authentication outages cascade. See December 2020. If the identity layer goes, mail is one of a dozen things down at once, and your team can't even open the status page in the same browser session.
ITIC's Hourly Cost of Downtime survey is the figure quoted everywhere. In its most recent published wave, roughly 90% of mid-size and large enterprises put a single hour of downtime above $300,000, with about 41% placing it between $1 million and $5 million per hour. The Uptime Institute's Annual Outage Analysis lands in the same territory — a majority of surveyed operators said their most recent serious outage cost over $100,000.
Those numbers are real. They're also useless to you unless you run a bank.
Both surveys poll data-centre and enterprise IT operators. The unit being measured is a full infrastructure outage at an organisation with thousands of employees — not "the marketing team's Outlook was down until lunch." Different animal. Applying $300,000-per-hour to a 15-person agency isn't conservative estimating; it's nonsense, and it's the most common abuse of this data in vendor content. If someone quotes it without naming the survey population, you've learned something about them.
The honest SMB version is smaller and more useful: salary hours lost, deals delayed a day, support responses that missed their window, the credibility hit when a client emails twice and hears nothing. We ran that arithmetic in what an email outage costs a 15-person company — four figures, not seven. For the general (non-email) version, JustAnalytics has a fuller breakdown, and their uptime monitoring statistics for 2027 cover how often the average service checks itself.
Stop shopping on SLA percentage. Everyone says 99.9%. The number carries no information.
Monitor your own MX, independently. A provider's status page is a lagging indicator — it updates after their engineers confirm, often 20-40 minutes after you noticed. An external check against port 25 and your DNS records tells you first. Building a status page from that data takes an afternoon.
Treat DNS as the critical path, not the mail server. Low TTLs during migrations, a secondary MX, and never — genuinely never — let a domain lapse into NXDOMAIN.
Separate the urgency tiers. Inbound mailbox delivery can absorb a few hours. Password resets and OTPs cannot.
Do the self-hosting maths honestly. Run your own server and your uptime is your uptime — no credit, no status page, nobody to escalate to at 2am. Sometimes that's still the right call: both sides here.
File the claim anyway. Fourteen dollars and a form. But providers track claim volume, and a band nobody claims against is a band nobody has to improve.
On a 30-day month, 99.9% permits 43 minutes and 12 seconds of downtime. Over a full year that's 8 hours 45 minutes. A 99.99% SLA allows 4 minutes 19 seconds monthly, or about 52 minutes a year. Google Workspace, Microsoft 365, Amazon SES and Amazon WorkMail all publish 99.9% as their standard commitment, so roughly nine hours a year of unavailability is contractually normal, not a breach.
Very little. Using the entry list prices of $84/user/year for Google Workspace and $72/user/year for Microsoft 365, a five-hour outage lands both providers in their first credit band. Google awards 3 days of added service — about $0.70 per user. Microsoft credits 25% of the monthly fee — about $1.50 per user. On a 20-seat account that's $14 and $30 respectively, and you only get it if you file a claim yourself.
Usually not. SMTP is store-and-forward: when a receiving server is unreachable or returns a 4xx temporary error, the sending server queues the message and retries. Postfix defaults to a 5-day queue lifetime, Exim to roughly 4 days, and Gmail retries deferred mail for a couple of days before bouncing. Inbound mail during a multi-hour outage is typically delayed, not destroyed. DNS failures are the dangerous exception.
The most-cited incidents are Google's 20 August 2020 Workspace outage (Gmail sending and attachments broken for roughly six hours), Google's 14 December 2020 authentication failure (about 47 minutes, but it took Gmail, Drive and YouTube down together), and Microsoft's 25 January 2023 global Microsoft 365 outage traced to a WAN router configuration change, which ran roughly five hours for many tenants. Durations vary by region and tenant.
Unlimited custom domain email hosting for $49/year flat — unlimited domains, unlimited mailboxes, 10 GB storage, full IMAP/SMTP. Built for agencies, freelancers, and anyone managing email across more than one domain.