Back

Your Cloud Hosting SLA Refunds Less Than One Per Cent of the Year

TLDR: A cloud hosting SLA is a billing instrument. A six-hour outage earns a 10 per cent credit on one month from AWS, Google Cloud and Azure, worth 0.83 per cent of the annual bill, so the architecture built above the provider decides survival.

99.99 per cent uptime allows four minutes and nineteen seconds a month

At 11.48 PM Pacific on 19 October 2025, the Domain Name System (DNS) records for Amazon DynamoDB’s regional endpoint in Northern Virginia went empty. A cloud hosting SLA, the service level agreement that fixes what a provider owes when it misses, priced the whole day at a rounding error. Amazon Web Services (AWS) later attributed the failure to a latent race condition in DNS management, where two automated components raced to write a plan and one deleted the other’s work. DynamoDB stayed unreachable for roughly two hours and fifty-two minutes. New Elastic Compute Cloud (EC2) instance launches stayed broken until 1.50 PM the next afternoon, about fourteen hours. Network Load Balancer errors ran from 5.30 AM to 2.09 PM.

Nobody signs a hosting contract expecting that. Buyers read the number on the front as a promise about survival, so work out what it permits. In a thirty-day month, 99.99 per cent uptime allows four minutes and nineteen seconds of downtime, and across a year, fifty-two minutes and thirty-four seconds. Drop one nine and the allowance jumps to forty-three minutes a month and eight hours forty-six minutes a year. The gap between those two figures is a bad coffee break against a full working day offline.

Read what a cloud hosting SLA calls Unavailable, then read the number

The Amazon Compute Service Level Agreement, last updated 25 May 2022, defines Region-Level Unavailability as the state where all of a customer’s running instances across two or more availability zones “concurrently have no external connectivity”. That definition does real work. During the October event, running instances kept their connectivity. What broke was the ability to launch new ones for fourteen hours. A company whose autoscaling group could not add capacity all day has a serious incident and, under that clause, no qualifying unavailability.

Google draws a similar boundary and adds a floor. Google’s Compute Engine agreement, last modified 4 March 2025, defines a Downtime Period as one or more consecutive minutes, and states plainly that “partial minutes or intermittent Downtime for a period of less than one minute will not count towards any Downtime Periods”. Error rates that climb without fully severing a connection stay invisible to the meter. The same goes for anything Google attributes to customer software, third-party hardware or applied quotas.

Then there is the claims process, which almost nobody runs. AWS requires the credit request “by the end of the second billing cycle after which the incident occurred”. Google gives sixty days and demands log files for each downtime period. Cloudflare’s Business plan agreement wants notification within five business days and a formal claim by the end of the following month. A credit nobody requests inside the window is forfeited, and each document states that the credit is the sole and exclusive remedy.

Cloudflare promises 100 per cent uptime and pays 0.8 per cent

The headline commitments differ across vendors, and the payouts converge anyway.

Take a six-hour outage in a thirty-day month. That is 99.17 per cent uptime. Under AWS Compute, that falls in the band “less than 99.99% but equal to or greater than 99.0%”, which pays a 10 per cent service credit. Google Compute Engine multi-zone instances pay 10 per cent across the same band. Microsoft Azure virtual machines deployed across two or more availability zones carry a 99.99 per cent commitment and pay 10 per cent below that threshold. That schedule is read from the 21Vianet edition, the only Microsoft-published copy that renders as text. Amazon S3 standard storage classes pay 10 per cent between 99.0 and 99.9 per cent.

Ten per cent of one month is 0.83 per cent of an annual bill.

Cloudflare is the outlier, promising more and paying less. Its Business plan commits to “100% Uptime. The Service will serve Customer Content 100% of the time without qualification.” The credit is pro rata rather than banded, calculated as outage minutes multiplied by the affected customer ratio, divided by scheduled availability minutes. Six hours out of a thirty-day month works out at 0.83 per cent of that month’s fee, or 0.069 per cent of the year, assuming the outage touched every customer, which is the ceiling the formula allows. Cloudflare’s own post-mortem of 18 November 2025 records core traffic failing from 11.20 Coordinated Universal Time (UTC) until 14.30 and all systems normal at 17.06, five hours forty-six minutes, which the schedule values at 0.80 per cent of one month’s fee. Total credits in twelve months are capped at one month of fees.

Exhibit 1
A six-hour outage refunds under one per cent of the year, whichever provider you signed with
Service credit earned by a six-hour outage in a thirty-day month, expressed as a share of the annual bill for the affected service.
AWS EC2, region-level, 99.99% commitment
Google Compute Engine, multi-zone, 99.99% commitment
Azure VMs across availability zones, 99.99% commitment
Amazon S3 standard classes, 99.9% commitment
Cloudflare Business, 100% commitment, pro rata
Bars are scaled against a 1.00 per cent ceiling. Four providers pay 0.83 per cent of the annual bill. Cloudflare, which promises the most, pays 0.069 per cent.
Source. Amazon Compute SLA (25 May 2022), Amazon S3 SLA (28 Nov 2023), Google Compute Engine SLA (4 Mar 2025), Microsoft Azure Virtual Machines SLA (21Vianet edition), Cloudflare Business SLA. Credit bands read from each published schedule, annualised by Wirecode. Wirecode exhibit.

Four dependencies at 99.99 per cent compose to 99.96 per cent

A single figure on a single service says almost nothing about the application a customer loads. Microsoft’s own architecture guidance is more candid than its sales collateral. The Azure Well-Architected reliability material states that SLAs “don’t guarantee an offering as a whole” and shows composite availability as a product distribution, working an example at 99.95 per cent multiplied by 99.99999 per cent. Its worked scenario lands on 99.45 per cent for a workload built from services with higher published numbers.

Run that yourself. Four independent dependencies, each at a genuine 99.99 per cent, compose to 99.96 per cent. That is seventeen minutes and seventeen seconds a month, four times the allowance the brochure implied. Ten dependencies at 99.99 per cent compose to 99.90 per cent, which is the eight-hours-and-forty-six-minutes-a-year figure that everyone treats as the cheap tier. No component in that chain missed its published target, and the monthly allowance still widened tenfold.

AWS is equally direct. The AWS shared responsibility model puts the provider in charge of the infrastructure and leaves the customer holding the guest operating system, patches, application software and firewall configuration. Every reliability decision above the hypervisor belongs to the buyer. The SLA covers the half of the stack that rarely takes a business down.

Delta lost thirteen times what CrowdStrike spent on the whole incident

The CrowdStrike Falcon sensor update of 19 July 2024 measures the gap between a vendor’s liability and a customer’s loss, because both sides filed numbers with the Securities and Exchange Commission.

Delta Air Lines told investors in its September-quarter 2024 release that the incident caused approximately 380 million dollars of direct revenue impact, driven by refunds and customer compensation, plus 170 million dollars of non-fuel expense, offset by 50 million dollars of fuel the airline did not burn across 7,000 cancelled flights over five days. Roughly 500 million dollars pre-tax. Two and three tenths of a point of operating margin, forty-five cents of earnings per share.

CrowdStrike, meanwhile, booked 39.1 million dollars of incident-related costs, net, across the first nine months of fiscal 2025, for every affected customer combined. One airline absorbed close to thirteen times the vendor’s entire nine-month cost of the event. CrowdStrike’s fiscal 2025 annual report sets out the July 19 Incident at length in its risk factors, and records that it agreed to provide incentives including subscription period extensions, discounts or promotional modules. Subscription extensions do not refund cancelled flights.

That asymmetry shows up in the aggregate. Uptime Institute has surveyed this ground for nine years, and its Annual Outage Analysis 2026 reports the sharpest version. Uptime Institute’s Figure 16, drawn from 94 responses, puts 43 per cent of respondents’ most recent significant outage under 100,000 dollars, 37 per cent between 100,000 dollars and a million, and 20 per cent above a million. Uptime Institute’s announcement of the 2026 edition notes that 57 per cent put their most recent major outage above 100,000 dollars, and that one in five exceeded a million for the second consecutive year.

The cause distribution matters more than the totals. Uptime Institute’s Figure 5, covering 403 respondents on information technology service outages, attributes 23 per cent to a third-party information technology service, the largest category, ahead of networking at 21 per cent and software at 19 per cent. Across nine years of publicly reported incidents, Uptime Institute attributes about two-thirds to third-party information technology and data centre providers.

Exhibit 2
Across 1,432 host-hours of a live fleet, processor load never came close to the capacity purchased
Processor utilisation across the two production hosts carrying an anonymised fleet of 19 production applications Wirecode manages, measured in 358 four-hour windows over thirty days.
Median four-hour window, 11.6%
Mean four-hour window, 12.5%
95th percentile window, 20.1%
Busiest window observed, 34.0%
Provisioned capacity, 100%
Three of 358 windows exceeded 30 per cent. Free memory on the larger host never fell below 2.76 GB of 8 GB.
Source. Wirecode first-party fleet telemetry pulled 25 August 2026 via the Cloudways monitoring application programming interface, aggregated and anonymised. Wirecode exhibit.

Wirecode’s own fleet averages 12.5 per cent processor load

That exhibit is Wirecode’s own fleet, and it undercuts the sales pitch.

Nineteen production applications on two hosts, one on DigitalOcean in Amsterdam, one on Vultr in Frankfurt. Across 1,432 host-hours of Wirecode’s fleet, mean processor load ran at 12.5 per cent and the busiest four-hour window on either machine touched 34.0 per cent. Capacity has never been the binding constraint here, and a larger instance or an extra nine would change nothing.

Here is the part that cuts against a tidy story. Seventeen of those nineteen applications sit on a single host. If that machine goes, the provider’s uptime commitment is beside the point, because the concentration was an architectural choice made on this side of the contract. No SLA in the world addresses it. A second host does, and it costs a fraction of what one day of correlated downtime across seventeen properties would cost their owners. That is the whole argument in one server bill.

HM Treasury named AWS, Microsoft, Google and Oracle critical third parties

The clearest evidence that service credits fail as resilience is that supervisors gave up on them and built direct oversight.

The European Union’s Digital Operational Resilience Act, Regulation (EU) 2022/2554, has applied since 17 January 2025. Its Article 30 sets mandatory contractual provisions for arrangements supporting critical or important functions. Point 3(a) requires full service level descriptions “with precise quantitative and qualitative performance targets“, point 3(e) grants unrestricted rights of access, inspection and audit, and point 3(f) requires exit strategies with mandatory transition periods. Credits appear nowhere in that list, where exit capability does.

The United Kingdom went further. The Financial Conduct Authority published Policy Statement 24/16 on critical third parties in November 2024, with rules in force from 1 January 2025 under powers granted by the Financial Services and Markets Act 2023. On 10 July 2026 HM Treasury used them, designating four cloud providers as critical third parties to the financial sector. The four are Microsoft Ireland Operations Limited, Google Cloud EMEA Limited, Amazon Web Services EMEA SARL and Oracle Corporation UK Limited. The announcement is careful about where responsibility stays. Financial firms, it says, remain responsible for managing risks arising from their third-party suppliers.

The nines still work as a filter on engineering quality

Three things cut against all of this.

The nines still separate real engineering populations. A provider committing to 99.5 per cent and one committing to 99.99 per cent describe different operations, and the gap is forty-three hours a year against fifty-two minutes. The number works poorly as compensation and well as a filter. Uptime Institute also reports outage rates falling per site for the fifth consecutive year, so platforms are improving even as each remaining failure costs more.

Architecture has hard limits of its own. During the October 2025 event, AWS recorded that Redshift customers in every region lost the ability to use identity credentials because of a defect that called an identity application programming interface in Northern Virginia, and that federated console sign-in failed outside that region. A multi-region deployment would have absorbed some of that day and none of the parts that ran through global endpoints. Anyone selling multi-region as an outage vaccine is overselling.

And the credit carries information even when the money is trivial. A provider whose schedule pays 100 per cent of the monthly fee below 95 per cent uptime is telling you where it thinks catastrophe begins. Read the bands as a disclosure of the vendor’s risk model, then stop expecting them to fund a recovery.

Do these four things before the next cloud hosting SLA renewal

Open the agreement and find the definition of Unavailable, then read it before you read the percentage. Work out whether the failure mode that would hurt you most, degraded latency, failed launches, an authentication outage, registers on the meter at all. Usually it will not, and that answer should change what you build.

Compose your own number. List every service a customer request touches, multiply the published commitments, and compare that with what your board has been told. Four dependencies at 99.99 per cent give you 99.96 per cent, and the difference between those two figures is where incident reviews start.

Read the incident history rather than the status badge. AWS, Cloudflare and Azure publish detailed post-event summaries describing failure modes, blast radius and the remediation that followed with a candour no sales deck matches. Two hours with three years of post-mortems beats any comparison table, and the same discipline applies when agent tooling reprices the software stack.

Then price the second host. Leave the second region for later. The second database replica, the tested restore, the failover rehearsed under real load. Against Uptime Institute’s finding that one outage in five now costs more than a million dollars, redundancy that removes a single point of failure is the cheapest insurance on the list, and unlike a service credit it pays out in availability rather than in next month’s invoice. That trade, and the discipline behind it, is what Wirecode builds and runs for growing companies.

References

  1. Amazon Web Services. Amazon Compute Service Level Agreement, last updated 25 May 2022. https://aws.amazon.com/compute/sla/
  2. Amazon Web Services. Amazon S3 Service Level Agreement, last updated 28 November 2023. https://aws.amazon.com/s3/sla/
  3. Google Cloud. Compute Engine Service Level Agreement, last modified 4 March 2025. https://cloud.google.com/compute/sla
  4. Microsoft. SLA for Virtual Machines, the edition published for Azure operated by 21Vianet, which is the copy Microsoft renders as readable text. The commitment and the credit bands quoted in this article are read from that edition. Microsoft’s global licensing portal serves the equivalent worldwide document only as a downloadable file. https://www.azure.cn/en-us/support/sla/virtual-machines/
  5. Cloudflare. Business Plan Service Level Agreement. https://www.cloudflare.com/business-sla/
  6. Amazon Web Services. Summary of the Amazon DynamoDB Service Disruption in the Northern Virginia (US-EAST-1) Region, October 2025. https://aws.amazon.com/message/101925/
  7. Cloudflare. Cloudflare outage on November 18, 2025. https://blog.cloudflare.com/18-november-2025-outage/
  8. Uptime Institute. Annual Outage Analysis 2026, May 2026. https://datacenter.uptimeinstitute.com/rs/711-RIA-145/images/2026.AnnualOutageAnalysis.pdf
  9. Uptime Institute. Uptime Announces Annual Outage Analysis Report 2026, 13 May 2026. https://uptimeinstitute.com/about-ui/press-releases/uptime-announces-annual-outage-analysis-report-2026
  10. Delta Air Lines, Inc. September Quarter 2024 Financial Results, Exhibit 99.1 filed with the US Securities and Exchange Commission. https://www.sec.gov/Archives/edgar/data/0000027904/000168316824007033/delta_ex9901.htm
  11. CrowdStrike Holdings, Inc. Third Quarter Fiscal Year 2025 Financial Results, Exhibit 99.1 filed with the US Securities and Exchange Commission. https://www.sec.gov/Archives/edgar/data/1535527/000153552724000024/crwd-20241126xex991.htm
  12. CrowdStrike Holdings, Inc. Form 10-K, filed 10 March 2025. https://ir.crowdstrike.com/static-files/26d52307-1e24-44b7-bbc9-498b534480e2
  13. European Union. Regulation (EU) 2022/2554 on digital operational resilience for the financial sector. https://eur-lex.europa.eu/eli/reg/2022/2554/oj/eng
  14. Regulation (EU) 2022/2554, Article 30, Key contractual provisions, consolidated text. https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32022R2554
  15. Financial Conduct Authority. PS24/16, Operational resilience, critical third parties to the UK financial sector, 12 November 2024. https://www.fca.org.uk/publications/policy-statements/ps24-16-operational-resilience-critical-third-parties-uk-financial-sector
  16. HM Government. UK financial system strengthened with new safeguards for major technology providers, 10 July 2026. https://www.gov.uk/government/news/uk-financial-system-strengthened-with-new-safeguards-for-major-technology-providers
  17. Microsoft. Recommendations for defining reliability targets, Azure Well-Architected Framework. https://learn.microsoft.com/en-sg/azure/well-architected/reliability/metrics
  18. Amazon Web Services. Shared Responsibility Model. https://aws.amazon.com/compliance/shared-responsibility-model/
  19. European Banking Authority. Interactive Single Rulebook, Regulation (EU) 2022/2554, Article 30, the supervisor’s landing page for the article, which links out to the full text. https://www.eba.europa.eu/regulation-and-policy/single-rulebook/interactive-single-rulebook/17755
ivan
ivan

Leave a Reply

Your email address will not be published. Required fields are marked *