— THE ELIS RELIABILITY INDEX
An annual benchmark of cloud-provider incident transparency, written by the engineers who carry the pager.
Five years of measurement across AWS, GCP, and Azure — published every quarter so SRE leaders have a single, independent number to put in front of procurement, the board, and their own on-call rotation.
— HEADLINE FIGURES
Four numbers that anchor the Q4 2024 Index.
Drawn from 180+ managed production environments and a public-incident audit of the three hyperscalers. Each figure is reproducible from the methodology appendix.
— RECENT FROM THE ENGINEERING DESK
Five long-form essays, curated for the engineer currently on call.
No release notes, no product news. Just the post-mortems, design notes, and architecture decisions our senior engineers wrote down because the next operator needed to read them.
Failover without the fiction: what we learned running active-active across three regions on AWS and GCP
A field report from four regulated customers on the assumptions that quietly broke when DR drills stopped being annual theater and became a Tuesday-morning audit finding.
By Lena Okafor, Principal Cloud Engineer · Issue 04Zero-downtime rollouts at 4,000 pods: how ELIS-Forge replaced our canary playbook
The reference architecture we open-sourced last quarter, the three failure modes it explicitly defends against, and the migration path we used to roll it to a Fortune 100 logistics carrier.
By Marcus Vela, Staff SRE · Issue 05Where the 67% goes: a line-by-line autopsy of a $4.2M annual AWS bill
The three categories of waste we find in roughly nine out of ten inherited environments — and the one category every CFO asks about that almost never matters.
By Priya Anand, Director of FinOps · Issue 05HIPAA, PCI, and FedRAMP Moderate in a single engagement: the control matrix that makes it possible
Our in-house compliance desk shares the deduplicated control map that lets a single architecture pass three audits without re-engineering the data plane between them.
By David Reinholt, Compliance Architect · Issue 04The 11-minute outage: a public post-mortem of an ELIS-operated payments platform
A customer-approved, NDA-cleaned reconstruction of the Q3 incident that drove our MTTR benchmark down to 11 minutes — including the two decisions we would change.
By the ELIS NOC · Issue 05— METHODOLOGY
CHAPTER 01 / 02
What we measure.
The Reliability Index tracks four pillar metrics each quarter. Measured uptime is calculated from ELIS-operated environments using independent synthetic checks from three geographic vantage points — never from a provider's self-reported SLA credit table. Mean time to resolution is timed from the first customer-impacting alert to the moment the last dependency returned its green health check, and excludes the time a ticket spends waiting on a human to read it. Incident-transparency scores provider post-mortems on a 0–100 weighted rubric covering disclosure latency, root-cause specificity, customer-impact quantification, and remediation commitment. Cloud-spend variance is the percentage reduction in monthly invoice value achieved within the first 90 days of an ELIS FinOps engagement, measured against the customer's pre-engagement baseline.
CHAPTER 02 / 02
How it's sourced.
The dataset is drawn from three independent sources: telemetry from the 180+ production environments ELIS actively operates, a quarterly audit of public provider post-mortems and status-page histories, and the anonymized bills submitted by customers participating in our FinOps cohort. The methodology appendix — including the weighting rubric, the geographic distribution of vantage points, and the exact formula for the 99.995% uptime figure — is published in full at the back of each annual report and is independently reviewed by the engineering advisory board before release.
The supporting reference architecture, ELIS-Forge — a zero-downtime Kubernetes rollout pattern downloaded 38,000+ times from GitHub — is released under the same methodology chapter so readers can audit our operational defaults against the claims in the Index. Where a number in this issue cannot be reproduced from public sources, we mark it as internal-only and exclude it from the headline figures.
SUPPORTING DOCUMENTS
- A. Reliability Index Methodology Appendix, v3.2 — published Q4 2024
- B. ELIS-Forge reference architecture — github.com/elisdc/forge
- C. Incident-Transparency Rubric, weighted scoring sheet
— FOR THE SRE READING THIS ON A MONDAY MORNING
Five questions we get from the people who actually carry the pager.
Answers written by the engineering desk, not marketing. If yours isn't here, write to us at the address in the footer — a senior engineer reads every inbound.
01 Who funds the Reliability Index? Is this a marketing vehicle for ELIS?
ELIS funds the Index directly out of engineering budget — there is no sponsor logo, no sponsored placement, and no paid placement in the ranking. The dataset is drawn from environments we operate and from public provider disclosures; the editorial is written by the engineering desk and reviewed by an external advisory board before release. We disclose this prominently because the alternative — an "independent" benchmark secretly funded by a vendor — is exactly the kind of report we built this to replace.
02 Where can I download the prior years' Index reports?
Issues 01 through 04 are available as a single PDF bundle on the Insights index. Each annual issue is a self-contained report — the methodology changed only twice in five years (v2.0 in 2022, v3.0 in 2024), and the changelog is included so you can compare numbers across years without false precision. If you need a specific quarter's raw dataset for a board deck, email the desk and we will send it.
03 What changed in the methodology between Issue 04 and Issue 05?
Two things. First, we replaced the provider self-report weighting in the incident-transparency score with a three-vantage-point independent probe — this brings the score closer to what an SRE actually experiences during an outage. Second, we tightened the FinOps variance definition to exclude one-time credits and refunds; the headline 67% number is therefore slightly lower in Issue 05 than it would have been under the v2.0 rubric. The full changelog, with worked examples, is in the methodology appendix.
04 Can I cite the Index in a board deck or procurement RFP?
Yes. Cite it as ELIS Reliability Index, Issue 05, Q4 2024 — methodology v3.2. The appendices include a citation block formatted for both academic and procurement use, and the dataset is license-cleared for internal corporate use. If you need a signed rights letter for a regulated RFP, the desk can issue one within a business day.
05 How do I contribute an incident report to the next issue?
Open a pull request against the ELIS-Forge repository or email the desk with a sanitized post-mortem. We accept submissions from any operator — ELIS customer or not — provided the report is technically specific (root cause, blast radius, remediation, time-to-detect, time-to-resolve) and free of customer-identifying detail. Submissions for Issue 06 close on the last business day of February 2025; selected entries are reviewed by the advisory board and credited in the issue.
— NEXT STEP
Bring your own incident data, cloud bill, or migration plan.
A 30-minute Architecture Review with a Principal Cloud Engineer — no sales deck, no qualification script, no follow-the-sun handoff to a junior AE. You leave the call with a written assessment of the three things we would change in your environment this quarter.
FOUR OFFICES · ONE 24/7 NOC
- AUSTIN1801 Lavaca Street, Suite 410 · TX 78701 · HQ
- TORONTO120 Adelaide Street West, 26th Floor · ON M5H 1T1
- BERLINFriedrichstraße 68 · 10117 Berlin · Mitte
- SINGAPORE8 Marina Boulevard, Tower 1 · 018981