<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/how-to-keep-your-mission-critical-cloud-workloads-running-da0hrpqbm" -->

---
title: How to keep your mission-critical cloud workloads running
description: Building resilient, highly available mission-critical cloud workloads requires understanding the cloud shared responsibility model and addressing four failure...
canonical: https://daily.dev/posts/how-to-keep-your-mission-critical-cloud-workloads-running-da0hrpqbm
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: How to keep your mission-critical cloud workloads running | daily.dev
og:description: Building resilient, highly available mission-critical cloud workloads requires understanding the cloud shared responsibility model and addressing four failure...
og:url: https://daily.dev/posts/how-to-keep-your-mission-critical-cloud-workloads-running-da0hrpqbm
og:image: https://api.daily.dev/og/posts/dA0hrpqbM.png
og:image:alt: How to keep your mission-critical cloud workloads running
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# How to keep your mission-critical cloud workloads running

**[InfoWorld](https://daily.dev/sources/infoworld)** · 11 min read · 0 upvotes · 0 comments

## Summary

Building resilient, highly available mission-critical cloud workloads requires understanding the cloud shared responsibility model and addressing four failure categories: single points of failure, excessive load, misconfigurations/bugs, and shared fate. Four architectural pillars are needed: clustering (increasingly software-based SANless clustering rather than hardware SANs), data replication (sync or async between nodes), automated failover, and disaster recovery with geographically distant sites using asynchronous replication. The piece uses SIOS LifeKeeper and SIOS DataKeeper as examples of application-aware SANless clustering software, and notes downtime can cost $1 million or more per hour for large enterprises, with 90% of organizations reporting at least $300,000 per hour in costs. Software-based clustering also enables safer rolling patch updates by allowing secondary nodes to be patched and verified before touching primary systems.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://www.infoworld.com/article/4211691/how-to-keep-your-mission-critical-cloud-workloads-running.html>

## Questions this post answers

### What is the difference between RPO and RTO in disaster recovery planning?

Recovery point objective (RPO) is the maximum acceptable amount of data loss measured in time, while recovery time objective (RTO) is the maximum acceptable amount of time a service can remain down after a disaster. Both metrics are used to measure and evaluate how well a disaster recovery plan restores service and data after an incident.

_Teams weighing RPO and RTO targets for their systems can track architecture guidance like this on daily.dev._

### How much does an hour of downtime typically cost an enterprise?

An hour of downtime can cost $1 million or more depending on company size and industry, and 90% of organizations report costs of at least $300,000 per hour. Beyond direct revenue loss, downtime also causes soft costs like reputation damage, lost productivity, and missed transactions.

_Anyone building a business case for high-availability investment can follow cost data like this on daily.dev._

### What is SANless clustering and why is it used instead of hardware-based SAN clustering in the cloud?

SANless clustering is a software-based approach to connecting systems for high availability, replacing the traditional hardware storage area network (SAN) approach. It works better in cloud and hybrid environments because SANs are expensive and inflexible there, and application-aware SANless software like SIOS LifeKeeper can properly fail over workloads such as SQL Server, SAP, or Postgres by understanding how those applications need to restart cleanly.

_Engineers comparing clustering approaches for cloud resilience can find architecture writeups like this on daily.dev._

## Similar posts on daily.dev

- [DevOps & SaaS Downtime: The High \(and Hidden\) Costs for Cloud-First Businesses](https://daily.dev/posts/devops-saas-downtime-the-high-and-hidden-costs-for-cloud-first-businesses-hfvfar8hj) · The Hacker News · 0 upvotes · 0 comments
- [Reliability is a Product Decision](https://daily.dev/posts/reliability-is-a-product-decision-wpmua7ypy) · Erlang Solutions · 1 upvotes · 0 comments
- [Why cloud outages are such a stubborn problem](https://daily.dev/posts/why-cloud-outages-are-such-a-stubborn-problem-kh6xbgqoz) · InfoWorld · 0 upvotes · 0 comments
- [Don’t waste your next cloud outage](https://daily.dev/posts/don-t-waste-your-next-cloud-outage-cm5rgaojt) · InfoWorld · 0 upvotes · 0 comments

---

Tags: [#cloud](https://daily.dev/tags/cloud)

[View this post on daily.dev](https://daily.dev/posts/how-to-keep-your-mission-critical-cloud-workloads-running-da0hrpqbm)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"How to keep your mission-critical cloud workloads running","url":"https://daily.dev/posts/how-to-keep-your-mission-critical-cloud-workloads-running-da0hrpqbm","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/how-to-keep-your-mission-critical-cloud-workloads-running-da0hrpqbm"},"datePublished":"2026-09-03T09:05:43.915Z","dateModified":"2026-09-03T09:18:07.527Z","description":"Building resilient, highly available mission-critical cloud workloads requires understanding the cloud shared responsibility model and addressing four failure...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/58e61bba04d8837cf8cbf5f273fad5e5?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/58e61bba04d8837cf8cbf5f273fad5e5?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"InfoWorld","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"InfoWorld","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/bf6d68a999064029b0bb09aa6268f1f3","url":"https://daily.dev/sources/infoworld"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/how-to-keep-your-mission-critical-cloud-workloads-running-da0hrpqbm","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"cloud","timeRequired":"PT11M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"InfoWorld","item":"https://daily.dev/sources/infoworld"},{"@type":"ListItem","position":3,"name":"How to keep your mission-critical cloud workloads running"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/how-to-keep-your-mission-critical-cloud-workloads-running-da0hrpqbm#faq","mainEntity":[{"@type":"Question","name":"What is the difference between RPO and RTO in disaster recovery planning?","acceptedAnswer":{"@type":"Answer","text":"Recovery point objective (RPO) is the maximum acceptable amount of data loss measured in time, while recovery time objective (RTO) is the maximum acceptable amount of time a service can remain down after a disaster. Both metrics are used to measure and evaluate how well a disaster recovery plan restores service and data after an incident. Teams weighing RPO and RTO targets for their systems can track architecture guidance like this on daily.dev."}},{"@type":"Question","name":"How much does an hour of downtime typically cost an enterprise?","acceptedAnswer":{"@type":"Answer","text":"An hour of downtime can cost $1 million or more depending on company size and industry, and 90% of organizations report costs of at least $300,000 per hour. Beyond direct revenue loss, downtime also causes soft costs like reputation damage, lost productivity, and missed transactions. Anyone building a business case for high-availability investment can follow cost data like this on daily.dev."}},{"@type":"Question","name":"What is SANless clustering and why is it used instead of hardware-based SAN clustering in the cloud?","acceptedAnswer":{"@type":"Answer","text":"SANless clustering is a software-based approach to connecting systems for high availability, replacing the traditional hardware storage area network (SAN) approach. It works better in cloud and hybrid environments because SANs are expensive and inflexible there, and application-aware SANless software like SIOS LifeKeeper can properly fail over workloads such as SQL Server, SAP, or Postgres by understanding how those applications need to restart cleanly. Engineers comparing clustering approaches for cloud resilience can find architecture writeups like this on daily.dev."}}]}
```

