A payment authorization process that stops for an hour costs a bank money and customers. A monthly reporting process that starts an hour late costs almost nothing. Both can run on Camunda, and they shouldn't share the same recovery plan or the same infrastructure bill.
With the 8.10 release, Camunda's high availability and disaster recovery options span every level, from a single failed broker to the loss of an entire region. Each option comes with published recovery targets and a clear cost profile, in both Camunda 8 Software as a Service (SaaS) and Self-Managed deployments. You protect each process as much as it needs and pay for no more than that. This is enterprise readiness in practice: the platform beneath your agents is as resilient as the processes running on it.
Two numbers behind every recovery plan
Every recovery plan comes down to two numbers. The recovery time objective (RTO) is how long a service can be down before it must be running again. The recovery point objective (RPO) is how much recent data you can afford to lose, measured in time.
If your last backup ran at 1 a.m. and a failure hits at 3 a.m., that event costs you two hours of data. An RPO of zero means no committed data is lost.
High availability keeps a system running through everyday failures, such as a crashed server or a restarted container. Disaster recovery covers the rarer, larger events: a lost region, corrupted data, or an operator mistake. Auditors increasingly ask you to state both numbers and prove them in a test.
High availability inside the engine
Zeebe, the workflow engine at the heart of Camunda, splits work into partitions and keeps copies of each partition on several brokers. The brokers form a peer-to-peer network with no single point of failure. A record counts as written only once a majority of copies holds it, so when a broker goes down, another takes over automatically, without human intervention or data loss.
That covers the failures every large deployment sees each week. It doesn't cover losing a whole location, and until now, Zeebe didn't know where its copies physically lived.
In Camunda 8.10 Self-Managed, operators can label each broker with its region, availability zone, or data center. Zeebe then spreads the copies so no single location holds a majority for any partition. In a correctly configured region with three availability zones, losing one zone leaves every partition able to keep processing, without paying for a second region. The same labels make three-region deployments possible.
Multi-region options in Self-Managed
Self-Managed trades that simplicity for control over where everything runs. Camunda documents three multi-region strategies for the Orchestration Cluster (Zeebe plus Operate, Tasklist, and the other core components), and 8.10 adds the third, backed by a relational database (RDBMS) or ElasticSearch.
Treat these as planning targets, not contractual commitments that will depend on your infrastructure and configurations. Cold Recovery depends on data volume, backup frequency, and how quickly your team restores, and the Dual-Region figure comes from internal operational tests.
Two changes in 8.10 make the third column possible.
- The first is the broker labeling described above: with brokers spread across three labeled regions, losing one leaves the survivors with the majority they need, so processing continues without a failover step.
- The second is support for asynchronously replicated relational databases, such as Amazon Aurora and PostgreSQL, as secondary storage. That's the database Operate, Tasklist, and your queries read from, while the authoritative record stays in Zeebe. Until now, zero data loss for that database across regions meant synchronous replication, which adds latency to every write.
- In 8.10, Zeebe keeps its own log history until the required number of database replicas confirm they have it. If the database fails over before replication catches up, Camunda replays the missing records from the Zeebe log with no manual data repair, and the engine keeps processing throughout. The cost is disk, since longer history increases broker disk usage, so size volumes for it and monitor replication lag.
Teams standardized on Amazon Elastic Container Service (Amazon ECS) get a new dual-region reference architecture on AWS Fargate, with Amazon Aurora Global Database for secondary storage and step-by-step failover and failback procedures. And restoring a Self-Managed cluster from backup no longer means restarting brokers. In-process restore runs through two API requests with no deployment changes, though processing pauses until it completes.
Recovery options in Camunda 8 SaaS
In SaaS, you choose resilience when you create a cluster. The cluster type sets uptime and recovery targets, and the cluster size, from 1x to 4x, sets capacity.
The uptime percentages are guaranteed, as defined in the Camunda Enterprise General Terms. The RTO and RPO figures are targets, provided on a best-effort basis. A customer-facing process belongs on Advanced, while a back-office process that can wait two hours runs comfortably on Standard.
Backups add a second layer. Enterprise customers can back up a whole cluster (Zeebe, Operate, Tasklist, and Optimize) as one consistent snapshot while it keeps running, on demand or on a schedule. Choose a dual-region backup location when you create the cluster, and backups replicate to a second region at no additional cost.
Since September, organization admins can restore a cluster from a completed backup themselves without opening a support ticket. Restore happens in place, so it overwrites the cluster's current data, and the cluster is unavailable until it finishes.
For a regional outage, cross-region disaster recovery now covers supported region pairs on Amazon Web Services (AWS) and Google Cloud. When you initiate recovery, Camunda provisions a new cluster in the secondary region and restores it from the replicated backups. No standby cluster runs before the outage, so your recovery point depends on how often you back up.
Matching resilience to your budget
The most resilient option may not be the right choice for every process. Start with the process: how long can it stop, and how much recent work can you afford to redo? Then you can pick the least expensive option that meets both answers.
That often means mixing tiers. A payment process that must survive a regional outage belongs on Dual-Region or three regions, while the reporting process from the beginning of this post can run on a Standard SaaS cluster or a single Self-Managed region with scheduled backups. With physical tenants, a Self-Managed platform team can host several business domains in one cluster, each with its own data store, identity provider, and backup and restore, instead of running a cluster per team. Compute is shared across tenants, so size for peak load.
The least expensive levers come first: dual-region backups in SaaS, availability zone labels within one region, and asynchronous instead of synchronous database replication. Existing deployments without topology labels keep their current placement, so you can move a process up a tier when it becomes more critical, and not before.
Whichever option you choose, rehearse the failure. Run the restore or the failover in a non-production environment, time it, and put the measured number in your runbook.
Why resilience matters more as agents take on real work
According to the 2026 Camunda State of Agentic Orchestration and Automation Report, 71% of organizations use AI agents, yet only 11% of use cases reached production last year. Enterprise readiness is a big part of that gap. A pilot project can tolerate a bad day, but a process that approves claims and moves money cannot. As companies re-engineer their processes around AI, the orchestration layer around those agents has to keep running through disruptions without losing work in flight.
That’s what 8.10 hardens. Camunda coordinates agents, people, and systems across the end-to-end process through outer orchestration (outside the agent). Inner orchestration (inside the agent) puts enforceable steps between an agent's reasoning and its actions, such as a human approval before a payment goes out. Camunda records every one of those steps as it happens. Our series on production agentic trust covers testing and observing the agents themselves; resilience is the enterprise-readiness pillar that keeps that record, and the in-flight work behind it, intact when infrastructure fails.
Ready to orchestrate? Compare the SaaS cluster types and the Self-Managed multi-region strategies, then talk to us about mapping your processes to the option that fits your recovery targets and your budget.



