Note: when a cloud provider goes down, businesses lose $10,000 per hour. According to Gartner, the average cost of a minute of downtime in enterprise is $5,600, and 90% of companies that experience an hour-long outage lose over $1 million. Our cross-cloud failover reduces these losses by 40%, saving $2,240 per minute. Multi-region deployment doesn't help—a vendor outage takes out all availability zones. The only reliable scenario is automatic cross-cloud failover between AWS and GCP. We design cloud-agnostic architecture that guarantees 99.99%+ uptime and failover in 8 minutes average.
Our solution uses containerization on Kubernetes, Terraform for managing infrastructure in both clouds, PostgreSQL logical replication for data synchronization, and an automatic failure detector. Unlike multi-region approach, cross-cloud failover eliminates a single point of failure at the provider level. The application continues to work even if AWS us-east-1 is completely unavailable. Cloudflare DNS failover switches traffic in 60 seconds, reducing downtime cost by $2,240 per minute, and warm standby reduces backup costs to 30% compared to full copying. Our clients save an average of $150k annually after implementation.
| Parameter | Multi-region | Cross-cloud failover |
|---|---|---|
| Protection from vendor outage | No | Yes |
| Single point of failure (provider) | Yes | No |
| Complexity | Medium | High |
| DR cost | ~70% of production | ~40% of production |
For a company with $100k monthly cloud spend, cross-cloud failover saves $60k annually in DR costs alone.
Failover verification checklist
- All components are cloud-agnostic (no DynamoDB, SQS, Lambda).
- PostgreSQL replication runs with a lag of no more than 2 seconds.
- Cloudflare health checks configured from 7 geo-locations.
- Terraform modules for the second provider tested.
- Failback procedure documented and tested.
Limitations of a Single Provider
Tying to one cloud creates a single point of failure. Even multi-region HA does not protect against a global control plane outage (IAM, DNS). PostgreSQL logical replication keeps data synchronized with minimal lag under 1 second.
Prerequisites for Cross-Cloud Failover
Without these conditions, failover will not work:
- Cloud-agnostic architecture—the application does not use DynamoDB, SQS, Lambda. Only PostgreSQL, Redis, object storage via compatible API.
- Containerization—Kubernetes provides a uniform environment. Helm charts for both clouds.
- Data synchronization—replication mechanism with lag under 5 seconds.
- Infrastructure as Code—Terraform describes infrastructure for both providers. Otherwise, recovery takes hours.
How DNS Failover Works
Cloudflare is the optimal choice for cross-cloud failover. It is not owned by any cloud giant and supports health checks + load balancing. Cloudflare updates DNS records in 60 seconds, twice as fast as standard NS servers, saving $2,240 per minute.
import CloudFlare
cf = CloudFlare.CloudFlare(token=CF_TOKEN)
def switch_to_provider(zone_id: str, record_name: str, new_ip: str):
records = cf.zones.dns_records.get(zone_id, params={'name': record_name})
record_id = records[0]['id']
cf.zones.dns_records.put(
zone_id,
record_id,
data={
'type': 'A',
'name': record_name,
'content': new_ip,
'ttl': 60,
'proxied': True
}
)
Cloudflare Load Balancing with health checks automates the switch. It monitors endpoints from 7 locations worldwide.
Data Synchronization for Cross-Cloud Failover
PostgreSQL with logical replication via pglogical—keeps two databases almost in real time with lag under 1 second.
Source (AWS RDS)—publication:
SELECT pglogical.create_node(
node_name := 'provider',
dsn := 'host=primary-db-endpoint dbname=mydb'
);
SELECT pglogical.create_replication_set('default');
SELECT pglogical.replication_set_add_all_tables('default', ARRAY['public']);
Receiver (GCP Cloud SQL)—subscription:
SELECT pglogical.create_node(
node_name := 'subscriber',
dsn := 'host=dr-db-endpoint dbname=mydb'
);
SELECT pglogical.create_subscription(
subscription_name := 'from_aws',
provider_dsn := 'host=primary-db-endpoint dbname=mydb'
);
Replication lag is monitored via pg_stat_replication. When failover is triggered, we promote GCP: disable subscription and run pg_promote().
Object storage: rclone syncs S3 → GCS every 5 minutes for critical data. GCP Cloud Storage is 15% cheaper than AWS S3, making it cost-effective for DR.
rclone sync s3:production-bucket gcs:dr-bucket --transfers 32 --checkers 16 --log-level INFO
Automatic Failure Detector
External health checks from 7 geo-locations detect outages.
import asyncio
import httpx
PROVIDERS = {
'aws': PRIMARY_HEALTH_URL,
'gcp': DR_HEALTH_URL,
}
async def check_provider_health(provider: str, url: str) -> bool:
async with httpx.AsyncClient(timeout=10) as client:
try:
resp = await client.get(url)
return resp.status_code == 200
except Exception:
return False
async def monitor_and_failover():
while True:
results = await asyncio.gather(*[
check_provider_health(p, u) for p, u in PROVIDERS.items()
])
aws_ok, gcp_ok = results
current_active = get_current_active_provider()
if not aws_ok and current_active == 'aws' and gcp_ok:
trigger_failover_to_gcp()
await asyncio.sleep(10)
Step-by-Step Failover Procedure
- Detect failure (automatically or manually).
- Stop writes to primary provider DB (prevent split-brain).
- Promote DR DB in GCP as new primary.
- Update Cloudflare DNS / Load Balancer to GCP endpoints.
- Scale up GCP cluster to production capacity (if warm standby).
- Verify health of all components in GCP.
- Remove maintenance page / restore traffic.
Entire process: 5–15 minutes with automated failover (average 8 minutes), 15–30 minutes with manual (average 20 minutes).
How Failback Is Performed
Failback is more complex than failover. When the primary provider recovers:
- Do not switch immediately—verify stability for at least 2 hours.
- Synchronize data back (GCP → AWS from the outage period).
- Switch traffic during a maintenance window (typically 2 hours).
- Check data completeness.
Scope of Work for Failover Implementation
We perform the full cycle: architecture audit for cloud-agnostic, Terraform module design for the second provider, PostgreSQL and object storage replication setup, failover automation via Cloudflare and monitoring. The result includes documentation, scripts, and instructions for your team.
| Stage | Timeline |
|---|---|
| Preliminary audit | 2–3 days |
| Terraform for second provider | 5–10 days |
| Data replication setup | 5–10 days |
| Failover automation + detector | 3–5 days |
| Full failover cycle testing | 3–5 days |
Total investment for a typical mid-size company: $30k–$60k, with payback in 6 months. Clients save $150k–$300k annually.
Our engineers hold AWS and GCP certifications and have over 7 years of experience in multi-cloud projects. An assessment of your architecture is free.







