惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

L
LINUX DO - 最新话题
I
InfoQ
V
V2EX
博客园 - 叶小钗
博客园 - 三生石上(FineUI控件)
爱范儿
爱范儿
T
Tailwind CSS Blog
大猫的无限游戏
大猫的无限游戏
IT之家
IT之家
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
阮一峰的网络日志
阮一峰的网络日志
博客园 - 聂微东
月光博客
月光博客
Hugging Face - Blog
Hugging Face - Blog
小众软件
小众软件
Last Week in AI
Last Week in AI
量子位
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
G
Google Developers Blog
MongoDB | Blog
MongoDB | Blog
博客园 - 司徒正美
腾讯CDC
罗磊的独立博客
有赞技术团队
有赞技术团队
Jina AI
Jina AI
T
The Exploit Database - CXSecurity.com
C
CERT Recently Published Vulnerability Notes
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
V
Visual Studio Blog
H
Help Net Security
AWS News Blog
AWS News Blog
C
Cybersecurity and Infrastructure Security Agency CISA
Recorded Future
Recorded Future
GbyAI
GbyAI
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Attack and Defense Labs
Attack and Defense Labs
Webroot Blog
Webroot Blog
S
SegmentFault 最新的问题
Know Your Adversary
Know Your Adversary
The GitHub Blog
The GitHub Blog
I
Intezer
博客园 - Franky
云风的 BLOG
云风的 BLOG
N
News and Events Feed by Topic
P
Proofpoint News Feed
D
DataBreaches.Net
Google DeepMind News
Google DeepMind News
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
G
GRAHAM CLULEY

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
OCI Full Stack Disaster Recovery (FSDR) Deep Dive: Architecture, Switchover, Failover, and Recovery Workflows
Bonthu Durga · 2026-05-20 · via DEV Community

Introduction

Disaster recovery in cloud environments is no longer limited to restoring virtual machines or recovering storage volumes. Modern enterprise applications depend on tightly coupled compute, networking, databases, load balancers, DNS, and application dependencies.

OCI Full Stack Disaster Recovery (FSDR) introduces orchestration-driven recovery workflows that coordinate infrastructure and application recovery across regions while minimizing operational risk and downtime.

FSDR IS NOT BACKUP
Backup protects data.
Disaster recovery restores application continuity.

FSDR focuses on orchestrating complete application recovery, not only restoring individual resources.

This blog explains the deeper architecture and operational concepts behind OCI FSDR, including recovery orchestration, dependency sequencing, traffic redirection, resiliency engineering, and enterprise recovery design patterns.

Traditional backups help restore files or databases, but enterprise applications require coordinated recovery across multiple infrastructure layers.

Example:

Database restored successfully
→ application services unavailable
→ load balancer returns errors
→ business outage continues

Architecture Overview

FSDR setup follows a simple two-region design. The primary region hosts the live application stack, including compute, load balancer, database, and storage components. The secondary region keeps the standby resources ready for recovery.

All these resources are placed into Disaster Recovery Protection Groups, which help FSDR understand what belongs together. Once the groups are created, recovery plans can be built to define the exact order of actions during switchover or failover. This makes disaster recovery far more predictable and much easier to test.

Enterprise Multi-Region Disaster Recovery Architecture

Primary and DR Region Design

Users


Primary OCI Region

├── Public Load Balancer
├── Web Tier
├── Application Tier
├── Database Tier
└── Storage Layer

Replication / Synchronization


Disaster Recovery Region

├── Standby Infrastructure
├── Recovery Workflows
├── Replicated Data
└── Traffic Redirection

Understanding Recovery Orchestration

One of the most important concepts in FSDR is orchestration.

FSDR does not recover everything simultaneously.

Instead, recovery occurs in dependency-aware orchestration stages.

Example Recovery Workflow

  1. Validate DR environment
  2. Attach replicated storage
  3. Recover database services
  4. Validate database health
  5. Start application services
  6. Start web services
  7. Update load balancer routing
  8. Redirect traffic
  9. Validate application response

This sequencing reduces operational failures during recovery events.

Why Dependency Order Matters

Application continuity depends heavily on startup sequencing.

Incorrect startup order is one of the most common disaster recovery failures.

Example:

Web tier starts before database recovery
→ application connection failures
→ unstable service state

OCI FSDR helps coordinate these dependencies through orchestrated recovery execution.

Traffic Flow During Disaster Recovery

Understanding traffic movement during failover is critical.

Normal Traffic Flow
Users


Primary Load Balancer


Application Stack

Disaster Event
Primary region unavailable
Recovery Flow
FSDR initiates recovery workflows
→ DR region activated
→ services validated
→ traffic redirected
→ application restored

Switchover vs Failover

Although these terms are often used interchangeably, operationally they are very different.

Switchover

Switchover is a controlled transition between regions.

Controlled migration with synchronized application state.

Typical use cases:

✔ Planned maintenance
✔ DR drills
✔ Infrastructure migration
✔ Region transition testing

Failover

Failover occurs during an actual disruption.

Emergency recovery during infrastructure failure.

Typical use cases:

✔ Region outage
✔ Critical disaster
✔ Connectivity failure
✔ Infrastructure incident

Key Operational Insight

Switchover focuses on continuity.
Failover focuses on survivability.

Recovery Objectives in Enterprise DR

Disaster recovery design is heavily influenced by two key metrics.

RTO (Recovery Time Objective)
Maximum acceptable downtime.

Example:

Application must recover within 15 minutes.
RPO (Recovery Point Objective)
Maximum acceptable data loss window.

Example:

5-minute replication lag accepted.

Important Design Insight

Lower RTO and RPO increase infrastructure complexity and operational cost.

This is one of the biggest design tradeoffs in enterprise disaster recovery.

Observability During Disaster Recovery

Recovery orchestration without observability creates blind operational recovery.

Monitoring and validation are essential during DR events.

Critical observability areas include:

✔ Replication health
✔ Recovery progress
✔ Application validation
✔ Service health
✔ Traffic routing
✔ Error monitoring

Without proper validation, infrastructure may recover while applications remain unavailable.

Real Enterprise Scenario

Consider a multi-tier banking application deployed across OCI regions.

Architecture:

Internet


Public Load Balancer


Web Tier


Application Tier


Database Tier

Disaster Recovery Deployment Models

One of the most important architectural decisions in disaster recovery design is selecting the appropriate DR deployment model.

The choice depends on:

✔ Recovery speed requirements
✔ Business criticality
✔ Infrastructure cost
✔ Operational complexity
✔ Acceptable downtime
✔ Recovery objectives (RTO/RPO)

Enterprise DR strategies are commonly divided into:

✔ Cold DR
✔ Warm DR
✔ Hot DR

old Disaster Recovery (Cold DR)
What is Cold DR?

Cold DR is the most cost-optimized disaster recovery model.

Simple explanation:

Infrastructure is created only during disaster recovery events.

In this model, the DR region does not continuously run the full application stack.

Instead:

✔ Backups are stored
✔ Configurations are maintained
✔ Infrastructure is provisioned during disaster
Cold DR Architecture
Primary Region

├── Running Production Environment


DR Region

├── Backup Storage
├── Infrastructure Templates
└── Minimal Active Resources

**Cold DR Workflow

During disaster:**

  1. Disaster detected
  2. Infrastructure provisioned in DR region
  3. Storage restored
  4. Database recovered
  5. Application deployed
  6. Traffic redirected

Warm Disaster Recovery (Warm DR)

What is Warm DR?

Warm DR provides a balance between recovery speed and infrastructure cost.

Simple explanation:

A partially running standby environment exists in the DR region.

Some infrastructure components remain active continuously.

Example:

✔ Database replication active
✔ Standby compute available
✔ Networking preconfigured
✔ Application services partially ready
Warm DR Architecture
Primary Region

├── Fully Active Environment

Replication


DR Region

├── Standby Database
├── Preconfigured Networking
├── Minimal Compute
└── Recovery Automation

Warm DR Workflow

During disaster:

  1. DR database promoted
  2. Additional compute started
  3. Application services activated
  4. Load balancer updated
  5. Traffic redirected

Hot Disaster Recovery (Hot DR)

What is Hot DR?

Hot DR is the most advanced disaster recovery model.

Simple explanation:

A fully active standby environment continuously runs in the DR region.

Both regions remain operational simultaneously.

The DR region is always ready for immediate failover.

Hot DR Architecture
Primary Region

├── Active Production Stack

Real-Time Replication


DR Region

├── Fully Active Standby Stack
├── Running Applications
├── Active Networking
└── Immediate Traffic Readiness

**Hot DR Workflow

During disaster:**

  1. Primary outage detected
  2. Traffic immediately redirected
  3. DR environment already operational
  4. Minimal recovery delay

During disaster:

Primary region unavailable
→ FSDR executes recovery orchestration
→ DR database activated
→ application services recovered
→ traffic redirected
→ banking services restored

Common Disaster Recovery Failures

Many DR failures occur during orchestration and validation rather than infrastructure provisioning.

Common issues include:

✔ Missing dependency mapping
✔ DNS still pointing to failed region
✔ Replication lag ignored
✔ Application validation skipped
✔ Untested DR workflows
✔ Incorrect startup sequencing

Critical Operational Insight
Most DR failures occur during orchestration and validation, not infrastructure provisioning.
Why OCI FSDR Matters

Cloud resiliency is no longer only an infrastructure recovery problem.

Modern disaster recovery is an application orchestration challenge.

OCI FSDR helps organizations move from:

Manual recovery

Automated resiliency engineering

through coordinated recovery workflows across regions.

Production Best Practices
✔ Perform regular DR drills
✔ Validate application dependencies
✔ Continuously monitor replication
✔ Test traffic failover procedures
✔ Maintain updated recovery documentation
✔ Validate application health after recovery
✔ Separate production and DR environments

Oracle FSDR official Doc :

Conclusion

OCI Full Stack Disaster Recovery enables organizations to orchestrate application-aware disaster recovery workflows across OCI regions.

By coordinating dependency sequencing, traffic routing, recovery validation, and service orchestration, FSDR helps reduce downtime and operational complexity during disaster events.

Modern disaster recovery is no longer just about recovering infrastructure — it is about restoring complete business continuity through intelligent orchestration and resiliency engineering.