惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
Tailwind CSS Blog
C
CERT Recently Published Vulnerability Notes
V
Visual Studio Blog
O
OpenAI News
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
The Cloudflare Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Last Week in AI
Last Week in AI
Y
Y Combinator Blog
博客园 - 聂微东
L
Lohrmann on Cybersecurity
P
Proofpoint News Feed
Simon Willison's Weblog
Simon Willison's Weblog
G
GRAHAM CLULEY
AI
AI
S
Security @ Cisco Blogs
TaoSecurity Blog
TaoSecurity Blog
Jina AI
Jina AI
W
WeLiveSecurity
大猫的无限游戏
大猫的无限游戏
腾讯CDC
K
Kaspersky official blog
Hugging Face - Blog
Hugging Face - Blog
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
宝玉的分享
宝玉的分享
AWS News Blog
AWS News Blog
月光博客
月光博客
P
Palo Alto Networks Blog
小众软件
小众软件
V2EX - 技术
V2EX - 技术
罗磊的独立博客
V
Vulnerabilities – Threatpost
J
Java Code Geeks
H
Heimdal Security Blog
S
SegmentFault 最新的问题
博客园 - 【当耐特】
Cyberwarzone
Cyberwarzone
S
Schneier on Security
博客园_首页
T
The Exploit Database - CXSecurity.com
Attack and Defense Labs
Attack and Defense Labs
Forbes - Security
Forbes - Security
N
News | PayPal Newsroom
IT之家
IT之家
Project Zero
Project Zero
Help Net Security
Help Net Security
P
Privacy International News Feed
爱范儿
爱范儿
D
Darknet – Hacking Tools, Hacker News & Cyber Security
T
Threat Research - Cisco Blogs

OneUptime Blog

How to Monitor Azure App Services (PaaS) with OpenTelemetry Grafana Stack vs OneUptime: DIY Observability or Unified Platform? Your AI Workloads Are About to Blow Up Your Observability Bill The Great Observability Consolidation Is Here How to Write Custom Object Classes for Ceph How to Write Custom Ceph Manager Modules How to Write a ceph.conf Configuration File How to Use Rook-Ceph with OpenShift How to Use Rook-Ceph with Longhorn for Comparison How to Configure Volume Snapshot Class for RBD in Rook How to Configure VolumeReplicationClass Scheduling Intervals in Rook How to Set Up Volume Replication with Rook-Ceph How to Create Volume Group Snapshots with Rook CSI How to Visualize Ceph Network Performance in Grafana How to Enable Virtual Host-Style Bucket Access in Rook How to View Runtime Configuration via Admin Socket How to View Quota Settings and Update Stats in Ceph RGW How to View PG Scaling Recommendations with autoscale-status How to View PG Distribution via Admin Socket How to View Performance Metrics in the Ceph Dashboard How to View OSD Performance Counters in Ceph How to View Connection Status via Admin Socket How to View Ceph Cluster Summary Dashboard via CLI How to Version Control Rook-Ceph Configuration How to Version Control Ceph Infrastructure with Terraform How to Verify Kubernetes Node Requirements for Rook-Ceph Deployment How to Verify Health Before and After Rook Upgrades How to Verify Data Integrity with Deep Scrubbing How to Verify Complete Rook-Ceph Cleanup How to Verify Backup Integrity from Ceph Snapshots How to Use Rook-Ceph with Velero for Kubernetes Backup How to Integrate HashiCorp Vault with Rook-Ceph (Token Auth) How to Configure TLS for Vault Integration in Rook How to Integrate HashiCorp Vault with Rook-Ceph (Kubernetes Auth) How to Validate Ceph Cluster Configuration After Deployment How to Understand User Type and ID Notation (TYPE.ID) in Ceph How to Configure User Management in the Ceph Dashboard How to Use Rook-Ceph with Kubernetes Operators How to Use Rook-Ceph with Helm Chart Deployments How to Use the Swift API with Ceph RGW How to Use SQLite Databases Stored on Ceph How to Use s3cmd with Ceph RGW How to Use the S3 API with Ceph RGW How to Use Red Hat Ceph with RHEL Virtualization How to Use RBD with QEMU How to Use RBD with Nomad How to Use RBD with CloudStack How to Use RBD Snapshot Rollback How to Use rados bench for Object Storage Benchmarking How to Secure Rook-Ceph with Pod Security Admission How to Use pg-upmap for PG Mapping in Ceph How to Use Multipath Devices with Ceph OSDs How to Use MinIO Client (mc) with Ceph RGW How to Use fs swap for CephFS How to Use fio for Ceph Block Storage Benchmarking How to Use the CephFS Shell How to Use Ceph RGW for Media Asset Management How to Use Ceph RGW for Log Storage and Archival How to Use Ceph RGW for Data Lake Storage How to Use Ceph RGW for Backup Repository Storage How to Use the ceph-authtool Utility How to Use boto3 (Python) with Ceph RGW S3 How to Use AWS CLI with Ceph RGW S3 How to Use the Admin Ops API with Ceph RGW How to Configure Usage Log Key Transition in Ceph RGW How to Handle Rook-Ceph Upgrades in GitOps Pipelines How to Create a Ceph Upgrade Runbook How to Upgrade the Rook Operator from v1.18 to v1.19 How to Upgrade the Rook Operator on Kubernetes How to Upgrade External Cluster Connections in Rook How to Upgrade the Ceph Version in Rook How to Upgrade from Ceph Reef to Squid How to Upgrade from Ceph Quincy to Reef How to Upgrade Ceph Clusters in Stretch Mode How to Update Kernel for CephFS Feature Compatibility How to Update Ceph Configuration on a Running Rook Cluster How to Create Unique Kubernetes Services per NFS Server in Rook How to Understand When Compression Helps vs Hurts in Ceph How to Understand User Types (Individual vs System) in Ceph How to Understand the undersized PG State in Ceph How to Understand the stale PG State in Ceph How to Understand the repair PG State in Ceph How to Understand the remapped PG State in Ceph How to Understand Red Hat Ceph Storage vs Upstream Ceph How to Understand Placement Groups in Ceph How to Understand PG Splitting in Ceph How to Understand the peering PG State in Ceph How to Understand OSD Recovery Process in Ceph How to Understand the OSD Map in Ceph How to Understand New Features in Each Ceph Release How to Understand Monitor Leadership in Ceph How to Understand MDS States in CephFS How to Understand Deprecated Features in Ceph Reef How to Understand the degraded PG State in Ceph How to Understand D3N in Ceph How to Understand the creating PG State in Ceph How to Understand the clean PG State in Ceph How to Understand CephX Authentication Protocol How to Understand CephX Authentication Flow How to Understand What Data Ceph Telemetry Collects
How to Upgrade Rook-Ceph with Zero Downtime
Nawaz Dhandala · 2026-03-31 · via OneUptime Blog

Zero-downtime upgrades in Rook require upgrading the operator first, then allowing it to rolling-update Ceph daemons in the correct order. PVCs remain accessible throughout if the cluster is healthy before starting.

Upgrade Order

flowchart TD
    A[Pre-upgrade validation] --> B[Upgrade Rook operator]
    B --> C[Operator upgrades MON daemons]
    C --> D[Operator upgrades MGR daemons]
    D --> E[Operator upgrades OSD daemons]
    E --> F[Operator upgrades MDS daemons]
    F --> G[Operator upgrades RGW daemons]
    G --> H[Post-upgrade validation]
    H --> I[Cluster HEALTH_OK]

Pre-Upgrade Checklist

# 1. Confirm cluster is HEALTH_OK
kubectl exec -n rook-ceph deploy/rook-ceph-tools -- ceph status
kubectl exec -n rook-ceph deploy/rook-ceph-tools -- ceph health detail

# 2. Check all OSDs are up and in
kubectl exec -n rook-ceph deploy/rook-ceph-tools -- ceph osd stat
# "X osds: X up, X in" (no down or out)

# 3. Confirm no PGs degraded
kubectl exec -n rook-ceph deploy/rook-ceph-tools -- ceph pg stat
# active+clean for all PGs

# 4. Check current Rook version
kubectl get deployment rook-ceph-operator -n rook-ceph \
  -o jsonpath='{.spec.template.spec.containers[0].image}'

# 5. Check current Ceph version
kubectl exec -n rook-ceph deploy/rook-ceph-tools -- ceph version

Disable PG Auto-Scrub During Upgrade (Optional)

kubectl exec -n rook-ceph deploy/rook-ceph-tools -- \
  ceph osd set noscrub

kubectl exec -n rook-ceph deploy/rook-ceph-tools -- \
  ceph osd set nodeep-scrub

Upgrade the Rook Operator

Using Helm

# Check current chart version
helm list -n rook-ceph

# Update repo
helm repo update rook-release

# Check available versions
helm search repo rook-release/rook-ceph --versions | head -10

# Upgrade operator (example: v1.13.x to v1.14.x)
helm upgrade rook-ceph rook-release/rook-ceph \
  --namespace rook-ceph \
  --version v1.14.0 \
  --reuse-values

# Monitor operator pod
kubectl rollout status deployment rook-ceph-operator -n rook-ceph

Using kubectl (manifest)

# Download the new operator manifest
curl -O https://raw.githubusercontent.com/rook/rook/v1.14.0/deploy/examples/operator.yaml

# Review changes before applying
kubectl diff -f operator.yaml

# Apply the new operator
kubectl apply -f operator.yaml

# Watch operator rollout
kubectl rollout status deployment rook-ceph-operator -n rook-ceph

Upgrade the Ceph Version

Update the Ceph image in the CephCluster spec:

apiVersion: ceph.rook.io/v1
kind: CephCluster
metadata:
  name: rook-ceph
  namespace: rook-ceph
spec:
  cephVersion:
    image: quay.io/ceph/ceph:v18.2.4   # New Ceph Reef version
    allowUnsupported: false
kubectl apply -f cephcluster.yaml

# Watch daemon rolling update
kubectl get pods -n rook-ceph -w | grep -v Running

Monitor the Rolling Upgrade

# Watch MONs update first
kubectl get pods -n rook-ceph -l app=rook-ceph-mon -w

# Then MGR
kubectl get pods -n rook-ceph -l app=rook-ceph-mgr -w

# Then OSDs (one at a time)
kubectl get pods -n rook-ceph -l app=rook-ceph-osd -w

# Cluster should remain HEALTH_OK or HEALTH_WARN during OSD rolling updates
watch kubectl exec -n rook-ceph deploy/rook-ceph-tools -- ceph status

Verify Each Daemon After Upgrade

# Check MON versions
kubectl exec -n rook-ceph deploy/rook-ceph-tools -- ceph versions

# Should show:
# {
#   "mon": {"ceph version 18.2.4 ...": 3},
#   "mgr": {"ceph version 18.2.4 ...": 2},
#   "osd": {"ceph version 18.2.4 ...": 9},
#   "mds": {"ceph version 18.2.4 ...": 2}
# }

# Upgrade is complete when all daemons show the new version

Post-Upgrade Validation

# Re-enable scrub
kubectl exec -n rook-ceph deploy/rook-ceph-tools -- \
  ceph osd unset noscrub
kubectl exec -n rook-ceph deploy/rook-ceph-tools -- \
  ceph osd unset nodeep-scrub

# Final health check
kubectl exec -n rook-ceph deploy/rook-ceph-tools -- ceph status
kubectl exec -n rook-ceph deploy/rook-ceph-tools -- ceph health detail

# Verify PVCs still accessible
kubectl get pvc --all-namespaces | grep Bound

Summary

Zero-downtime Rook-Ceph upgrades require a healthy cluster as the starting point, followed by upgrading the Rook operator, then updating the Ceph version image in the CephCluster CR. The operator handles the rolling update of all daemons in the correct order (MON, MGR, OSD, MDS, RGW). Monitor ceph status throughout and disable scrub during the OSD rolling update to reduce overhead. Always validate cluster health before and after each stage.