惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

MyScale Blog
MyScale Blog
G
Google Developers Blog
B
Blog
Microsoft Azure Blog
Microsoft Azure Blog
博客园_首页
人人都是产品经理
人人都是产品经理
B
Blog RSS Feed
A
About on SuperTechFans
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
aimingoo的专栏
aimingoo的专栏
N
News and Events Feed by Topic
L
LINUX DO - 最新话题
V
Vulnerabilities – Threatpost
H
Hacker News: Front Page
T
Tor Project blog
P
Proofpoint News Feed
P
Privacy International News Feed
Recorded Future
Recorded Future
F
Fortinet All Blogs
量子位
博客园 - 聂微东
月光博客
月光博客
博客园 - Franky
SecWiki News
SecWiki News
G
GRAHAM CLULEY
腾讯CDC
Know Your Adversary
Know Your Adversary
宝玉的分享
宝玉的分享
The Cloudflare Blog
美团技术团队
小众软件
小众软件
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
T
Threatpost
爱范儿
爱范儿
A
Arctic Wolf
博客园 - 叶小钗
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Project Zero
Project Zero
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
J
Java Code Geeks
C
Cyber Attacks, Cyber Crime and Cyber Security
博客园 - 三生石上(FineUI控件)
PCI Perspectives
PCI Perspectives
Latest news
Latest news
V
V2EX
罗磊的独立博客
T
Threat Research - Cisco Blogs
Scott Helme
Scott Helme
S
Security Affairs
S
SegmentFault 最新的问题

OneUptime Blog

How to Monitor Azure App Services (PaaS) with OpenTelemetry Grafana Stack vs OneUptime: DIY Observability or Unified Platform? Your AI Workloads Are About to Blow Up Your Observability Bill The Great Observability Consolidation Is Here How to Write Custom Object Classes for Ceph How to Write Custom Ceph Manager Modules How to Write a ceph.conf Configuration File How to Use Rook-Ceph with OpenShift How to Use Rook-Ceph with Longhorn for Comparison How to Configure Volume Snapshot Class for RBD in Rook How to Configure VolumeReplicationClass Scheduling Intervals in Rook How to Set Up Volume Replication with Rook-Ceph How to Create Volume Group Snapshots with Rook CSI How to Visualize Ceph Network Performance in Grafana How to Enable Virtual Host-Style Bucket Access in Rook How to View Runtime Configuration via Admin Socket How to View Quota Settings and Update Stats in Ceph RGW How to View PG Scaling Recommendations with autoscale-status How to View PG Distribution via Admin Socket How to View Performance Metrics in the Ceph Dashboard How to View OSD Performance Counters in Ceph How to View Connection Status via Admin Socket How to View Ceph Cluster Summary Dashboard via CLI How to Version Control Rook-Ceph Configuration How to Version Control Ceph Infrastructure with Terraform How to Verify Kubernetes Node Requirements for Rook-Ceph Deployment How to Verify Health Before and After Rook Upgrades How to Verify Data Integrity with Deep Scrubbing How to Verify Complete Rook-Ceph Cleanup How to Verify Backup Integrity from Ceph Snapshots How to Use Rook-Ceph with Velero for Kubernetes Backup How to Integrate HashiCorp Vault with Rook-Ceph (Token Auth) How to Configure TLS for Vault Integration in Rook How to Integrate HashiCorp Vault with Rook-Ceph (Kubernetes Auth) How to Validate Ceph Cluster Configuration After Deployment How to Understand User Type and ID Notation (TYPE.ID) in Ceph How to Configure User Management in the Ceph Dashboard How to Use Rook-Ceph with Kubernetes Operators How to Use Rook-Ceph with Helm Chart Deployments How to Use the Swift API with Ceph RGW How to Use SQLite Databases Stored on Ceph How to Use s3cmd with Ceph RGW How to Use the S3 API with Ceph RGW How to Use Red Hat Ceph with RHEL Virtualization How to Use RBD with QEMU How to Use RBD with Nomad How to Use RBD with CloudStack How to Use RBD Snapshot Rollback How to Use rados bench for Object Storage Benchmarking How to Secure Rook-Ceph with Pod Security Admission How to Use pg-upmap for PG Mapping in Ceph How to Use Multipath Devices with Ceph OSDs How to Use MinIO Client (mc) with Ceph RGW How to Use fs swap for CephFS How to Use fio for Ceph Block Storage Benchmarking How to Use the CephFS Shell How to Use Ceph RGW for Media Asset Management How to Use Ceph RGW for Log Storage and Archival How to Use Ceph RGW for Data Lake Storage How to Use Ceph RGW for Backup Repository Storage How to Use the ceph-authtool Utility How to Use boto3 (Python) with Ceph RGW S3 How to Use AWS CLI with Ceph RGW S3 How to Use the Admin Ops API with Ceph RGW How to Configure Usage Log Key Transition in Ceph RGW How to Handle Rook-Ceph Upgrades in GitOps Pipelines How to Upgrade Rook-Ceph with Zero Downtime How to Create a Ceph Upgrade Runbook How to Upgrade the Rook Operator from v1.18 to v1.19 How to Upgrade the Rook Operator on Kubernetes How to Upgrade External Cluster Connections in Rook How to Upgrade the Ceph Version in Rook How to Upgrade from Ceph Reef to Squid How to Upgrade from Ceph Quincy to Reef How to Upgrade Ceph Clusters in Stretch Mode How to Update Kernel for CephFS Feature Compatibility How to Update Ceph Configuration on a Running Rook Cluster How to Create Unique Kubernetes Services per NFS Server in Rook How to Understand When Compression Helps vs Hurts in Ceph How to Understand User Types (Individual vs System) in Ceph How to Understand the undersized PG State in Ceph How to Understand the stale PG State in Ceph How to Understand the repair PG State in Ceph How to Understand the remapped PG State in Ceph How to Understand Red Hat Ceph Storage vs Upstream Ceph How to Understand Placement Groups in Ceph How to Understand PG Splitting in Ceph How to Understand OSD Recovery Process in Ceph How to Understand the OSD Map in Ceph How to Understand New Features in Each Ceph Release How to Understand Monitor Leadership in Ceph How to Understand MDS States in CephFS How to Understand Deprecated Features in Ceph Reef How to Understand the degraded PG State in Ceph How to Understand D3N in Ceph How to Understand the creating PG State in Ceph How to Understand the clean PG State in Ceph How to Understand CephX Authentication Protocol How to Understand CephX Authentication Flow How to Understand What Data Ceph Telemetry Collects
How to Understand the peering PG State in Ceph
Nawaz Dhandala · 2026-03-31 · via OneUptime Blog

Peering is the process by which all OSDs in a PG's acting set agree on the complete and consistent state of every object in the PG. A PG must complete peering before it can serve any I/O. Understanding peering is fundamental to diagnosing cluster startup and OSD failure recovery.

What Peering Is

When a PG starts or any OSD in its acting set changes, the PG must peer. The primary OSD:

  1. Contacts all replica OSDs and requests their PG logs
  2. Determines the authoritative object version history
  3. Identifies any objects that need recovery or removal
  4. Marks the PG as active when all replicas agree

Until peering completes, the PG is peering and cannot serve I/O.

Checking Peering PGs

ceph status
# HEALTH_WARN: X pgs not active

ceph pg stat | grep peering

# List PGs currently in peering
ceph pg dump | grep "peering"

How Long Peering Takes

Peering should complete in seconds. If it is taking minutes, something is wrong. Check:

# PGs that have been peering for a long time
ceph health detail | grep -i "stuck"

# Per-PG peering state
ceph pg <pg-id> query | jq '.recovery_state.name'

Why PGs Get Stuck in Peering

Not enough OSDs up

For a size 3 pool with min_size 2, at least 2 OSDs must be up:

# Check acting set for stuck PG
ceph pg <pg-id> query | jq '{acting: .acting, up: .up}'

# Check which OSDs are down
ceph osd tree | grep down

pg_temp map issues

Old pg_temp entries can confuse peering:

ceph osd dump | grep pg_temp
ceph osd pg-temp <pg-id>  # clear pg_temp by passing no OSDs

Stuck in WaitUpThru

The PG may wait for the monitor's epoch to advance:

ceph pg <pg-id> query | jq '.recovery_state'

If stuck in WaitUpThru, the OSD is waiting for the monitor to acknowledge its up_thru value. Restarting the affected OSD or triggering an OSD map update can help:

# Check current OSD map epoch
ceph osd stat

# Restart the affected OSD to trigger a new map update
systemctl restart ceph-osd@<osd-id>

Viewing Peering Log

ceph pg <pg-id> query | jq '.recovery_state.prior_set'
ceph pg <pg-id> query | jq '.peer_info'

Forcing Peering

If a PG is stuck peering because not enough OSDs are available, you can lower the pool's min_size to allow the PG to become active with fewer replicas:

ceph osd pool set <pool-name> min_size 1

This allows PGs to become active with only one copy and may serve stale data. Restore the original min_size once the missing OSDs are recovered. Use only when you need to restore I/O and accept the risk of reduced redundancy.

Peering After Cluster Restart

After a full cluster restart, all PGs start peering simultaneously. This is normal and usually resolves quickly:

watch ceph pg stat
# peering count decreases as OSDs complete startup

Summary

Peering is the negotiation phase where all acting OSDs for a PG agree on its authoritative state. It is a brief transitional state that precedes active. PGs stuck in peering indicate a problem with OSD availability or the peering protocol. Resolve by ensuring all acting OSDs are reachable or by adjusting pool min_size if permanent OSD loss has occurred.