惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
CERT Recently Published Vulnerability Notes
G
Google Developers Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
宝玉的分享
宝玉的分享
Microsoft Security Blog
Microsoft Security Blog
Jina AI
Jina AI
L
LangChain Blog
博客园_首页
有赞技术团队
有赞技术团队
The Register - Security
The Register - Security
GbyAI
GbyAI
Blog — PlanetScale
Blog — PlanetScale
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
J
Java Code Geeks
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Security Archives - TechRepublic
Security Archives - TechRepublic
量子位
雷峰网
雷峰网
Security Latest
Security Latest
博客园 - 【当耐特】
V2EX - 技术
V2EX - 技术
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - 聂微东
IT之家
IT之家
爱范儿
爱范儿
S
Schneier on Security
N
News | PayPal Newsroom
H
Help Net Security
Recent Announcements
Recent Announcements
Martin Fowler
Martin Fowler
N
News and Events Feed by Topic
C
Cyber Attacks, Cyber Crime and Cyber Security
U
Unit 42
博客园 - 司徒正美
Forbes - Security
Forbes - Security
P
Proofpoint News Feed
W
WeLiveSecurity
Cisco Talos Blog
Cisco Talos Blog
小众软件
小众软件
The Cloudflare Blog
AWS News Blog
AWS News Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
P
Palo Alto Networks Blog
Google DeepMind News
Google DeepMind News
H
Heimdal Security Blog
V
Vulnerabilities – Threatpost
Microsoft Azure Blog
Microsoft Azure Blog
T
Tailwind CSS Blog
G
GRAHAM CLULEY

OneUptime Blog

How to Monitor Azure App Services (PaaS) with OpenTelemetry Grafana Stack vs OneUptime: DIY Observability or Unified Platform? Your AI Workloads Are About to Blow Up Your Observability Bill The Great Observability Consolidation Is Here How to Write Custom Object Classes for Ceph How to Write Custom Ceph Manager Modules How to Write a ceph.conf Configuration File How to Use Rook-Ceph with OpenShift How to Use Rook-Ceph with Longhorn for Comparison How to Configure Volume Snapshot Class for RBD in Rook How to Configure VolumeReplicationClass Scheduling Intervals in Rook How to Set Up Volume Replication with Rook-Ceph How to Create Volume Group Snapshots with Rook CSI How to Visualize Ceph Network Performance in Grafana How to Enable Virtual Host-Style Bucket Access in Rook How to View Runtime Configuration via Admin Socket How to View Quota Settings and Update Stats in Ceph RGW How to View PG Scaling Recommendations with autoscale-status How to View PG Distribution via Admin Socket How to View Performance Metrics in the Ceph Dashboard How to View OSD Performance Counters in Ceph How to View Connection Status via Admin Socket How to View Ceph Cluster Summary Dashboard via CLI How to Version Control Rook-Ceph Configuration How to Version Control Ceph Infrastructure with Terraform How to Verify Kubernetes Node Requirements for Rook-Ceph Deployment How to Verify Health Before and After Rook Upgrades How to Verify Data Integrity with Deep Scrubbing How to Verify Complete Rook-Ceph Cleanup How to Verify Backup Integrity from Ceph Snapshots How to Use Rook-Ceph with Velero for Kubernetes Backup How to Integrate HashiCorp Vault with Rook-Ceph (Token Auth) How to Configure TLS for Vault Integration in Rook How to Integrate HashiCorp Vault with Rook-Ceph (Kubernetes Auth) How to Validate Ceph Cluster Configuration After Deployment How to Understand User Type and ID Notation (TYPE.ID) in Ceph How to Configure User Management in the Ceph Dashboard How to Use Rook-Ceph with Kubernetes Operators How to Use Rook-Ceph with Helm Chart Deployments How to Use the Swift API with Ceph RGW How to Use SQLite Databases Stored on Ceph How to Use s3cmd with Ceph RGW How to Use the S3 API with Ceph RGW How to Use Red Hat Ceph with RHEL Virtualization How to Use RBD with QEMU How to Use RBD with Nomad How to Use RBD with CloudStack How to Use RBD Snapshot Rollback How to Use rados bench for Object Storage Benchmarking How to Secure Rook-Ceph with Pod Security Admission How to Use pg-upmap for PG Mapping in Ceph How to Use Multipath Devices with Ceph OSDs How to Use MinIO Client (mc) with Ceph RGW How to Use fs swap for CephFS How to Use fio for Ceph Block Storage Benchmarking How to Use the CephFS Shell How to Use Ceph RGW for Media Asset Management How to Use Ceph RGW for Log Storage and Archival How to Use Ceph RGW for Backup Repository Storage How to Use the ceph-authtool Utility How to Use boto3 (Python) with Ceph RGW S3 How to Use AWS CLI with Ceph RGW S3 How to Use the Admin Ops API with Ceph RGW How to Configure Usage Log Key Transition in Ceph RGW How to Handle Rook-Ceph Upgrades in GitOps Pipelines How to Upgrade Rook-Ceph with Zero Downtime How to Create a Ceph Upgrade Runbook How to Upgrade the Rook Operator from v1.18 to v1.19 How to Upgrade the Rook Operator on Kubernetes How to Upgrade External Cluster Connections in Rook How to Upgrade the Ceph Version in Rook How to Upgrade from Ceph Reef to Squid How to Upgrade from Ceph Quincy to Reef How to Upgrade Ceph Clusters in Stretch Mode How to Update Kernel for CephFS Feature Compatibility How to Update Ceph Configuration on a Running Rook Cluster How to Create Unique Kubernetes Services per NFS Server in Rook How to Understand When Compression Helps vs Hurts in Ceph How to Understand User Types (Individual vs System) in Ceph How to Understand the undersized PG State in Ceph How to Understand the stale PG State in Ceph How to Understand the repair PG State in Ceph How to Understand the remapped PG State in Ceph How to Understand Red Hat Ceph Storage vs Upstream Ceph How to Understand Placement Groups in Ceph How to Understand PG Splitting in Ceph How to Understand the peering PG State in Ceph How to Understand OSD Recovery Process in Ceph How to Understand the OSD Map in Ceph How to Understand New Features in Each Ceph Release How to Understand Monitor Leadership in Ceph How to Understand MDS States in CephFS How to Understand Deprecated Features in Ceph Reef How to Understand the degraded PG State in Ceph How to Understand D3N in Ceph How to Understand the creating PG State in Ceph How to Understand the clean PG State in Ceph How to Understand CephX Authentication Protocol How to Understand CephX Authentication Flow How to Understand What Data Ceph Telemetry Collects
How to Use Ceph RGW for Data Lake Storage
Nawaz Dhandala · 2026-03-31 · via OneUptime Blog

Why Ceph RGW for Data Lake Storage?

Data lakes require scalable, cost-effective object storage with S3-compatible APIs. Ceph RGW provides:

  • S3 and Swift compatible API
  • Horizontal scalability to petabytes
  • Multi-tenancy via bucket policies
  • On-premises data sovereignty
  • Integration with major analytics frameworks

Setting Up a Data Lake Bucket

# Configure the AWS CLI to point to Ceph RGW
aws configure set default.endpoint_url https://rgw.example.com
export AWS_ACCESS_KEY_ID=<access-key>
export AWS_SECRET_ACCESS_KEY=<secret-key>

# Create a data lake bucket
aws s3 mb s3://datalake --endpoint-url https://rgw.example.com

# Create a structured folder hierarchy
aws s3api put-object --bucket datalake --key raw/ --endpoint-url https://rgw.example.com
aws s3api put-object --bucket datalake --key processed/ --endpoint-url https://rgw.example.com
aws s3api put-object --bucket datalake --key curated/ --endpoint-url https://rgw.example.com

Configuring Bucket Versioning

Enable versioning to maintain history of data changes:

aws s3api put-bucket-versioning \
  --bucket datalake \
  --versioning-configuration Status=Enabled \
  --endpoint-url https://rgw.example.com

Setting a Lifecycle Policy for Data Tiers

Move data from raw to archive after 90 days:

{
  "Rules": [{
    "ID": "archive-raw-data",
    "Filter": { "Prefix": "raw/" },
    "Status": "Enabled",
    "Transitions": [{
      "Days": 90,
      "StorageClass": "GLACIER"
    }]
  }]
}
aws s3api put-bucket-lifecycle-configuration \
  --bucket datalake \
  --lifecycle-configuration file://lifecycle.json \
  --endpoint-url https://rgw.example.com

Integrating with Apache Spark

Configure Spark to read from Ceph RGW:

from pyspark.sql import SparkSession

spark = SparkSession.builder \
    .appName("DataLake") \
    .config("spark.hadoop.fs.s3a.endpoint", "https://rgw.example.com") \
    .config("spark.hadoop.fs.s3a.access.key", "my-access-key") \
    .config("spark.hadoop.fs.s3a.secret.key", "my-secret-key") \
    .config("spark.hadoop.fs.s3a.path.style.access", "true") \
    .config("spark.hadoop.fs.s3a.impl", "org.apache.hadoop.fs.s3a.S3AFileSystem") \
    .getOrCreate()

# Read Parquet from data lake
df = spark.read.parquet("s3a://datalake/processed/events/")
df.show()

Integrating with Trino (Presto)

Configure Trino catalog for Ceph S3:

connector.name=hive
hive.metastore.uri=thrift://hive-metastore:9083
hive.s3.endpoint=https://rgw.example.com
hive.s3.aws-access-key=my-access-key
hive.s3.aws-secret-key=my-secret-key
hive.s3.path-style-access=true
hive.s3.ssl.enabled=true

Query data lake tables:

SELECT date_trunc('hour', event_time) AS hour,
       count(*) AS events
FROM datalake.processed.events
WHERE event_date = CURRENT_DATE
GROUP BY 1
ORDER BY 1;

Enabling Multipart Upload for Large Files

Large data files (>100MB) should use multipart uploads. Configure the chunk size first, then upload:

# Set multipart chunk size to 64MB
aws configure set default.s3.multipart_chunksize 64MB

# Upload large file (multipart upload is used automatically)
aws s3 cp large-dataset.parquet s3://datalake/raw/datasets/ \
  --endpoint-url https://rgw.example.com

Summary

Ceph RGW provides a scalable, self-hosted S3-compatible backend for data lake architectures. Configure structured bucket hierarchies with raw, processed, and curated zones, enable versioning for data lineage, and set lifecycle policies to automate data tiering. Integrate with Spark using the S3A connector and Trino using the Hive catalog, both of which support path-style access required by Ceph RGW.