Ceph RGW S3 Object Gateway: Realm, Zonegroup and Zone - 夜莺博客

Ceph RGW S3 Object Gateway: Realm, Zonegroup and Zone

The Ceph Object Gateway (radosgw) turns a RADOS cluster into an S3- and Swift-compatible object store. The part that surprises people is the configuration hierarchy: objects live in buckets, buckets live in zones, zones live in zonegroups, and zonegroups live in realms. Getting that stack right from the beginning is what makes multisite replication possible later without a rebuild. This article walks the four levels in order, gives the radosgw-admin commands for each, and ends with the two checks that prove the gateway is serving S3 correctly.

The Four Levels

  • Realm — a globally unique namespace. It is the container that makes multisite possible and it enforces namespace uniqueness across all its zonegroups.
  • Zonegroup — a geographic or administrative grouping of zones. It holds the endpoints advertised to clients. Formerly called a "region".
  • Zone — a set of radosgw instances backed by the same pools. All gateways in a zone serve the same objects from the same pools.
  • Period — the time-versioned state of the zonegroup/zone configuration. Every configuration change should be followed by a period update and commit.

A multi-site configuration requires one master zonegroup and, inside it, one master zone. Other zonegroups may exist but each needs its own master zone.

Step 1 — Create the Realm

radosgw-admin realm create --rgw-realm=prod --default

Specify --default if this is the only realm on the cluster. Without it you must pass --rgw-realm or --realm-id on every subsequent zonegroup and zone command, which is easy to forget.

Step 2 — Create the Master Zonegroup

radosgw-admin zonegroup create --rgw-zonegroup=shared \
    --endpoints=http://rgw.example.com:80 \
    --rgw-realm=prod --master --default

Step 3 — Create the Master Zone

radosgw-admin zone create --rgw-zonegroup=shared \
    --rgw-zone=us-east --master --default \
    --endpoints=http://rgw.example.com:80

radosgw-admin period update --commit

The period commit is not optional. Skipping it is the single most common reason a freshly created zone does not appear to exist when you query it back.

Step 4 — Add a Secondary Zone (Multisite)

# On a host in the secondary zone - note: no --master, no --default
radosgw-admin zone create --rgw-zonegroup=shared \
    --rgw-zone=us-west \
    --access-key= --secret= \
    --endpoints=http://rgw-west.example.com:80
radosgw-admin period update --commit

Zones run active-active by default, so a client may write to either zone and the zone replicates to its peers. Add --read-only to make the secondary zone passive if you want active-passive instead.

Step 5 — Point the Gateway at the Zone

# ceph.conf
[client.rgw.us-east]
rgw_realm = prod
rgw_zonegroup = shared
rgw_zone = us-east
rgw_enable_apis = s3,s3website

Serving with the s3 API enabled is mandatory for any radosgw instance that will participate in multisite. For S3-style virtual-host bucket access, add a wildcard to the DNS record the gateway resolves against and set the zonegroup hostnames accordingly.

Verification

radosgw-admin realm list
radosgw-admin zonegroup list
radosgw-admin zone list
radosgw-admin sync status
radosgw-admin period get

# Prove S3 actually works end to end
aws --endpoint-url http://rgw.example.com s3 mb s3://test-bucket
aws --endpoint-url http://rgw.example.com s3 cp ./file.bin s3://test-bucket/
aws --endpoint-url http://rgw.example.com s3 ls s3://test-bucket/

radosgw-admin sync status is the command that tells you whether replication is healthy: it prints the metadata sync state and, per data pool, whether the shards are caught up or lagging. If buckets replicate but objects do not, the usual cause is a data-sync policy mismatch or a bucket that was created before the zone was added to the zonegroup.

Design Notes

  • Zone placement uses a placement_target with storage classes. If you need different pools for different workloads, define a second placement target rather than mixing them.
  • Bucket index sharding is set per zone. The default is adequate for small deployments but too low for millions of objects per bucket — raise it before loading data, not after.
  • Gateways in a multi-site configuration retrieve their configuration from a radosgw in the master zone. If that zone is unreachable during a cold start, a secondary gateway may come up with stale configuration.

Related reading on this site: Ceph CRUSH Map: Device Classes and Custom Rules for the placement rules behind the pools, Ceph OSD Recovery and Backfill Throttling recovery tuning to keep backfill from starving gateway traffic, and MinIO Distributed Erasure-Coded Cluster Deployment if you want an S3-compatible gateway that is not Ceph-based.

原文链接:https://docs.ceph.com/en/reef/radosgw/multisite/