PostgreSQL on Kubernetes · Design, migration, operations and 24×7 support · Any operator, any cluster, any cloud

PostgreSQL on Kubernetes, Engineered to Production Standards: Operators, Storage, Failover and Backups You Can Drill

MinervaDB designs, migrates and operates PostgreSQL on Kubernetes for enterprises whose platform teams already run Kubernetes in production, and tells the others honestly when a managed service is the better fit. The work covers operator selection (CloudNativePG, Percona, Crunchy PGO, Zalando, StackGres), storage and scheduling engineering, failover and restore drills that produce measured RPO and RTO, and a 24×7 senior watch with S1 acknowledged in 15 minutes. Every change is declared in Git, reviewed, reversible and rehearsed before it reaches a production cluster.

5operators supported: CloudNativePG, Percona, Crunchy, Zalando, StackGres
900+enterprises supported across every major engine
46cities with on-site delivery presence
15 minS1 acknowledgement, 24×7×365
200+years of combined leadership experience

01 · Why PostgreSQL on Kubernetes needs database engineering

The operator automates the mechanics; the outcomes still have to be engineered

Kubernetes is the default control plane for application platforms, and running PostgreSQL on Kubernetes lets one team provision, upgrade and observe databases the same way it does everything else. The operator does not choose your storage class, size your memory, or time your restore.

A mature operator turns a Cluster manifest into pods, persistent volumes, Services, replication and backups, and it will promote a replica when the primary disappears. What it cannot do is decide whether the persistent volume can follow the pod to another zone, whether the memory limit leaves room for the page cache, whether a node drain will take the primary and its synchronous replica down together, or whether the backup in object storage restores in ten minutes or four hours.

Those are database engineering decisions, and they are where PostgreSQL on Kubernetes projects succeed or stall. MinervaDB brings two decades of PostgreSQL production experience to the manifests, the storage layer and the runbooks, so the cluster behaves the way the diagram promises when a node actually dies.

The practice is vendor-neutral by principle. We sell no operator, no distribution and no cloud, so the recommendation can be CloudNativePG, Percona's operator, a commercial vendor, or an honest "use RDS for this one". Every recommendation is justified by a metric, a drill result or a licence term you can check yourself.

Whatever the answer, the same senior engineers who design the platform carry the 24×7 watch afterwards, with severity targets written into the agreement rather than into a slide.

02 · PostgreSQL on Kubernetes reference architecture

What a production PostgreSQL on Kubernetes cluster looks like when it is engineered

One namespace per cluster, an operator reconciling a declarative Cluster resource, a primary and replicas spread across zones on their own persistent volumes, poolers in front, WAL flowing to object storage, and observability wired from day one.

PostgreSQL on Kubernetes reference architecture: operator, rw and ro Services, primary and replicas across zones on CSI persistent volumes, PgBouncer, WAL archiving to object storage, TLS and Prometheus

Figure 1. PostgreSQL on Kubernetes reference architecture: operator control plane, rw and ro Services, primary and replicas across three zones on CSI-backed persistent volumes, PgBouncer poolers, WAL archiving to object storage, secrets, TLS and Prometheus, with the pieces that live outside the cluster.

Declarative and reviewable

PostgreSQL on Kubernetes is declarative by nature: the Cluster resource, pooler, backup schedule and scheduled restores live in Git and are applied by Argo CD or Flux. A change to shared_buffers, a replica count or a storage class is a pull request with a diff, a review and a rollback, not a command typed on a node.

Topology that survives a zone

Topology spread constraints and anti-affinity place the primary and each replica in different zones, a storage class per zone with WaitForFirstConsumer keeps the volume where the pod lands, and a PodDisruptionBudget stops a routine node drain from taking two members at once.

Services, not IPs

Applications connect to the rw Service for writes and the ro Service for replica reads, through PgBouncer poolers with per-role pool sizes and TLS. A failover moves endpoints; the application only needs a driver that reconnects and retries idempotently.

03 · PostgreSQL on Kubernetes operator selection

Choosing the right PostgreSQL on Kubernetes operator, by evidence

No operator is best for every organisation. The choice is weighted on how it behaves in a failover drill, how long its restore takes on your data, its upgrade path across PostgreSQL majors, and the licence terms of its images as well as its code.

PostgreSQL on Kubernetes operator selection matrix: CloudNativePG, Percona, Crunchy PGO, Zalando and StackGres compared on HA, backup, pooling, licence and best fit

Figure 2. Operator selection matrix for PostgreSQL on Kubernetes: HA mechanism, backup and PITR tooling, pooling, licence and stewardship, and best fit, with the weighting MinervaDB applies. Verified 28 September 2026.

OperatorHow it does HAWhat we check before recommending it
CloudNativePG 1.29Operator-native: an instance manager in each pod, primary election through the Kubernetes API, no Patroni or external DCSBehaviour when the operator itself is down, volume-snapshot backup support on your CSI driver, barman-cloud plugin retention, major-version upgrade procedure, supported Kubernetes versions
Percona Operator for PostgreSQL 2.xPatroni inside the pods with Kubernetes as the DCS; pgBackRest for backups and PITR; PgBouncer built inFit with an estate already on Percona tooling for MySQL or MongoDB, pgBackRest repository layout, pgvector and extension images
Crunchy Postgres for Kubernetes (PGO) 5.xPatroni inside the pods; pgBackRest with multiple repositories; PgBouncer built inContainer image licence terms for production use versus the Apache-licensed operator code, commercial support scope, monitoring stack integration
Zalando postgres-operatorPatroni in the Spilo image; WAL-G or WAL-E to object storage; connection poolerCommunity cadence and PostgreSQL major support in Spilo, upgrade path, fit for teams already standardised on Patroni outside Kubernetes
StackGresPatroni; pgBackRest through SGBackup; PgBouncer and Envoy; management UIAGPL v3 obligations, commercial edition boundaries, UI versus GitOps workflow for the platform team

PostgreSQL on Kubernetes also runs on any conformant distribution: Amazon EKS, Azure AKS, Google GKE, Red Hat OpenShift, Rancher, upstream Kubernetes on-premises, and air-gapped clusters. The operator choice is independent of the distribution; the storage class and node pool design are not.

04 · PostgreSQL on Kubernetes failover and self-healing

What happens when a node dies, second by second

High availability for PostgreSQL on Kubernetes is a sequence of timeouts and endpoint changes, each of which can be measured. We measure them in a drill before production does it for us.

PostgreSQL on Kubernetes failover sequence: node loss, detection, fencing, promotion, Service endpoint switch, rejoin, with the RTO components measured in a drill

Figure 3. Failover sequence for PostgreSQL on Kubernetes, the four components of RTO that a drill measures, and the scheduling and storage mistakes that turn a ten-second failover into a ten-minute outage.

Detection tuned, not defaulted

For PostgreSQL on Kubernetes, readiness-probe periods, failure thresholds and lease TTLs are set against the false-positive rate your nodes actually produce under memory pressure. Too aggressive and a busy checkpoint triggers a failover; too lax and RTO is measured in minutes.

Volumes that follow the pod

A persistent volume still attached to a dead node is the most common cause of a stalled recovery. Zonal storage classes, CSI attach and detach timeouts and, where the platform allows, volume-attachment force-detach are configured and then tested by killing the node, not by reading the docs.

Rejoin without a full re-clone

The old primary returns as a replica through pg_rewind where the timeline allows, or a fresh clone from the backup repository where it does not. Both paths are in the runbook with the expected duration on your data size.

05 · Backup, WAL archiving and PITR

PostgreSQL on Kubernetes backups that have been restored, on a schedule

Every operator can ship base backups and WAL to object storage. Whether the archive is complete, encrypted, retained for the compliance window and restorable inside the RTO is engineering, and it is proved by a timed restore into a new namespace.

PostgreSQL on Kubernetes backup, WAL archiving and point-in-time recovery to object storage with a timed restore drill

Figure 4. Backup, WAL archiving and point-in-time recovery for PostgreSQL on Kubernetes: base backups and continuous WAL to object storage, restore into a new Cluster resource with a recovery target, and the quarterly drill that produces the RTO the runbook quotes.

# CloudNativePG 1.29: a recovery cluster bootstrapped from
# the object-store backup of "orders-db" to a point in time
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: orders-db-restore
spec:
  instances: 1
  imageName: ghcr.io/cloudnative-pg/postgresql:18.6
  storage:
    storageClass: gp3-zone-a
    size: 500Gi
  bootstrap:
    recovery:
      source: orders-db
      recoveryTarget:
        targetTime: "2026-09-28 06:10:00+00"
  externalClusters:
    - name: orders-db
      plugin:
        name: barman-cloud.cloudnative-pg.io
        parameters:
          barmanObjectName: minervadb-backups
          serverName: orders-db

The PostgreSQL on Kubernetes restore drill loads production-size data into a new namespace, replays WAL to a timestamp, runs pg_amcheck, compares row counts and checksums against the source and smoke-tests the application through the ro Service. The elapsed time is written into the runbook as the RTO; the WAL archive lag at any moment is the RPO, and it is alerted at a fraction of the budget.

Retention, encryption keys and object-lock immutability are set against the compliance window rather than the tool's default, and the cross-region copy is tested by restoring in a second cluster. The failure cases rehearsed include namespace deletion, persistent-volume loss, an operator upgrade gone wrong and region loss.

06 · PostgreSQL on Kubernetes resources, storage and scheduling

Sizing PostgreSQL on Kubernetes so the kernel never kills the database

An OOM-killed primary is a design failure. Memory, CPU, storage and scheduling are engineered from measured workload behaviour, and the evidence that proves the sizing is collected from PostgreSQL, the kubelet and the storage layer together.

PostgreSQL on Kubernetes resource, storage and scheduling engineering: Guaranteed QoS, CPU pinning, zonal storage classes, node pools and disruption budgets with the evidence used

Figure 5. Resource, storage and scheduling engineering for PostgreSQL on Kubernetes: Guaranteed QoS, shared_buffers versus the memory limit, CPU pinning without throttling, zonal storage classes and provisioned IOPS, node pools and disruption budgets, with the evidence used to size and prove each decision.

Memory

PostgreSQL on Kubernetes memory: requests equal limits for Guaranteed QoS. shared_buffers at roughly a quarter of the limit, with headroom for work_mem times the pool size, maintenance work and the page cache that lives inside the same cgroup. Huge pages where the node pool allows them.

CPU

Whole-core requests and a static CPU manager policy for latency-sensitive primaries. CPU limits are avoided or set generously, because CFS throttling during a checkpoint or an anti-wraparound vacuum is exactly the wrong moment to slow PostgreSQL down.

Storage

A storage class per zone with WaitForFirstConsumer, IOPS and throughput provisioned to the measured write rate, pg_test_fsync run against the class before go-live, and a separate WAL volume or local NVMe where the operator and the workload justify it.

Scheduling

A dedicated node pool with taints and tolerations, topology spread across zones, a PodDisruptionBudget that keeps a quorum through drains and upgrades, and a priority class above application pods so the database is never the first thing evicted.

07 · Services

PostgreSQL on Kubernetes services across the full lifecycle

Design and migrate, optimise and audit, or hand the clusters to a 24×7 senior watch. Each service is delivered by the same engineers and hands back versioned manifests and runbooks.

Design and migrate

Operator and distribution selection, node pool and storage design, the Cluster resources and GitOps pipeline, and migration from VMs, RDS or another cluster by logical replication with parity checks and a rehearsed cutover through the rw Service, reverse replication held as the rollback. Data modernization →

Optimise and audit

Already running PostgreSQL on Kubernetes? A read-only audit of manifests, storage, scheduling, failover behaviour, backup completeness and PostgreSQL configuration, returned as findings ranked P0 to P2 with the metric each fix moves and the drill that proves it. PostgreSQL consulting →

24×7 support and operations

Senior engineers on watch with S1 acknowledged in 15 minutes, S2 in 12 hours, S3 in 24 hours and S4 in 48 hours; operator and PostgreSQL upgrades on a calendar; quarterly restore and failover drills; a monthly SLO report. 24×7 consultative support →

Security and compliance

TLS everywhere with cert-manager or operator-issued certificates, credentials from Secrets or Vault with rotation, network policies scoped to the namespace, pgaudit retained for the compliance window, and the evidence pack for GDPR, HIPAA, PCI DSS and SOC 2.

Observability

PostgreSQL on Kubernetes observability: operator exporters and custom queries into Prometheus, Grafana dashboards for replication lag, WAL archive status, persistent-volume usage, cgroup throttling and vacuum progress, and alerts tuned to the workload rather than a vendor default.

Upgrades

PostgreSQL on Kubernetes upgrades: operator upgrades rehearsed on a staging cluster first; PostgreSQL major-version upgrades by logical replication into a new cluster, or the operator's in-place path where it exists, with the restore timed before either starts. Current targets: PostgreSQL 18.6, or 17 where an extension lags.

08 · PostgreSQL on Kubernetes engagement lifecycle

Assess, design and pilot, migrate, operate

Every PostgreSQL on Kubernetes engagement follows the same four phases. The pilot is where the drills run for the first time, so nothing is discovered in production.

PostgreSQL on Kubernetes engagement lifecycle: assess, design and pilot, migrate, operate, and the Kubernetes versus managed DBaaS decision gate

Figure 6. The MinervaDB PostgreSQL on Kubernetes engagement lifecycle and the decision gate that decides between Kubernetes and a managed service before design starts.

PhaseWhat happensGate to the next phase
AssessWorkload and HA/DR targets captured; node, storage and Kubernetes inventory; operator fit and licence review; honest comparison with the managed service the cloud already offers, including costWritten recommendation with the metric behind it, signed off by the platform and database owners
Design and pilotCluster resources, storage classes, node pool, disruption budgets, TLS, backups and observability built in a pilot namespace; failover, restore and upgrade drills run and timedDrill results within the RPO and RTO targets; GitOps pipeline in place
MigrateLogical replication from the existing primary, parity checks (row counts, checksums, replayed queries), rehearsed cutover through the rw Service with reverse replication held openParity suite green; observation window passed; source decommission approved in writing
Operate24×7 senior watch, upgrade calendar, quarterly drills, monthly SLO and cost report, quarterly architecture reviewOngoing; every drill produces a runbook diff

Standing caveat: every procedure on this page is rehearsed on a non-production cluster with production-representative data before it is applied to production, with a verified backup taken first and a restore that has been timed. Operator versions and licence terms were verified on 28 September 2026 against the projects' own release pages; confirm image terms with the vendor before production use.

09 · The decision gate

PostgreSQL on Kubernetes or a managed service: we will tell you which

Because MinervaDB sells neither, the assessment can end with "not Kubernetes for this workload". That is a recommendation we make regularly, and it is the reason the ones that do go ahead succeed.

PostgreSQL on Kubernetes fits when

The platform team already runs Kubernetes in production with on-call and upgrade discipline; there are many clusters or tenants that benefit from consistent, GitOps-driven provisioning; data residency, licence terms or air-gap requirements rule out a managed service; or portability across clouds is a real, exercised requirement rather than a slide.

A managed service fits better when

There is no Kubernetes operations depth on call; the estate is one or two clusters with a standard topology; the provider's HA, backup and upgrade automation meets the RPO and RTO targets at acceptable cost; or the team's time is more valuable spent on the application than on storage classes. Cloud database FinOps →

10 · FAQ

PostgreSQL on Kubernetes questions we are asked most

Short answers to what platform and database leaders ask before the first call.

Is PostgreSQL production-ready on Kubernetes?

Yes, with a mature operator, zonal storage engineered for the write rate, scheduling that survives a node drain, and drills that prove failover and restore inside the targets. PostgreSQL on Kubernetes fails in production when those four things are assumed rather than engineered; that engineering is what MinervaDB delivers.

Which PostgreSQL on Kubernetes operator should we use?

It depends on how each behaves in a failover drill on your cluster, how long its restore takes on your data, its upgrade path across PostgreSQL majors, its licence terms for images as well as code, and the platform team's GitOps and observability stack. CloudNativePG, the Percona operator, Crunchy PGO, Zalando and StackGres each fit different estates; MinervaDB recommends one per estate with the evidence written down.

How is high availability achieved for PostgreSQL on Kubernetes?

A primary and replicas spread across zones with topology constraints, synchronous replication where zero data loss is required, an operator or Patroni that detects failure, fences the old primary and promotes a replica, Services whose endpoints move to the new primary, and poolers and drivers that reconnect and retry. Each step's duration is measured in a drill and the sum is the RTO the runbook quotes.

How do backups and point-in-time recovery work on Kubernetes?

Base backups and continuous WAL archiving to object storage through the operator (barman-cloud plugin, pgBackRest or WAL-G), with retention, encryption and immutability set to the compliance window. Recovery bootstraps a new Cluster resource from the repository to a timestamp or LSN, and the restore is timed quarterly on production-size data so the RTO is an observed number.

Can you migrate our existing PostgreSQL into Kubernetes without downtime?

Yes. The new cluster is built and drilled first, logical replication from the existing primary (VM, RDS or another cluster) drives lag to zero, parity is verified with row counts, checksums and replayed queries, and traffic is switched through the rw Service in a rehearsed window with reverse replication held open as the rollback.

Which Kubernetes platforms and clouds do you support?

Amazon EKS, Azure AKS, Google GKE, Red Hat OpenShift, Rancher, upstream Kubernetes on-premises and air-gapped clusters. The operator choice is independent of the distribution; storage class, node pool and CSI behaviour are engineered per platform.

When would you recommend a managed service instead of PostgreSQL on Kubernetes?

When there is no Kubernetes operations depth on call, the estate is one or two clusters with a standard topology, the provider's HA, backup and upgrade automation meets the RPO and RTO at acceptable cost, or the team's time is better spent on the application. MinervaDB sells neither, so the assessment says so when that is the answer.

What does 24x7 support for PostgreSQL on Kubernetes include?

A senior engineer on watch across APAC, EMEA and the Americas with S1 acknowledged in 15 minutes, S2 in 12 hours, S3 in 24 hours and S4 in 48 hours; operator and PostgreSQL upgrades on a calendar; quarterly restore and failover drills; observability tuned to the workload; and a monthly SLO and cost report.

Talk to a senior engineer about PostgreSQL on Kubernetes

Bring your Cluster manifests, the storage class definitions and the last incident timeline to the first call. We will tell you what would fail first in a drill, what it would cost to fix, and whether Kubernetes is the right home for this workload at all.