PostgreSQL on Kubernetes · Design, migration, operations and 24×7 support · Any operator, any cluster, any cloud
PostgreSQL on Kubernetes, Engineered to Production Standards: Operators, Storage, Failover and Backups You Can Drill
MinervaDB designs, migrates and operates PostgreSQL on Kubernetes for enterprises whose platform teams already run Kubernetes in production, and tells the others honestly when a managed service is the better fit. The work covers operator selection (CloudNativePG, Percona, Crunchy PGO, Zalando, StackGres), storage and scheduling engineering, failover and restore drills that produce measured RPO and RTO, and a 24×7 senior watch with S1 acknowledged in 15 minutes. Every change is declared in Git, reviewed, reversible and rehearsed before it reaches a production cluster.
01 · Why PostgreSQL on Kubernetes needs database engineering
The operator automates the mechanics; the outcomes still have to be engineered
Kubernetes is the default control plane for application platforms, and running PostgreSQL on Kubernetes lets one team provision, upgrade and observe databases the same way it does everything else. The operator does not choose your storage class, size your memory, or time your restore.
A mature operator turns a Cluster manifest into pods, persistent volumes, Services, replication and backups, and it will promote a replica when the primary disappears. What it cannot do is decide whether the persistent volume can follow the pod to another zone, whether the memory limit leaves room for the page cache, whether a node drain will take the primary and its synchronous replica down together, or whether the backup in object storage restores in ten minutes or four hours.
Those are database engineering decisions, and they are where PostgreSQL on Kubernetes projects succeed or stall. MinervaDB brings two decades of PostgreSQL production experience to the manifests, the storage layer and the runbooks, so the cluster behaves the way the diagram promises when a node actually dies.
The practice is vendor-neutral by principle. We sell no operator, no distribution and no cloud, so the recommendation can be CloudNativePG, Percona's operator, a commercial vendor, or an honest "use RDS for this one". Every recommendation is justified by a metric, a drill result or a licence term you can check yourself.
Whatever the answer, the same senior engineers who design the platform carry the 24×7 watch afterwards, with severity targets written into the agreement rather than into a slide.
02 · PostgreSQL on Kubernetes reference architecture
What a production PostgreSQL on Kubernetes cluster looks like when it is engineered
One namespace per cluster, an operator reconciling a declarative Cluster resource, a primary and replicas spread across zones on their own persistent volumes, poolers in front, WAL flowing to object storage, and observability wired from day one.
Figure 1. PostgreSQL on Kubernetes reference architecture: operator control plane, rw and ro Services, primary and replicas across three zones on CSI-backed persistent volumes, PgBouncer poolers, WAL archiving to object storage, secrets, TLS and Prometheus, with the pieces that live outside the cluster.
Declarative and reviewable
PostgreSQL on Kubernetes is declarative by nature: the Cluster resource, pooler, backup schedule and scheduled restores live in Git and are applied by Argo CD or Flux. A change to shared_buffers, a replica count or a storage class is a pull request with a diff, a review and a rollback, not a command typed on a node.
Topology that survives a zone
Topology spread constraints and anti-affinity place the primary and each replica in different zones, a storage class per zone with WaitForFirstConsumer keeps the volume where the pod lands, and a PodDisruptionBudget stops a routine node drain from taking two members at once.
Services, not IPs
Applications connect to the rw Service for writes and the ro Service for replica reads, through PgBouncer poolers with per-role pool sizes and TLS. A failover moves endpoints; the application only needs a driver that reconnects and retries idempotently.
03 · PostgreSQL on Kubernetes operator selection
Choosing the right PostgreSQL on Kubernetes operator, by evidence
No operator is best for every organisation. The choice is weighted on how it behaves in a failover drill, how long its restore takes on your data, its upgrade path across PostgreSQL majors, and the licence terms of its images as well as its code.
Figure 2. Operator selection matrix for PostgreSQL on Kubernetes: HA mechanism, backup and PITR tooling, pooling, licence and stewardship, and best fit, with the weighting MinervaDB applies. Verified 28 September 2026.
| Operator | How it does HA | What we check before recommending it |
|---|---|---|
| CloudNativePG 1.29 | Operator-native: an instance manager in each pod, primary election through the Kubernetes API, no Patroni or external DCS | Behaviour when the operator itself is down, volume-snapshot backup support on your CSI driver, barman-cloud plugin retention, major-version upgrade procedure, supported Kubernetes versions |
| Percona Operator for PostgreSQL 2.x | Patroni inside the pods with Kubernetes as the DCS; pgBackRest for backups and PITR; PgBouncer built in | Fit with an estate already on Percona tooling for MySQL or MongoDB, pgBackRest repository layout, pgvector and extension images |
| Crunchy Postgres for Kubernetes (PGO) 5.x | Patroni inside the pods; pgBackRest with multiple repositories; PgBouncer built in | Container image licence terms for production use versus the Apache-licensed operator code, commercial support scope, monitoring stack integration |
| Zalando postgres-operator | Patroni in the Spilo image; WAL-G or WAL-E to object storage; connection pooler | Community cadence and PostgreSQL major support in Spilo, upgrade path, fit for teams already standardised on Patroni outside Kubernetes |
| StackGres | Patroni; pgBackRest through SGBackup; PgBouncer and Envoy; management UI | AGPL v3 obligations, commercial edition boundaries, UI versus GitOps workflow for the platform team |
PostgreSQL on Kubernetes also runs on any conformant distribution: Amazon EKS, Azure AKS, Google GKE, Red Hat OpenShift, Rancher, upstream Kubernetes on-premises, and air-gapped clusters. The operator choice is independent of the distribution; the storage class and node pool design are not.
04 · PostgreSQL on Kubernetes failover and self-healing
What happens when a node dies, second by second
High availability for PostgreSQL on Kubernetes is a sequence of timeouts and endpoint changes, each of which can be measured. We measure them in a drill before production does it for us.
Figure 3. Failover sequence for PostgreSQL on Kubernetes, the four components of RTO that a drill measures, and the scheduling and storage mistakes that turn a ten-second failover into a ten-minute outage.
Detection tuned, not defaulted
For PostgreSQL on Kubernetes, readiness-probe periods, failure thresholds and lease TTLs are set against the false-positive rate your nodes actually produce under memory pressure. Too aggressive and a busy checkpoint triggers a failover; too lax and RTO is measured in minutes.
Volumes that follow the pod
A persistent volume still attached to a dead node is the most common cause of a stalled recovery. Zonal storage classes, CSI attach and detach timeouts and, where the platform allows, volume-attachment force-detach are configured and then tested by killing the node, not by reading the docs.
Rejoin without a full re-clone
The old primary returns as a replica through pg_rewind where the timeline allows, or a fresh clone from the backup repository where it does not. Both paths are in the runbook with the expected duration on your data size.
05 · Backup, WAL archiving and PITR
PostgreSQL on Kubernetes backups that have been restored, on a schedule
Every operator can ship base backups and WAL to object storage. Whether the archive is complete, encrypted, retained for the compliance window and restorable inside the RTO is engineering, and it is proved by a timed restore into a new namespace.
Figure 4. Backup, WAL archiving and point-in-time recovery for PostgreSQL on Kubernetes: base backups and continuous WAL to object storage, restore into a new Cluster resource with a recovery target, and the quarterly drill that produces the RTO the runbook quotes.
# CloudNativePG 1.29: a recovery cluster bootstrapped from
# the object-store backup of "orders-db" to a point in time
apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
name: orders-db-restore
spec:
instances: 1
imageName: ghcr.io/cloudnative-pg/postgresql:18.6
storage:
storageClass: gp3-zone-a
size: 500Gi
bootstrap:
recovery:
source: orders-db
recoveryTarget:
targetTime: "2026-09-28 06:10:00+00"
externalClusters:
- name: orders-db
plugin:
name: barman-cloud.cloudnative-pg.io
parameters:
barmanObjectName: minervadb-backups
serverName: orders-db The PostgreSQL on Kubernetes restore drill loads production-size data into a new namespace, replays WAL to a timestamp, runs pg_amcheck, compares row counts and checksums against the source and smoke-tests the application through the ro Service. The elapsed time is written into the runbook as the RTO; the WAL archive lag at any moment is the RPO, and it is alerted at a fraction of the budget.
Retention, encryption keys and object-lock immutability are set against the compliance window rather than the tool's default, and the cross-region copy is tested by restoring in a second cluster. The failure cases rehearsed include namespace deletion, persistent-volume loss, an operator upgrade gone wrong and region loss.
06 · PostgreSQL on Kubernetes resources, storage and scheduling
Sizing PostgreSQL on Kubernetes so the kernel never kills the database
An OOM-killed primary is a design failure. Memory, CPU, storage and scheduling are engineered from measured workload behaviour, and the evidence that proves the sizing is collected from PostgreSQL, the kubelet and the storage layer together.
Figure 5. Resource, storage and scheduling engineering for PostgreSQL on Kubernetes: Guaranteed QoS, shared_buffers versus the memory limit, CPU pinning without throttling, zonal storage classes and provisioned IOPS, node pools and disruption budgets, with the evidence used to size and prove each decision.
Memory
PostgreSQL on Kubernetes memory: requests equal limits for Guaranteed QoS. shared_buffers at roughly a quarter of the limit, with headroom for work_mem times the pool size, maintenance work and the page cache that lives inside the same cgroup. Huge pages where the node pool allows them.
CPU
Whole-core requests and a static CPU manager policy for latency-sensitive primaries. CPU limits are avoided or set generously, because CFS throttling during a checkpoint or an anti-wraparound vacuum is exactly the wrong moment to slow PostgreSQL down.
Storage
A storage class per zone with WaitForFirstConsumer, IOPS and throughput provisioned to the measured write rate, pg_test_fsync run against the class before go-live, and a separate WAL volume or local NVMe where the operator and the workload justify it.
Scheduling
A dedicated node pool with taints and tolerations, topology spread across zones, a PodDisruptionBudget that keeps a quorum through drains and upgrades, and a priority class above application pods so the database is never the first thing evicted.
07 · Services
PostgreSQL on Kubernetes services across the full lifecycle
Design and migrate, optimise and audit, or hand the clusters to a 24×7 senior watch. Each service is delivered by the same engineers and hands back versioned manifests and runbooks.
Design and migrate
Operator and distribution selection, node pool and storage design, the Cluster resources and GitOps pipeline, and migration from VMs, RDS or another cluster by logical replication with parity checks and a rehearsed cutover through the rw Service, reverse replication held as the rollback. Data modernization →
Optimise and audit
Already running PostgreSQL on Kubernetes? A read-only audit of manifests, storage, scheduling, failover behaviour, backup completeness and PostgreSQL configuration, returned as findings ranked P0 to P2 with the metric each fix moves and the drill that proves it. PostgreSQL consulting →
24×7 support and operations
Senior engineers on watch with S1 acknowledged in 15 minutes, S2 in 12 hours, S3 in 24 hours and S4 in 48 hours; operator and PostgreSQL upgrades on a calendar; quarterly restore and failover drills; a monthly SLO report. 24×7 consultative support →
Security and compliance
TLS everywhere with cert-manager or operator-issued certificates, credentials from Secrets or Vault with rotation, network policies scoped to the namespace, pgaudit retained for the compliance window, and the evidence pack for GDPR, HIPAA, PCI DSS and SOC 2.
Observability
PostgreSQL on Kubernetes observability: operator exporters and custom queries into Prometheus, Grafana dashboards for replication lag, WAL archive status, persistent-volume usage, cgroup throttling and vacuum progress, and alerts tuned to the workload rather than a vendor default.
Upgrades
PostgreSQL on Kubernetes upgrades: operator upgrades rehearsed on a staging cluster first; PostgreSQL major-version upgrades by logical replication into a new cluster, or the operator's in-place path where it exists, with the restore timed before either starts. Current targets: PostgreSQL 18.6, or 17 where an extension lags.
08 · PostgreSQL on Kubernetes engagement lifecycle
Assess, design and pilot, migrate, operate
Every PostgreSQL on Kubernetes engagement follows the same four phases. The pilot is where the drills run for the first time, so nothing is discovered in production.
Figure 6. The MinervaDB PostgreSQL on Kubernetes engagement lifecycle and the decision gate that decides between Kubernetes and a managed service before design starts.
| Phase | What happens | Gate to the next phase |
|---|---|---|
| Assess | Workload and HA/DR targets captured; node, storage and Kubernetes inventory; operator fit and licence review; honest comparison with the managed service the cloud already offers, including cost | Written recommendation with the metric behind it, signed off by the platform and database owners |
| Design and pilot | Cluster resources, storage classes, node pool, disruption budgets, TLS, backups and observability built in a pilot namespace; failover, restore and upgrade drills run and timed | Drill results within the RPO and RTO targets; GitOps pipeline in place |
| Migrate | Logical replication from the existing primary, parity checks (row counts, checksums, replayed queries), rehearsed cutover through the rw Service with reverse replication held open | Parity suite green; observation window passed; source decommission approved in writing |
| Operate | 24×7 senior watch, upgrade calendar, quarterly drills, monthly SLO and cost report, quarterly architecture review | Ongoing; every drill produces a runbook diff |
Standing caveat: every procedure on this page is rehearsed on a non-production cluster with production-representative data before it is applied to production, with a verified backup taken first and a restore that has been timed. Operator versions and licence terms were verified on 28 September 2026 against the projects' own release pages; confirm image terms with the vendor before production use.
09 · The decision gate
PostgreSQL on Kubernetes or a managed service: we will tell you which
Because MinervaDB sells neither, the assessment can end with "not Kubernetes for this workload". That is a recommendation we make regularly, and it is the reason the ones that do go ahead succeed.
PostgreSQL on Kubernetes fits when
The platform team already runs Kubernetes in production with on-call and upgrade discipline; there are many clusters or tenants that benefit from consistent, GitOps-driven provisioning; data residency, licence terms or air-gap requirements rule out a managed service; or portability across clouds is a real, exercised requirement rather than a slide.
A managed service fits better when
There is no Kubernetes operations depth on call; the estate is one or two clusters with a standard topology; the provider's HA, backup and upgrade automation meets the RPO and RTO targets at acceptable cost; or the team's time is more valuable spent on the application than on storage classes. Cloud database FinOps →
10 · FAQ
PostgreSQL on Kubernetes questions we are asked most
Short answers to what platform and database leaders ask before the first call.
Is PostgreSQL production-ready on Kubernetes?
Yes, with a mature operator, zonal storage engineered for the write rate, scheduling that survives a node drain, and drills that prove failover and restore inside the targets. PostgreSQL on Kubernetes fails in production when those four things are assumed rather than engineered; that engineering is what MinervaDB delivers.
Which PostgreSQL on Kubernetes operator should we use?
It depends on how each behaves in a failover drill on your cluster, how long its restore takes on your data, its upgrade path across PostgreSQL majors, its licence terms for images as well as code, and the platform team's GitOps and observability stack. CloudNativePG, the Percona operator, Crunchy PGO, Zalando and StackGres each fit different estates; MinervaDB recommends one per estate with the evidence written down.
How is high availability achieved for PostgreSQL on Kubernetes?
A primary and replicas spread across zones with topology constraints, synchronous replication where zero data loss is required, an operator or Patroni that detects failure, fences the old primary and promotes a replica, Services whose endpoints move to the new primary, and poolers and drivers that reconnect and retry. Each step's duration is measured in a drill and the sum is the RTO the runbook quotes.
How do backups and point-in-time recovery work on Kubernetes?
Base backups and continuous WAL archiving to object storage through the operator (barman-cloud plugin, pgBackRest or WAL-G), with retention, encryption and immutability set to the compliance window. Recovery bootstraps a new Cluster resource from the repository to a timestamp or LSN, and the restore is timed quarterly on production-size data so the RTO is an observed number.
Can you migrate our existing PostgreSQL into Kubernetes without downtime?
Yes. The new cluster is built and drilled first, logical replication from the existing primary (VM, RDS or another cluster) drives lag to zero, parity is verified with row counts, checksums and replayed queries, and traffic is switched through the rw Service in a rehearsed window with reverse replication held open as the rollback.
Which Kubernetes platforms and clouds do you support?
Amazon EKS, Azure AKS, Google GKE, Red Hat OpenShift, Rancher, upstream Kubernetes on-premises and air-gapped clusters. The operator choice is independent of the distribution; storage class, node pool and CSI behaviour are engineered per platform.
When would you recommend a managed service instead of PostgreSQL on Kubernetes?
When there is no Kubernetes operations depth on call, the estate is one or two clusters with a standard topology, the provider's HA, backup and upgrade automation meets the RPO and RTO at acceptable cost, or the team's time is better spent on the application. MinervaDB sells neither, so the assessment says so when that is the answer.
What does 24x7 support for PostgreSQL on Kubernetes include?
A senior engineer on watch across APAC, EMEA and the Americas with S1 acknowledged in 15 minutes, S2 in 12 hours, S3 in 24 hours and S4 in 48 hours; operator and PostgreSQL upgrades on a calendar; quarterly restore and failover drills; observability tuned to the workload; and a monthly SLO and cost report.
Talk to a senior engineer about PostgreSQL on Kubernetes
Bring your Cluster manifests, the storage class definitions and the last incident timeline to the first call. We will tell you what would fail first in a drill, what it would cost to fix, and whether Kubernetes is the right home for this workload at all.