The cloud console says Available. The dashboard is green. The provider’s status page has not had an incident in months. And yet the monthly invoice has doubled, a restore that “should have worked” did not, and a security researcher has just emailed about a database endpoint reachable from the internet. None of these failures came from the provider. They came from a belief: that a managed or serverless database takes care of everything, so the business no longer needs anyone who owns its data architecture.
That belief is one of the most expensive DBaaS risks a company can carry, precisely because it feels like a saving. This post explains where the provider’s responsibility ends, how serverless database costs spiral without design ownership, why automated backups are not recoverability, and what a data architect or data management expert actually does on a managed platform. Every guardrail comes with code you can apply this week. Figures in examples are illustrative, not benchmarks.
The myth: “the platform handles it”
DBaaS platforms such as Amazon RDS and Aurora, Google Cloud SQL, AlloyDB and BigQuery, and Azure SQL are excellent engineering. They remove the toil of racking servers, patching operating systems, replicating storage and running failover machinery. Serverless options go further and scale capacity automatically. It is natural to conclude that the database is now “someone else’s problem”.
The providers themselves say otherwise. AWS describes it as the shared responsibility model: the provider is responsible for security “of” the cloud, and the customer remains responsible for security “in” the cloud, which for a database means its configuration, access, data and use. Every major provider draws the same line. The business still owns everything that determines cost, correctness and trust.
Look at the left column of Figure 1, the heart of most DBaaS risks. None of those items is server administration. They are design and governance decisions, and they are exactly what a data architect does. Removing that role does not remove the decisions; it means they are made by default, by accident, or by whoever happens to be writing the next migration.
Financial destruction, part 1: elastic capacity scales waste
Serverless database costs follow usage, and usage is driven by design. A missing index turns a 5 ms lookup into a full table scan. An ORM that issues one query per row turns one request into hundreds. A dashboard that runs SELECT * over a year of data every minute scans terabytes a day. On a fixed-size server, these mistakes show up as slowness, and someone investigates. On an elastic platform, they show up as capacity: Aurora Serverless v2 adds ACUs, an on-demand table consumes more request units, BigQuery bills more bytes scanned. The application stays fast. The invoice absorbs the problem.
The ceiling is higher than many teams realise: AWS has extended Aurora Serverless v2 to 256 ACUs per instance. That headroom is a gift for genuine growth and a liability for unbounded design.
An illustrative example of the arithmetic: a report that scans 2 TB per run, scheduled every 15 minutes, scans close to 200 TB a month. Multiply by your on-demand rate per terabyte and the number is rarely small. The same report with a partition filter and a pre-aggregated table might scan a few gigabytes. That difference is not a provider setting. It is a design decision.
The first line of defence is finding the waste. On any managed PostgreSQL, pg_stat_statements ranks queries by the work they do, which is what elastic platforms charge for:
-- DBaaS risks, the design gap: find the queries that make elastic capacity expensive.
-- PostgreSQL on RDS, Aurora, Cloud SQL, AlloyDB or Azure (pg_stat_statements enabled).
SELECT queryid,
calls,
round(total_exec_time::NUMERIC / 1000, 1) AS total_s,
round((total_exec_time / NULLIF(calls, 0))::NUMERIC, 2) AS mean_ms,
shared_blks_read + shared_blks_hit AS blocks_touched,
round(rows::NUMERIC / NULLIF(calls, 0), 1) AS rows_per_call,
left(query, 70) AS query
FROM pg_stat_statements
ORDER BY shared_blks_read + shared_blks_hit DESC -- work done, which is what you pay for
LIMIT 15;
-- Tables read mostly by sequential scans: usually a missing or unusable index.
SELECT relname,
seq_scan,
seq_tup_read,
idx_scan,
n_live_tup
FROM pg_stat_user_tables
WHERE seq_scan > 0
AND n_live_tup > 100000
ORDER BY seq_tup_read DESC
LIMIT 15;
On BigQuery, the equivalent discipline is estimating before running and refusing to run past a ceiling. Google documents both per-query limits and custom query quotas for projects and users:
# Serverless database costs on BigQuery on-demand: you pay for bytes scanned,
# not for the rows you return. Estimate first, then refuse to run past a ceiling.
from google.cloud import bigquery
client = bigquery.Client() # credentials from the environment, never in code
SQL = """
SELECT customer_id, SUM(amount) AS revenue
FROM `analytics.orders`
WHERE order_date BETWEEN '2026-09-01' AND '2026-09-30' -- partition filter
GROUP BY customer_id
"""
# 1. Dry run: BigQuery plans the query and reports bytes it WOULD scan, for free.
dry = client.query(SQL, job_config=bigquery.QueryJobConfig(dry_run=True, use_query_cache=False))
gib = dry.total_bytes_processed / 1024**3
print(f"Estimated scan: {gib:,.1f} GiB")
# 2. Hard ceiling: the job fails instead of billing more than 50 GiB.
cfg = bigquery.QueryJobConfig(maximum_bytes_billed=50 * 1024**3)
rows = client.query(SQL, job_config=cfg).result()
# Project-wide protection belongs in custom quotas (query usage per day / per user),
# set by an administrator, so one dashboard or notebook cannot spend the month's budget.
Financial destruction, part 2: nobody set the ceilings
Finding waste is reactive. The structural fix for serverless database costs is to make spend a designed property of the system: a maximum capacity per cluster, a storage autoscaling cap, a budget with forecast alerts, and a daily anomaly check. These are a few lines of infrastructure code, but someone has to decide the numbers, defend them in a review and change them deliberately when the business grows.
# DBaaS risks, guardrail 2: put a ceiling on elastic capacity and an alarm on spend.
# Aurora Serverless v2: capacity is billed in ACUs. Without a sensible maximum,
# a bad query or traffic spike scales the bill as happily as it scales the database.
resource "aws_rds_cluster" "events" {
cluster_identifier = "events-prod"
engine = "aurora-postgresql"
engine_mode = "provisioned"
engine_version = "16.6"
storage_encrypted = true
deletion_protection = true
manage_master_user_password = true
master_username = "dbadmin"
serverlessv2_scaling_configuration {
min_capacity = 0.5 # ACUs; the floor you pay for when idle (0 enables auto-pause on supported versions)
max_capacity = 16 # ACUs; the ceiling is a business decision, written down and reviewed
}
}
# Monthly budget for the database service with early warnings.
resource "aws_budgets_budget" "database" {
name = "database-monthly"
budget_type = "COST"
limit_amount = "4000" # illustrative; derive from your unit economics
limit_unit = "USD"
time_unit = "MONTHLY"
cost_filter {
name = "Service"
values = ["Amazon Relational Database Service"]
}
notification { # forecast to exceed: act before the money is spent
comparison_operator = "GREATER_THAN"
threshold = 80
threshold_type = "PERCENTAGE"
notification_type = "FORECASTED"
subscriber_email_addresses = ["data-platform@example.com"]
}
}
Then watch the trend of serverless database costs, not just the threshold. A daily query over the Cost and Usage Report catches a doubling in days rather than at month end:
-- Serverless database costs: daily database spend from the AWS Cost and Usage Report
-- in Athena, with a simple anomaly flag. Column names follow the legacy CUR schema;
-- adjust them for CUR 2.0 exports.
WITH daily AS (
SELECT date(line_item_usage_start_date) AS usage_day,
product_product_name AS service,
SUM(line_item_unblended_cost) AS cost_usd
FROM cur.cost_and_usage
WHERE product_product_name IN ('Amazon Relational Database Service', 'Amazon DynamoDB')
AND line_item_usage_start_date >= date_add('day', -60, current_date)
GROUP BY 1, 2
)
SELECT usage_day,
service,
round(cost_usd, 2) AS cost_usd,
round(avg(cost_usd) OVER (PARTITION BY service ORDER BY usage_day
ROWS BETWEEN 14 PRECEDING AND 1 PRECEDING), 2) AS trailing_14d_avg,
CASE WHEN cost_usd > 1.5 * avg(cost_usd) OVER (PARTITION BY service ORDER BY usage_day
ROWS BETWEEN 14 PRECEDING AND 1 PRECEDING)
THEN 'INVESTIGATE' ELSE '' END AS flag
FROM daily
ORDER BY usage_day DESC, service;
Reputation compromise, part 1: backups that were never restores
“Backups are automatic” is true and dangerously incomplete. Automated snapshots are only as good as their retention, their location and the last time anyone restored one. Common gaps we find in DBaaS risks reviews: retention shorter than the time it takes to notice silent corruption, snapshots stored in the same account as production so one compromised credential can delete both, no tested point-in-time restore, and no idea how long a restore of the real dataset takes. The recovery time objective written in a contract is a guess until it has been measured.
A restore drill is the only evidence that counts. The script below restores production to a new instance, measures the elapsed time, validates the data and then removes the drill copy behind a verification step and an explicit confirmation gate:
#!/usr/bin/env bash
# DBaaS risks, guardrail 3: a backup only counts once it has been restored.
# Quarterly point-in-time restore drill on Amazon RDS. Measures real RTO.
set -euo pipefail
: "${SRC:?source DB instance id}" "${PGPASSWORD:?}" "${PGUSER:?}"
DRILL="${SRC}-restore-drill-$(date -u +%Y%m%d)"
START=$(date +%s)
# 1. Restore to a NEW instance (never touches the source).
aws rds restore-db-instance-to-point-in-time \
--source-db-instance-identifier "${SRC}" \
--target-db-instance-identifier "${DRILL}" \
--use-latest-restorable-time \
--no-publicly-accessible \
--tags Key=purpose,Value=restore-drill
aws rds wait db-instance-available --db-instance-identifier "${DRILL}"
END=$(date +%s)
echo "Restore completed in $(( (END - START) / 60 )) minutes <- your measured RTO for this size"
# 2. Validate the data, not just the instance status.
HOST=$(aws rds describe-db-instances --db-instance-identifier "${DRILL}" \
--query 'DBInstances[0].Endpoint.Address' --output text)
psql -h "${HOST}" -d orders -At -c "
SELECT 'orders', count(*), max(created_at) FROM orders
UNION ALL
SELECT 'payments', count(*), max(created_at) FROM payments;"
# 3. DESTRUCTIVE cleanup of the drill copy only, behind verification and a gate.
TAG=$(aws rds list-tags-for-resource \
--resource-name "$(aws rds describe-db-instances --db-instance-identifier "${DRILL}" \
--query 'DBInstances[0].DBInstanceArn' --output text)" \
--query "TagList[?Key=='purpose'].Value" --output text)
[[ "${TAG}" == "restore-drill" && "${DRILL}" != "${SRC}" ]] || { echo "Not a drill instance. Stop."; exit 1; }
read -r -p "Type 'DELETE ${DRILL}' to remove the drill copy: " ANSWER
[[ "${ANSWER}" == "DELETE ${DRILL}" ]] || { echo "Kept ${DRILL}."; exit 0; }
aws rds delete-db-instance --db-instance-identifier "${DRILL}" --skip-final-snapshot
aws rds wait db-instance-deleted --db-instance-identifier "${DRILL}"
aws rds describe-db-instances --db-instance-identifier "${SRC}" \
--query 'DBInstances[0].DBInstanceStatus' --output text # validation: source still "available"
Run it quarterly, record the time, and copy snapshots to a separate account and region. When a regulator, an auditor or a customer asks how long recovery takes, the answer becomes a measured number instead of a hope.
Reputation compromise, part 2: secure platform, insecure database
A managed database runs on hardened infrastructure. Whether it is reachable from the internet, who can log in, which roles can read raw personal data and whether the encryption keys are under your control are all configuration choices. Defaults vary by provider and console path, and “it worked in the demo” configurations have a way of reaching production. A safe baseline should be codified, reviewed and enforced, not remembered:
# DBaaS risks, guardrail 1: a managed database that is safe by default.
# Terraform (AWS provider 5.x). Every line below is a decision the provider
# will NOT make for you; leave one out and the default may be the risky option.
resource "aws_db_instance" "orders" {
identifier = "orders-prod"
engine = "postgres"
engine_version = "17" # pin the major; track its end-of-support date
instance_class = "db.r7g.xlarge"
allocated_storage = 500
max_allocated_storage = 1000 # cap storage autoscaling: growth is a cost decision
# Data protection
storage_encrypted = true
kms_key_id = aws_kms_key.db.arn # customer-managed key you can audit and rotate
deletion_protection = true # a stray "terraform destroy" cannot drop prod
backup_retention_period = 14 # days; align with your documented RPO
copy_tags_to_snapshot = true
skip_final_snapshot = false
final_snapshot_identifier = "orders-prod-final"
# Exposure
publicly_accessible = false
db_subnet_group_name = aws_db_subnet_group.private.name
vpc_security_group_ids = [aws_security_group.db_from_app_only.id]
iam_database_authentication_enabled = true
# Availability and observability
multi_az = true
performance_insights_enabled = true
monitoring_interval = 30 # seconds, Enhanced Monitoring
# Credentials never live in code: RDS stores the master secret in Secrets Manager.
username = "dbadmin"
manage_master_user_password = true
auto_minor_version_upgrade = true
maintenance_window = "sun:03:00-sun:04:00" # your quietest hour, not the default
}
Inside the database, the same discipline applies: few superusers, masked views for analysts, and row-level isolation for multi-tenant data. Under regimes such as GDPR, India’s DPDP Act and HIPAA, the question after an incident is not which cloud you used but what controls you had and whether you can prove it.
-- DBaaS risks, guardrail 4: least privilege and PII protection inside the database.
-- 1. Who can do everything? (rds_superuser on RDS/Aurora; adjust per platform)
SELECT r.rolname,
r.rolcanlogin,
ARRAY(SELECT b.rolname
FROM pg_auth_members m
JOIN pg_roles b ON b.oid = m.roleid
WHERE m.member = r.oid) AS member_of
FROM pg_roles AS r
WHERE r.rolcanlogin
AND (r.rolsuper
OR pg_has_role(r.oid, 'rds_superuser', 'MEMBER'))
ORDER BY r.rolname;
-- 2. Analysts see masked PII through a view; the base table stays locked down.
REVOKE ALL ON customers FROM analyst_role;
CREATE OR REPLACE VIEW customers_masked AS
SELECT customer_id,
left(email, 2) || '***@' || split_part(email, '@', 2) AS email_masked,
country_code,
created_at
FROM customers;
GRANT SELECT ON customers_masked TO analyst_role;
-- 3. Multi-tenant SaaS: each session only sees its own tenant's rows.
ALTER TABLE invoices ENABLE ROW LEVEL SECURITY;
CREATE POLICY invoices_tenant_isolation ON invoices
USING (tenant_id = current_setting('app.tenant_id')::BIGINT);
The slow-burn risks: upgrades, end of support and lock-in
Managed engines still reach end of support. On AWS, databases running major versions past the end of standard support can be automatically enrolled in Amazon RDS Extended Support, which carries an additional charge. That is a reasonable bridge for a planned migration and an unpleasant surprise for a business that did not know which versions it was running. A version inventory is the starting point for an upgrade calendar:
#!/usr/bin/env bash
# DBaaS risks, guardrail 5: know every engine version you run before the provider tells you.
# Inventory of RDS and Aurora engine versions across regions, for the upgrade calendar.
for region in us-east-1 eu-west-1 ap-south-1; do
aws rds describe-db-instances --region "${region}" \
--query 'DBInstances[].[DBInstanceIdentifier,Engine,EngineVersion]' --output text \
| awk -v r="${region}" '{print r"\t"$0}'
done | sort -k3,3 -k4,4V
# Compare each major version with the provider's release calendar. Versions past the
# end of standard support can be enrolled in paid Extended Support automatically.
Lock-in is the other slow burn. Proprietary features, provider-specific SQL extensions, egress pricing and tightly coupled event integrations are often the right choice, but each one raises the cost of leaving. A data architect records those choices and their exit cost while they are being made, rather than discovering them during a contract renegotiation.
What a data architect actually does on a managed platform
The role has changed, not disappeared. On a managed or serverless platform a data architect or data management expert spends little time on servers and most of it on decisions: which engine fits which workload, how data is modelled and keyed, what each feature will cost at ten times today’s traffic, what the recovery objectives are and how they are proven, who may see which data, and how the business could leave a platform if it had to.
Not every company needs a full-time hire to cover this. Many need a fractional data architect, or a partner who brings the role and a team behind it. What no company should do is leave the role empty and assume the platform fills it. DBaaS risks are not caused by the providers; they are caused by the gap between what the provider does and what the business assumes it does.
Why “we will hire a data architect later” costs more
Data design debt compounds in a way application debt often does not. A table keyed badly at launch holds a year of data by the time anyone notices; fixing it means a migration under live traffic, dual writes, backfills and a cutover window. A tenant model without isolation has to be retrofitted while customers are using it. A cost structure built on scanning raw events has to be redesigned while the invoice keeps arriving. The earlier a data architect reviews these choices, the cheaper each correction is.
There is also a negotiating cost. A business that cannot say what it spends per feature, how fast it can recover or how it would leave a provider has no leverage in a renewal and no credible answer in a due-diligence questionnaire. Investors, enterprise buyers and auditors increasingly ask these questions directly, and “the platform handles it” is not an answer they accept. Treating DBaaS risks as a board-level topic, with named owners, is cheaper than explaining an incident afterwards.
Twelve questions for leadership
If the answer to any of these is “the platform handles it” or “we are not sure”, the business is carrying unpriced DBaaS risks:
- What is the maximum monthly database spend, and what technically stops us exceeding it?
- Which ten queries cost the most, and who reviewed them?
- When did we last restore production, and how long did it take?
- Are backups copied to a separate account and region?
- Is any database endpoint reachable from the internet?
- How many people can read raw personal data?
- Who controls the encryption keys?
- Which engine versions do we run, and when does each reach end of support?
- What are our RPO and RTO per system, and are they measured or assumed?
- Which provider-specific features would make leaving expensive?
- Who is on call for the database, and what is their runbook?
- Who owns the data model and approves schema changes?
How MinervaDB closes the gap
MinervaDB is a vendor-neutral, full-stack database infrastructure company. We act as the data architect and data management team for businesses on Amazon RDS and Aurora, DynamoDB, Google Cloud SQL, AlloyDB and BigQuery, Azure SQL and Cosmos DB, Snowflake and Databricks, as well as self-managed PostgreSQL, MySQL, SQL Server, MongoDB and ClickHouse. A typical engagement starts with a DBaaS risks review covering cost, recoverability, security and lifecycle, then puts the guardrails in this post into code.
We have supported more than 900 enterprises from 46 cities over 15+ years, with 24×7 consultative support targets of 15 minutes for Severity 1, 12 hours for Severity 2, 24 hours for Severity 3 and 48 hours for Severity 4. Because we sell no licences, we will also tell you when a managed service is the right choice and when it is not.
Frequently asked questions
What are the biggest DBaaS risks for a business?
Uncontrolled cost from elastic scaling of inefficient queries, backups that have never been restored, insecure configuration and access, unplanned end-of-support upgrades, and lock-in that makes leaving expensive.
Why do serverless database costs grow so quickly?
Serverless platforms bill by capacity used, requests or bytes scanned. Missing indexes, chatty application code and unfiltered analytical queries consume more of each, and the platform scales to absorb them instead of failing.
Do we still need a data architect if we use a managed database?
Yes. The provider operates infrastructure; the business still owns the data model, queries, cost limits, recovery testing, access control and lifecycle planning. That ownership can be a full-time hire, a fractional role or a partner.
Are automated DBaaS backups enough for disaster recovery?
Not on their own. Recovery needs adequate retention, copies in a separate account and region, and regular restore drills that measure the real recovery time.
How can we cap serverless database costs?
Set maximum capacity per cluster, cap storage autoscaling, use per-query byte limits and custom quotas on scan-billed engines, configure budgets with forecast alerts, and review the most expensive queries regularly.
All Terraform, SQL, Python and scripts in this post are illustrative; limits, versions and thresholds are examples, not recommendations for your workload, and provider features and pricing change over time. Test every change in a non-production environment first, confirm current provider documentation, and maintain a robust disaster-recovery posture, including verified backups and rehearsed failover, before applying anything to production.
Not sure who owns your data architecture today? Book a DBaaS risks review with a MinervaDB principal architect, or email contact@minervadb.com.
Running this in production?
MinervaDB provides BigQuery Consulting and Google Cloud Data Platform Engineering with 24x7 coverage and a 15-minute S1 response. Talk to an engineer.