MinervaDB × Amazon Web Services · AWS Data Platform Engineering · Aurora, RDS, Redshift, DynamoDB, S3 lakehouse, Glue, DMS
AWS Data Platform Engineering: Aurora, Redshift and S3 Lakehouse Built by Database Engineers
AWS makes provisioning effortless; it does not choose the store, size the cluster, set the distribution key or read the Cost and Usage Report for you. MinervaDB AWS Data Platform Engineering does that work: we design, migrate, tune and operate the AWS data estate with measured before-and-after numbers, and we stay on under a 24×7 SLA with a 15-minute Severity-1 response.
Reference platform
The estate AWS Data Platform Engineering designs and operates
AWS offers more data services than any other cloud, which is both its strength and its trap. A coherent estate uses few of them deliberately: an operational tier, a lake, a warehouse, pipelines between them and one governance plane over all of it. This is the shape our AWS Data Platform Engineering team builds, and every arrow in it has an owner, a latency target and a cost line.
Figure 1. The five layers of an AWS data platform and the governance and operations planes that span them, as delivered by AWS data platform engineering.
Ingest and orchestrate
In AWS data platform engineering we use AWS DMS for heterogeneous migration and CDC, zero-ETL integrations from Aurora, RDS and DynamoDB, Kinesis and Amazon MSK for streams, and Glue with Step Functions for batch. Every pipeline reports freshness and failures.
Store per workload
AWS Data Platform Engineering puts transactions on Aurora and RDS, key access at scale on DynamoDB, tolerant reads on ElastiCache for Valkey, the lake on Apache Iceberg in S3 and concurrent BI on Redshift.
Process and serve
For AWS data platform engineering serving layers we use Athena, EMR and Glue ETL for transformation, Redshift materialized views for marts, QuickSight for dashboards, and SageMaker or Bedrock on governed, well-modelled data.
Under AWS Data Platform Engineering, the relational tier gets the same depth as our dedicated Amazon RDS Support practice, and self-managed databases on EKS follow our cloud native database support standard. AWS’s own Well-Architected Data Analytics Lens is the checklist we start reviews from; the engineering judgment is in where an estate deviates from it and why.
Migration
Migration factory: Oracle and SQL Server into Aurora, RDS and Redshift
Enterprises rarely start clean. There is an Oracle or SQL Server estate nearing end of support, a warehouse straining under reporting and a deadline. AWS Data Platform Engineering runs migrations as a factory: repeatable steps, AWS DMS for full load and change data capture, and a cutover that is a gate with evidence rather than a date on a plan.
Figure 3. The six-step AWS data platform engineering migration factory and the lag gate that decides when cutover is allowed. The lag curve is an illustrative shape.
Schema and code conversion is where timelines slip, so it starts first. DMS Schema Conversion handles the mechanical share; PL/SQL packages, T-SQL procedures and vendor-specific data types get engineer review in Git. Large tables load in parallel segments, indexes are built after the load, and CDC then runs for weeks while the application team tests against the target.
Before cutover every table must show ValidationState = Validated under DMS data validation, CDC latency must hold under the agreed threshold, and a reverse replication task must already exist so rollback is a switch rather than a restore. The script alongside is the gate our AWS data platform engineering team runs: read-only checks, a typed confirmation, then validation after the change.
#!/usr/bin/env bash
# DMS cutover gate: read-only checks first, then an explicit confirmation before forward replication stops.
set -euo pipefail
TASK_ARN=${DMS_TASK_ARN:?set DMS_TASK_ARN}
TASK_ID=${DMS_TASK_ID:?set DMS_TASK_ID}; INSTANCE_ID=${DMS_INSTANCE_ID:?set DMS_INSTANCE_ID}
# 1. Verify: every table validated, no mismatched records
aws dms describe-table-statistics --replication-task-arn "$TASK_ARN" \
--query "TableStatistics[?ValidationState!='Validated'].[SchemaName,TableName,ValidationState,ValidationFailedRecords]" \
--output table
# 2. Verify: CDC latency on the target over the last 15 minutes (seconds)
aws cloudwatch get-metric-statistics --namespace AWS/DMS --metric-name CDCLatencyTarget \
--dimensions Name=ReplicationInstanceIdentifier,Value="$INSTANCE_ID" Name=ReplicationTaskIdentifier,Value="$TASK_ID" \
--start-time "$(date -u -d '-15 min' +%FT%TZ)" --end-time "$(date -u +%FT%TZ)" \
--period 60 --statistics Maximum --query "sort_by(Datapoints,&Timestamp)[-5:].[Timestamp,Maximum]" --output table
# 3. Confirmation gate: application writes must already be frozen at the source
read -r -p "Writes frozen, all tables Validated, latency at 0? Type CUTOVER to stop forward CDC: " ANSWER
[[ "$ANSWER" == "CUTOVER" ]] || { echo "aborted, replication still running"; exit 1; }
aws dms stop-replication-task --replication-task-arn "$TASK_ARN"
# 4. Validate: task stopped cleanly; start the pre-built reverse task that keeps the old source as rollback target
aws dms describe-replication-tasks --filters Name=replication-task-arn,Values="$TASK_ARN" \
--query "ReplicationTasks[0].[Status,StopReason]" --output text
Warehouse
Redshift physical design, workload management and Serverless
Redshift performance is decided by where rows live. A join on the distribution key stays local to each slice; a mismatched key shuffles data across the network on every query, and a stale sort order turns block skipping into full scans. AWS Data Platform Engineering designs tables from the join graph and the filter columns, then proves the design with system views.
Figure 4. Redshift design in AWS data platform engineering: co-located and redistributed joins across slices, the four distribution styles, and zone-map block skipping on a sort-key range filter.
We read DS_DIST_BOTH and DS_BCAST_INNER steps in EXPLAIN as findings, not details. The usual fix is a shared DISTKEY on the dominant join, DISTSTYLE ALL for small dimensions, and a sort key on the column most queries filter by range, usually a date. The Redshift table design guidance covers the mechanics.
Provisioned RA3 suits steady, predictable concurrency; Redshift Serverless suits bursty or intermittent load if the base RPU is set from measured demand rather than left at a default. Redshift reads and writes Apache Iceberg tables, with Iceberg v3 support added in August 2026, which lets the warehouse and the lake share one copy of curated data. That decision belongs in the AWS Data Platform Engineering design, not in a later clean-up.
-- Redshift table health: skew, unsorted share and stale statistics (works on provisioned and Serverless)
SELECT
"schema",
"table",
diststyle,
sortkey1,
tbl_rows,
size AS size_mb,
skew_rows, -- max rows per slice / min rows per slice; > 4 deserves a look
unsorted, -- % of rows not in sort-key order
stats_off -- 0 = current statistics, 100 = stale
FROM svv_table_info
WHERE tbl_rows > 1000000
ORDER BY size DESC
LIMIT 25;
Lakehouse
S3 and Apache Iceberg lakehouse engineering
On a lake, bytes scanned is the bill. Athena and Redshift Spectrum charge for data read, so file format, partitioning, file size and compaction set the cost of every query before it is written. AWS Data Platform Engineering builds the lake as zoned Iceberg tables with a maintenance schedule, not as a folder of files.
Figure 5. AWS Data Platform Engineering lake zones, Iceberg table anatomy and an illustrative comparison of bytes scanned under three layouts.
Iceberg’s hidden partitioning means queries filter on event_ts and the engine prunes partitions without users knowing the layout. Manifest statistics add file-level pruning on top, and columnar Parquet with ZSTD reads only the columns a query touches. Streaming writes create many small files, so compaction and snapshot expiry run on a schedule, per the Athena Iceberg documentation.
Where teams prefer a managed table store, Amazon S3 Tables provide Iceberg tables with automatic maintenance, and Glue zero-ETL can land DynamoDB and SaaS data directly into them. AWS Data Platform Engineering chooses between self-managed Iceberg and S3 Tables on cost and control, and catalogues both in Glue with Lake Formation permissions.
-- Athena (engine v3): curated Iceberg table with hidden partitioning and ZSTD Parquet
CREATE TABLE lake_curated.events (
event_id STRING,
customer_id BIGINT,
event_type STRING,
amount DECIMAL(18,2),
event_ts TIMESTAMP
)
PARTITIONED BY (day(event_ts), bucket(16, customer_id))
LOCATION 's3://${LAKE_BUCKET}/curated/events/'
TBLPROPERTIES (
'table_type' = 'ICEBERG',
'format' = 'parquet',
'write_compression' = 'zstd'
);
-- Weekly maintenance: compact small files from streaming writes, then expire old snapshots
OPTIMIZE lake_curated.events REWRITE DATA USING BIN_PACK
WHERE event_ts >= current_timestamp - INTERVAL '7' DAY;
VACUUM lake_curated.events; -- honours vacuum_max_snapshot_age_seconds; time travel before it is gone
Cost engineering
AWS Data Platform Engineering treats cost as a measured, attributed metric
An AWS bill that surprises the CFO almost always traces back to engineering choices: over-sized clusters, the wrong Aurora storage configuration, warehouses idle overnight, unpartitioned scans and retention nobody set. We work from the Cost and Usage Report so every dollar maps to a workload and an owner, and every change is verified with the same query a month later.
Figure 6. The monthly measure, attribute, act and verify loop, and where data-platform spend usually hides.
The first query ranks last month’s data-platform spend by service and workload tag from a CUR 2.0 data export. Untagged spend is itself a finding. The second computes each Aurora cluster’s I/O share: AWS guidance points to Aurora I/O-Optimized when I/O exceeds roughly a quarter of Aurora spend, and the switch is an engineering decision we test on the workload’s own numbers.
Typical AWS Data Platform Engineering actions follow from the evidence: right-sizing to measured load, Serverless v2 with auto-pause for dev and test, Redshift pause and resume or Serverless for intermittent use, S3 lifecycle and Intelligent-Tiering for cold data, VPC endpoints to remove NAT data-processing charges, and upgrades off RDS versions billing Extended Support. Deeper programmes run under our cloud database optimization and FinOps practice.
-- CUR 2.0 data export queried in Athena: last month's data-platform spend by service and workload tag
SELECT
line_item_product_code AS service,
COALESCE(resource_tags['user_workload'], '(untagged)') AS workload,
ROUND(SUM(line_item_unblended_cost), 2) AS cost_usd
FROM cur2.data_export
WHERE billing_period = '2026-08'
AND line_item_product_code IN ('AmazonRDS', 'AmazonRedshift', 'AmazonS3',
'AmazonDynamoDB', 'AmazonAthena', 'AWSGlue')
GROUP BY 1, 2
ORDER BY cost_usd DESC;
-- Aurora I/O share per cluster: AWS guidance points to I/O-Optimized when I/O exceeds ~25% of Aurora spend
SELECT
line_item_resource_id,
ROUND(SUM(CASE WHEN line_item_usage_type LIKE '%Aurora:StorageIOUsage%'
THEN line_item_unblended_cost ELSE 0 END), 2) AS io_usd,
ROUND(SUM(line_item_unblended_cost), 2) AS aurora_usd,
ROUND(100.0 * SUM(CASE WHEN line_item_usage_type LIKE '%Aurora:StorageIOUsage%'
THEN line_item_unblended_cost ELSE 0 END)
/ NULLIF(SUM(line_item_unblended_cost), 0), 1) AS io_pct
FROM cur2.data_export
WHERE billing_period = '2026-08'
AND line_item_product_code = 'AmazonRDS'
AND line_item_usage_type LIKE '%Aurora%'
GROUP BY 1
ORDER BY io_pct DESC;
Governance and security
Governance that auditors and engineers both accept
Governance fails when it is designed for one audience and resented by the other. The AWS governance stack, assembled well, gives security teams provable control and engineers a path to ship. AWS Data Platform Engineering maps each obligation to an enforced control, not a policy document.
| Obligation | AWS control we implement | How we prove it |
|---|---|---|
| Least-privilege data access | Lake Formation grants with LF-Tags; column and row filters | Grant inventory per role, reviewed quarterly |
| Identity | IAM Identity Center, permission sets, no long-lived keys | Credential report, access analyzer findings |
| Encryption | KMS customer-managed keys at rest; TLS enforced in transit | Key policies, rds.force_ssl / require_secure_transport |
| Network isolation | Private subnets, VPC endpoints, PrivateLink, no public endpoints | Config rules, security-group review |
| Audit | CloudTrail data events, database audit logs to CloudWatch | Queries that answer “who read this table” in minutes |
| Residency and regulation | Region placement, replication scoping; SOC 2, ISO 27001, HIPAA, GDPR, India’s DPDP Act | Control-to-evidence map per framework |
DataOps is the other half. Pipelines and schema changes live in Git and ship through CodePipeline or GitHub Actions with Terraform or CloudFormation, transformations are tested before they reach production, and lineage lets a team trace a wrong 9am dashboard to the upstream change within minutes. That is part of AWS data platform engineering, not an optional extra.
Engagement model
From assessment to managed operations
We meet the estate where it is: a greenfield build, a stalled migration, or a platform that works but costs too much. The front of every AWS data platform engineering engagement is deliberately short, because two weeks of measurement beat a quarter of strategy decks.
Phase 1
Assess
Workloads, access patterns, CUR spend, governance posture and risks, ending in a ranked, costed roadmap.
Phase 2
Architect
Store per workload, pipelines, lake and warehouse design, security and cost targets, reviewed with your engineers.
Phase 3
Engineer
Build and migration by senior engineers, with validation and a rollback path at every cutover.
Phase 4
Operate
24×7 operations under SLA: performance, cost, incidents and continuous optimization as data grows.
FAQ
Questions about AWS Data Platform Engineering with MinervaDB
The questions data leaders ask most often before an AWS data platform engineering engagement.
What does MinervaDB do on AWS that a generalist consultancy does not?
We are database engineers first. We reason about Aurora sizing and storage configuration, Redshift distribution and sort keys, Iceberg layout and the consumption patterns in the Cost and Usage Report, and we prove each change with before-and-after measurements.
Can you migrate Oracle and SQL Server estates to AWS?
Yes. AWS Data Platform Engineering migrates Oracle, SQL Server, MySQL and PostgreSQL into Aurora, RDS and Redshift with AWS DMS full load and CDC, DMS data validation on every table, and a gated cutover with reverse replication prepared for rollback.
Which AWS database should we use for our workload?
It depends on the access pattern. Aurora or RDS for transactional work, Aurora DSQL or DynamoDB global tables for multi-Region writes, DynamoDB for key access at scale, Redshift for concurrent BI, and Athena on Iceberg for lake scans. We confirm the choice from measured behaviour.
How do you reduce AWS data platform cost?
From the Cost and Usage Report: we attribute spend to workloads, then right-size, choose Aurora Standard or I/O-Optimized on measured I/O share, pause or go Serverless where load is intermittent, fix lake layout so Athena scans less, and verify savings with the same query a month later.
Should we use Redshift or Athena?
Redshift suits steady, concurrent BI and heavy ELT; Athena suits ad hoc and scheduled scans billed per terabyte. Many estates use both over one copy of Iceberg data, and we size each from query history rather than preference.
Do you only consult, or do you also run the platform?
Both. Most AWS data platform engineering engagements continue into 24×7 managed operations with a 15-minute Severity-1 response, monthly cost reviews and continuous performance work.
How do you handle governance and compliance?
Lake Formation grants and LF-Tags for data access, IAM Identity Center for identity, KMS customer keys, private networking and CloudTrail audit, mapped to SOC 2, ISO 27001, HIPAA, GDPR and India’s DPDP Act with evidence per control.
Will you tell us when AWS is the wrong answer?
Yes. We are vendor-neutral and do not resell AWS. When a workload is cheaper or safer self-managed, on another cloud or on another engine, we say so in writing before any project starts.
Build an AWS data estate that performs and pays off
Start with a two-week assessment: workloads, access patterns, CUR spend and risks, ranked by impact. AWS Data Platform Engineering for active incidents is available through the same page, 24×7.