Data analytics platform engineering · CDC, Kafka, ClickHouse, lakehouse, serving and AI retrieval · Engineered and operated to signed SLOs
Data Analytics Platform Engineering: One Engineered System From Source Commit to Queryable Row, With 5 Signed SLOs
MinervaDB designs, builds and operates the platform your analytics and AI run on: change-data-capture ingestion with Debezium, Kafka transport, a ClickHouse columnar core delivered through ChistaDATA, Trino federation over Iceberg and Delta lakehouse tables, a Valkey or Redis serving layer, and a private in-VPC retrieval layer on Milvus or pgvector. It is engineered as one system rather than a chain of tools, and the same principal engineers who design it carry 24×7 operational responsibility under five signed objectives: freshness, latency, availability, RPO and RTO, and the cost curve. Vendor-neutral by principle: no cloud, warehouse or licence resale, so the platform is built from what your workload measures, not from what a vendor sells.
01 · Why data analytics platform engineering
Analytics initiatives fail at the platform layer, not in the dashboards
The models are fine and the dashboards are fine. What fails is the pipeline that stalled overnight, the warehouse that costs more every month, the freshness nobody measured, and the outage nobody owned because the platform was five vendors and no engineer.
One system, one owner
Data analytics platform engineering treats ingestion, transport, storage, federation, serving and retrieval as a single engineered system with one accountable team. A stalled Debezium slot, a Kafka consumer that fell behind and a ClickHouse merge backlog are the same incident, and one engineer owns it end to end.
SLOs the operator signs
Freshness, latency, availability, RPO and RTO, and the cost curve are defined with their measurement source before the build starts, and reported monthly with evidence. A platform without signed objectives is a collection of components; this is not that.
Open source, no lock-in
Debezium, Kafka, ClickHouse, Trino, Iceberg, Valkey and Milvus are all open source, and every configuration lives in your Git. MinervaDB and ChistaDATA resell nothing, so the platform is portable across clouds and away from either of us.
Measurement first
Every design decision and every recommendation cites the source table behind it: system.query_log, connector metrics, consumer-group offsets, pg_stat_statements, billing exports. Estimates are labelled as estimates; nothing is called an outcome until the number has moved.
02 · Data analytics platform engineering reference architecture
The data analytics platform engineering reference architecture
Sources, transport, analytics core, lakehouse, serving and retrieval, with an SLO measured at each tier. The components are chosen per workload; the shape is the same on AWS, Azure, Google Cloud and on-premises.
Figure 1. The data analytics platform engineering reference architecture: OLTP sources and event producers, Debezium CDC into Kafka, the ClickHouse columnar core, Iceberg and Delta lakehouse tables federated by Trino, the Valkey and Redis serving layer, the Milvus and pgvector retrieval layer, and the SLO measured at each tier.
Sources and transport
Data analytics platform engineering starts at the log: Debezium reads the transaction log of PostgreSQL, MySQL, SQL Server, MongoDB and Oracle rather than polling tables; event producers publish through OpenTelemetry, Vector or Fluent Bit. Kafka 4.x on KRaft carries everything, keyed by entity, with schema registry and retention sized so any downstream tier can be rebuilt by replay. Kafka consulting →
Analytics core and lakehouse
ClickHouse on ReplicatedMergeTree serves dashboards and APIs at sub-second p95, tiered by TTL to S3; Iceberg or Delta tables on object storage hold long-horizon history and ML feature tables; Trino federates both with the OLTP sources under one SQL surface. ClickHouse is delivered by ChistaDATA. ClickHouse consulting →
Serving and retrieval
Valkey or Redis hold hot aggregates and sessions with TTLs matched to the freshness budget; APIs carry p99 budgets with a fallback to ClickHouse. Milvus or pgvector serve a private, in-VPC retrieval layer fed from the same Kafka topics, under GDPR, DPDP and HIPAA-class governance. Vector data engineering →
03 · Data analytics platform engineering for CDC and streaming ingestion
Freshness is engineered from the transaction log, and measured end to end
Most freshness problems are ingestion problems: a connector holding a replication slot, a schema change nobody told the sink about, or a consumer group that fell behind at month-end. Data analytics platform engineering designs those out and measures the result commit to queryable.
Figure 2. Data analytics platform engineering for CDC and streaming ingestion: capture, transport, transform, sink and verify, with the four failure modes engineered out and the metric that catches each.
In data analytics platform engineering, capture starts with an initial consistent snapshot and continues from the log: logical replication slots on PostgreSQL, the binlog on MySQL, change streams on MongoDB, CDC tables on SQL Server. Slot lag and binlog age are alerted because a stalled connector fills the source's disk, and a heartbeat table keeps slots advancing on quiet databases. Topics are keyed by primary key with Avro or Protobuf schemas under a registry compatibility mode, compacted for current-state topics and retained long enough to replay a rebuild.
Transforms live in Flink or Kafka Streams when they need state, joins or late-data handling, and in ClickHouse materialized views when they do not. Sinks land in ReplacingMergeTree with version columns so at-least-once delivery produces exactly-once results, and in Iceberg for the lakehouse. Verification is continuous: row counts and checksums per window against the source, and freshness measured from source commit timestamp to queryable row.
-- Freshness SLI: source commit to queryable row, last hour
-- (source_commit_ts travels in the Debezium envelope)
SELECT
pipeline,
quantile(0.95)(
dateDiff('second', source_commit_ts, insert_ts)) AS p95_lag_s,
max(dateDiff('second', source_commit_ts, insert_ts)) AS max_lag_s,
count() AS rows_landed
FROM analytics.cdc_events
WHERE insert_ts >= now() - INTERVAL 1 HOUR
GROUP BY pipeline
ORDER BY p95_lag_s DESC;
# The same SLI seen from Kafka: consumer-group lag
kafka-consumer-groups.sh \
--bootstrap-server ${KAFKA_BOOTSTRAP} \
--describe --group clickhouse-orders-sink \
| awk 'NR>1 {lag+=$6} END {print "total lag:", lag}' 04 · Data analytics platform engineering for the columnar core and lakehouse
Hot, warm and cold tiers placed by query class, not by data age alone
Data analytics platform engineering puts each query class where it is cheapest to serve within its latency budget: ClickHouse on local NVMe for dashboards and APIs, ClickHouse on S3 for ad hoc analysis, Iceberg or Delta through Trino for batch and long horizons.
Figure 3. Data analytics platform engineering tiering for the columnar core and lakehouse: hot ClickHouse on local NVMe with projections and materialized views, warm ClickHouse on S3, cold Iceberg or Delta on object storage federated by Trino, and the placement rule per query class.
ClickHouse as the core
Data analytics platform engineering derives sort keys and partitioning from the query log, projections and AggregatingMergeTree materialized views for the top query shapes, TTL moves to S3 by age, and a Keeper ensemble sized for the failure domain. Acceptance is p95 per query shape from system.query_log, delivered and operated by ChistaDATA.
Lakehouse and federation
Iceberg or Delta tables on object storage for history and ML feature tables, written from the same Kafka topics, queried by Trino with connector push-down verified through EXPLAIN, and protected by resource groups so an analyst's scan never starves the hot tier. Databricks and Snowflake are integrated where inherited rather than replaced for their own sake.
Moving data between tiers
Tier moves are TTL rules and materialized views, never hand-run copies. Every move is rehearsed on a copy, reversible, and measured for the latency and cost it changes, so the placement rule is evidence rather than folklore. Warehouse modernization →
05 · Data analytics platform engineering for serving and AI retrieval
The last millisecond and the private retrieval layer
Dashboards and APIs are only as fast as the serving layer in front of the core, and the AI features built on the platform inherit its freshness, availability and governance. Data analytics platform engineering treats both as part of the system.
Figure 4. Data analytics platform engineering for serving, caching and AI retrieval: precomputed rollups into Valkey or Redis with API latency budgets, and the in-VPC retrieval layer with an embeddings pipeline, vector store selection, retrieval-augmented generation and governance.
Serving and caching
Data analytics platform engineering for the last millisecond: AggregatingMergeTree materialized views produce rollups refreshed by the stream rather than by cron; Valkey or Redis hold them keyed for the API's access pattern with TTLs matched to the freshness budget and stampede protection on cache miss. The p99 is measured at the API, not at the cache, and a timeout falls back to ClickHouse. Redis and Valkey support →
AI retrieval, in-VPC
An embedding job consumes the same Kafka topics with in-VPC models versioned so re-embedding is a replay; Milvus for scale and hybrid search, pgvector where the data already lives in PostgreSQL, with HNSW or IVF chosen by a recall-versus-latency test on your corpus. Row-level filters apply before retrieval, prompts and chunks are logged for audit, deletions propagate to the vector store, and no data leaves the VPC. Data science and AI consulting →
06 · Data analytics platform engineering SLOs
Five platform SLOs we contractually own, each with its measurement source
A managed platform is only as real as the objectives its operator will sign. Every managed data analytics platform engineering engagement defines these five before the build and reports against them monthly with evidence.
Figure 5. The data analytics platform engineering SLO framework: query latency, pipeline freshness, availability, RPO and RTO, and the cost curve, each with its definition, measurement source and reporting, plus the error-budget policy.
| Objective | Definition | Measurement source | Reported as |
|---|---|---|---|
| Query latency | p95 and p99 per workload class: dashboards, APIs, ad hoc, batch | system.query_log in ClickHouse, Trino query statistics, pg_stat_statements | Monthly percentiles per class against agreed thresholds |
| Pipeline freshness | End-to-end lag from source commit to queryable row, per pipeline | Debezium connector metrics, Kafka consumer-group offsets, sink insert timestamps | p95 lag per pipeline; breaches counted against the error budget |
| Availability | Success rate at the query interface, not per component; a healthy cluster behind a failed load balancer counts as an outage | Synthetic probes through the load balancer plus real error rates | Monthly availability with error-budget burn |
| RPO and RTO | Data at risk and time to recover, per tier | Quarterly restore and failover drills with timestamped evidence | Drill report with observed numbers written into the runbook |
| Cost curve | Cost per query and per ingested terabyte | Billing exports joined to query and ingestion volumes | Release-over-release trend with regressions explained |
Fast error-budget burn pages the on-call senior engineer; slow burn opens a change freeze on the affected tier until root cause is closed. Every number in the monthly pack is traceable to its source table. Standing caveat: every change is rehearsed on a non-production copy with production-representative data first, with a verified backup and a restore that has been timed.
07 · Data analytics platform engineering engagement lifecycle
Assess, build and pilot, migrate or take over, operate
Data analytics platform engineering engagements follow four phases, and the drills run on the pilot so nothing is discovered in production. MinervaDB owns the platform; ChistaDATA delivers the ClickHouse workstream under one agreement and one severity matrix.
Figure 6. The data analytics platform engineering engagement lifecycle: assessment, build and pilot with drills, migration or takeover with parity gates, and managed operations under signed SLOs, with the MinervaDB, ChistaDATA and customer split.
Assess
A data analytics platform engineering assessment measures workload classes, freshness and latency targets, data volumes and growth from the existing pipelines and warehouses rather than surveyed; a target architecture with tiering; and the five SLOs drafted with their measurement sources.
Build and pilot
CDC, Kafka, ClickHouse, lakehouse, serving and retrieval built as one system in a pilot; restore, failover and replay drills run and timed; every configuration in Git; runbooks drafted with verification and validation queries.
Migrate or take over
Dual-run against the existing warehouse or pipelines with parity gates on counts, checksums, freshness and query results; a rehearsed cutover with the reverse path held open. Platforms MinervaDB did not build are taken over the same way, after a read-only assessment.
Operate
24×7 senior watch with S1 acknowledged in 15 minutes, S2 in 12 hours, S3 in 24 hours and S4 in 48 hours; the monthly SLO evidence pack and cost curve; quarterly drills and an architecture review. 24×7 consultative support →
08 · FAQ
Data analytics platform engineering questions we are asked most
Short answers to what data and platform leaders ask before the first call.
What is data analytics platform engineering?
The discipline of designing, building and operating the platform that analytics and AI run on, ingestion, transport, columnar storage, lakehouse federation, serving and retrieval, as one engineered system with measured objectives. It sits at the platform and data-engineering layer, below the decision-science and dashboard layer, and it is what decides whether insight ships in milliseconds or dies in a queue of stalled pipelines.
Which technologies does the platform build on?
Open source throughout: Debezium for change data capture, Kafka 4.x on KRaft for transport, Flink or Kafka Streams for stateful transforms, ClickHouse for the columnar core, Iceberg or Delta tables on object storage federated by Trino, Valkey or Redis for serving, and Milvus or pgvector for the retrieval layer. Databricks, Snowflake and the cloud warehouses are integrated where a customer already runs them. Nothing is resold, and every configuration lives in the customer's Git.
Do you take over platforms you did not build?
Yes. Takeover starts with a read-only assessment that measures freshness, latency, availability and cost from the platform's own telemetry, then a dual-run against the existing pipelines with parity gates before MinervaDB assumes operational responsibility. The same five SLOs are signed, with the first month's numbers taken as the baseline.
What does 'contractually own' mean for SLOs?
Each objective has a definition, a measurement source, a threshold and an error budget written into the agreement, and MinervaDB reports against it monthly with evidence traceable to the source table. Fast error-budget burn pages the on-call senior engineer; slow burn triggers a change freeze on the affected tier until root cause is closed. SLOs are renegotiated on evidence, never assumed.
How does ChistaDATA fit into ClickHouse engagements?
ChistaDATA is MinervaDB's dedicated ClickHouse practice. On a data analytics platform engineering engagement MinervaDB owns the platform and the SLO agreement, and ChistaDATA designs, engineers and operates the ClickHouse tier under the same severity matrix. The customer holds one agreement and one escalation path.
How is freshness actually measured?
End to end: the source commit timestamp travels in the Debezium envelope and is compared with the insert timestamp at the sink, giving p95 and maximum lag per pipeline; Kafka consumer-group offsets give the same view from the transport side. Freshness is a per-pipeline SLO with its own budget, and a stalled connector or a lagging consumer is an incident before a dashboard goes stale.
Can the platform run on-premises or across clouds?
Yes. The reference architecture is the same on AWS, Azure, Google Cloud and on-premises, with object storage, Kubernetes or VMs chosen per environment. Because every component is open source and every configuration is in Git, the platform is portable across clouds and away from MinervaDB and ChistaDATA.
How long does an assessment take, and what does it need?
A fixed-scope assessment typically runs a few weeks, sized by the number of pipelines, sources and workload classes. It needs read-only access to the existing pipelines' metrics, the warehouse's query history and billing exports. It returns the target architecture, the tiering plan, the five draft SLOs with measurement sources, and a build or takeover plan with rollback at every phase.
Let's engineer the platform your analytics deserve
Bring the pipeline that stalled last month, the warehouse invoice, and the dashboards whose freshness nobody can state to the first call. We will tell you which tier is the constraint, what the five SLOs would be for your workload, and what we would build first.