A data pipeline is a set of promises made to people who will never read its code. The analyst refreshing a dashboard at 09:00 assumes last night’s orders are in it; the finance model assumes the revenue table has no duplicates; the fraud service assumes the feature it reads was computed from events no older than a minute. None of those assumptions is written anywhere, which is why data engineering incidents so rarely look like software incidents: nothing crashed, every job reported success, and the number on the screen was wrong.
This page states the promises explicitly, as seven contracts a pipeline signs with its consumers, and for each one names the clause (what is promised, in a measurable form), the test that proves the clause is kept, and the archive post that shows the contract being honoured on a specific stack. The form is deliberate: a contract with no test is a hope, and a data engineering practice is the discipline of turning hopes into tests that run on every load.
The archive it introduces covers data quality SLOs, exactly-once semantics in Kafka, medallion-architecture governance, the single-source-of-truth lakehouse, demand forecasting and predictive maintenance pipelines, Trino anti-patterns, the legacy warehouse migration in BFSI, the retail stack, fractional data leadership, and the operational stores (MongoDB, MariaDB, MySQL connection pooling) whose behaviour the pipeline inherits.
Data engineering contract 1: freshness
The first clause is when the data will be there. Stated properly it is a number per table: “orders is complete through T minus 15 minutes, 99 percent of the time, measured by the max event timestamp against the wall clock at query time.” Stated the way it usually is (“the nightly job runs at 2 a.m.”) it is not a contract, because a job that ran does not imply that the data arrived, and a consumer who reads at 2:05 has no way to know whether the job that ran was the one that mattered.
data quality SLOs: five proven indicators is the archive’s post on turning freshness and the clauses that follow into service-level objectives with an error budget, and its first indicator is this one: freshness as a measured lag, alerted on the burn rate, not on the job scheduler’s exit code. The test is a query that runs after every load and writes the observed lag to a table that keeps history, so that the freshness SLO has a dashboard and the 09:00 analyst has a number to read before trusting the refresh.
Data engineering contract 2: completeness and exactly-once
The second clause is that every event arrives once. Both halves are hard, and they fail in opposite directions: at-most-once delivery drops events under a retry, at-least-once duplicates them, and exactly-once is a property of the whole path from producer to sink, not of any one component. A pipeline that promises exactly-once has to say where the idempotency lives (a producer id and sequence, a transactional consumer, an upsert keyed on a business identifier at the sink) and test for duplicates and gaps after every load.
Kafka exactly-once semantics: the ultimate guide is the archive’s post on the log side of that promise: idempotent producers, transactional writes across partitions, and the read_committed isolation a consumer needs to honour them. What the post is careful about is the boundary: Kafka’s guarantee ends at the consumer, and the sink’s upsert or merge is what carries it the rest of the way.
connection pooling best practices for MySQL at scale is in this archive for the same boundary in the other direction, where a pooler that retries a timed-out write is the most common source of a duplicate the log never saw.
-- contracts 1 and 2, tested after every load and written to history (PostgreSQL 12+; adapt for the warehouse)
CREATE TABLE IF NOT EXISTS pipeline_contract_check (
checked_at timestamptz NOT NULL DEFAULT now(),
table_name text NOT NULL,
contract text NOT NULL, -- 'freshness' | 'duplicates' | 'gaps' | 'schema' | 'volume'
observed numeric NOT NULL,
threshold numeric NOT NULL,
passed boolean GENERATED ALWAYS AS (threshold >= observed) STORED,
CONSTRAINT pk_pipeline_contract_check PRIMARY KEY (checked_at, table_name, contract)
);
-- freshness: lag in minutes between now and the newest event, threshold 15
INSERT INTO pipeline_contract_check (table_name, contract, observed, threshold)
SELECT 'orders', 'freshness',
EXTRACT(EPOCH FROM (now() - MAX(event_ts))) / 60,
15
FROM orders;
-- duplicates: rows sharing a business key in the last day, threshold 0
INSERT INTO pipeline_contract_check (table_name, contract, observed, threshold)
SELECT 'orders', 'duplicates',
COUNT(*) - COUNT(DISTINCT order_id),
0
FROM orders
WHERE event_ts >= now() - INTERVAL '1 day';
-- volume: today's row count against the trailing 7-day median, threshold 30 percent deviation
INSERT INTO pipeline_contract_check (table_name, contract, observed, threshold)
SELECT 'orders', 'volume',
ABS(today.n - med.n) * 100.0 / NULLIF(med.n, 0),
30
FROM (SELECT COUNT(*) AS n FROM orders WHERE event_ts >= date_trunc('day', now())) AS today,
(SELECT percentile_cont(0.5) WITHIN GROUP (ORDER BY n) AS n
FROM (SELECT date_trunc('day', event_ts) AS d, COUNT(*) AS n
FROM orders
WHERE event_ts >= now() - INTERVAL '8 days'
AND date_trunc('day', now()) > event_ts
GROUP BY 1) AS daily) AS med;
Three checks, one table, and a history that turns “the pipeline is fine” into a chart, which is what a data engineering test suite is for. The thresholds are illustrative; the real ones come from the consumer’s tolerance, which is the point of writing the contract down with the consumer in the room.
Data engineering contract 3: schema and meaning
The third clause is that a column means tomorrow what it meant today. Schema drift is the quiet failure: a producer adds a field, renames one, changes a unit from cents to dollars, or starts sending nulls where it sent zeros, and every job downstream keeps succeeding while every number downstream goes wrong.
The contract is a schema registry or a typed model with a compatibility rule (backward, forward, or full), a test that rejects an incompatible change at the producer, and a data dictionary that states the unit and the meaning of every column a consumer is allowed to read.
medallion architecture data governance: from data swamp to clarity is the archive’s post on layering that contract: bronze holds the raw event exactly as received, silver holds the conformed version with the schema enforced and the meaning fixed, gold holds the aggregate a consumer reads, and a change in meaning is a new silver table rather than an edit to the old one. a single source of truth on the Databricks lakehouse is the companion on where the one conformed copy lives and how every consumer is pointed at it rather than at a copy of a copy.
Data engineering contract 4: lineage and reproducibility
The fourth clause is that any number can be traced back to the events that produced it and reproduced from them. When the finance model shows a revenue figure that does not match the ledger, the question is which rows, from which load, transformed by which version of which job, and a pipeline that cannot answer it in an hour has no contract on lineage. The test is a reproduction: pick a figure from last week’s report, rebuild it from bronze with the job versions recorded at the time, and compare.
The archive’s migration post is the largest exercise of this clause. legacy data warehouse migration to Databricks in BFSI covers a regulated environment where every reported figure has to be reproducible for years, and the migration’s acceptance test was exactly this: the same figures, from the old warehouse and the new, reconciled to the row.
the top ten Trino performance anti-patterns is in this archive because a federated query engine is often the lineage tool of last resort, joining the source system to the warehouse to find where a figure diverged, and the anti-patterns are what make that join take hours instead of minutes.
Data engineering contract 5: latency for the consumers that act on data
The fifth clause is for the consumers that are not people. A fraud model, a pricing service and a predictive-maintenance alert all read features that a pipeline computed, and for them freshness is measured in seconds and latency in milliseconds, with the contract stated at a percentile. The batch pipeline’s freshness clause is not enough; these consumers need a streaming path with its own end-to-end latency budget, from the event’s timestamp to the feature’s availability, measured continuously.
predictive maintenance data platform: six layers and a demand forecasting pipeline for CPG are the archive’s two worked examples of the streaming and the batch contract living side by side: sensor events and daily sell-through, each with its own freshness and latency clauses, feeding models with different tolerances. MongoDB performance tuning from 1 ms to 100 microseconds is the operational-store side of the same clause, where the feature is served from a document store and its read latency is the last leg of the budget.
Data engineering contract 6: cost per delivered row
The sixth clause is what the pipeline costs to keep its other promises, stated as a unit cost so that it can be compared: compute and storage per million rows delivered to gold, per table, per month. A pipeline whose cost per row rises month over month is making a promise it will eventually be unable to afford, and the usual causes are the same ones the warehouse bill shows: a full reload where an incremental load would do, a small-file problem in the lake, a join that shuffles the largest table, a retention policy nobody set.
the modern retail data analytics stack is the archive’s post that prices the layers of a stack end to end, and tuning MariaDB for cloud and containerised environments is the source-side cost that a CDC pipeline imposes on the operational database: binary log volume, replica lag, and the connection the CDC connector holds open. The test is the unit cost trended, and the review question is which contract clause a rising cost is paying for and whether the consumer still needs that clause at that level.
Data engineering contract 7: recoverability
The seventh clause is what happens when any of the first six is broken: how a bad load is detected, quarantined, and replayed without the consumer seeing a partial state. The contract names the replay source (bronze, the log’s retention, or the source system), the replay time for a day and for a month, and the mechanism that keeps a half-loaded table from being read (a swap, a version pointer, or an atomic partition replace). A pipeline that cannot replay is a pipeline whose bronze layer is decorative.
The test is a drill, on the same footing as a database restore drill, and it is the data engineering practice’s equivalent of a restore: pick a day, corrupt the silver table for it on a copy, replay from bronze, reconcile against the untouched copy, and record the elapsed time. the fractional chief data officer and real-time analytics is the archive’s post on who owns that drill in an organisation without a full-time data leader, and the answer it gives is that ownership of the seven contracts is the job, whoever holds the title.
The seven data engineering contracts, side by side
| Contract | Clause, stated measurably | Test that proves it | Usual breach |
|---|---|---|---|
| 1. Freshness | Complete through T minus N minutes, P percent of the time | Observed lag after every load, written to history, alerted on burn rate | Job succeeded, data did not arrive |
| 2. Completeness, exactly-once | Every event once; idempotency located at a named component | Duplicate and gap counts on a business key after every load | Retry at a pooler or connector; consumer without read_committed |
| 3. Schema and meaning | Compatibility rule enforced; unit and meaning of every readable column fixed | Incompatible change rejected at the producer; conformed layer versioned | Unit change or renamed field, every job green |
| 4. Lineage, reproducibility | Any figure traceable to its rows and job versions; rebuildable | Reproduce last week’s figure from bronze with recorded versions | Cannot say which load produced a number |
| 5. Latency for machines | Event-to-feature latency at a percentile, on the streaming path | Continuous end-to-end timing from event timestamp to availability | Batch freshness offered to a real-time consumer |
| 6. Cost per delivered row | Compute plus storage per million rows to gold, per table, per month | Unit cost trended; rising cost mapped to the clause it pays for | Full reloads, small files, unset retention |
| 7. Recoverability | Replay source, replay time for a day and a month, partial-state guard | Quarterly replay drill on a copy, reconciled, timed | Bronze exists; nobody has ever replayed from it |
Every clause in the second column is a number or a named component, which is what makes the third column possible; a clause that cannot be tested is rewritten until it can.

How data engineering contracts are negotiated, and with whom
The contracts are written with the consumer, not for them, and the negotiation is where a data engineering practice earns its keep. A consumer asked for freshness will say “real time”; the engineer’s job is to ask what decision the data feeds and how often that decision is made, and to write the clause at the tolerance the decision actually has. A dashboard read twice a day does not need a fifteen-minute freshness clause, and a fraud model does not need a nightly one.
The cost clause is what keeps the negotiation honest: every tightening of clauses one to five has a price in clause six, and the consumer sees it.
The data engineering negotiation produces a document per table (or per domain, in a larger estate) that states the seven clauses and names the owner on each side. It is reviewed when a consumer changes, when a source changes, and when the unit cost moves, and the check history from the tests above is the evidence brought to the review. The archive’s leadership post above describes that review as the core of the fractional data officer’s calendar, and the reason is that it is the one meeting where the pipeline’s promises and its costs are in the same room.
Version notes: the mechanisms named here move with releases (Kafka’s transactional protocol and the KRaft-era defaults, Delta and Iceberg’s schema evolution and compatibility rules, the catalog views a warehouse exposes for volume and freshness checks), so every archive post’s specifics should be checked against the running versions.
Every contract test above should be run on a copy of the production tables before it is scheduled against production, since a badly written check is itself a load on the store it checks, and the replay drill in clause seven is the DR posture of a pipeline: a bronze layer that has never been replayed from is a backup that has never been restored.
Where this data engineering archive sits
This archive is the pipeline layer between the operational engines and the analytical homes on minervadb.com. The Kafka archive is the log that most of the contracts ride on, the data warehousing archive is where clause six is read line by line, the data strategy archive is the decision layer that chooses the homes the pipeline connects, and the monitoring archive is where the check history above becomes an alert. The open specification behind the compatibility rules in clause three is the schema evolution section of the Apache Iceberg table specification.
For a pipeline review that writes the seven contracts down and installs the tests, a migration whose acceptance criterion is reproducibility to the row, or 24×7 managed data operations in which the contracts are what the on-call engineer is paged on, the MinervaDB database consulting practice starts from this page, and states for every recommendation which contract it strengthens and what the test showed before and after.