A data strategy that cannot be written down as a short list of decisions is not a strategy; it is a slide. The test is simple: for each decision, can the organisation say what was chosen, what was rejected, on what measurement, and what would cause the choice to be revisited? Most enterprises can answer that for their cloud provider and for almost nothing else, which is why their database estates look the way they do: a dozen engines chosen by whichever team arrived first, three copies of the customer table, and a warehouse whose bill nobody can attribute.
This page states a data strategy as nine decision records. Each record has the same four fields (chosen, rejected, measured by, revisit when), and each section explains what the decision governs, what evidence it should be made on, and which archive posts show the decision being made for a specific engine or workload. The form is borrowed from architecture decision records because it forces the rejected option and the revisit trigger to be written down, which is what distinguishes a strategy from a preference.
The archive it introduces covers polyglot persistence, the retail data strategy, the metrics layer, fractional data leadership, Oracle Exadata exit, pgvector on RDS, Azure SQL caching, Kafka tiered storage, Snowflake cost and rightsizing, Databricks, Cosmos DB, TiDB, CockroachDB, Vertica, Milvus, Redis, MariaDB, SQL Server Query Store, MongoRocks, PostgreSQL checkpoints, and the CPG and retail industry stacks.
Data strategy decision 1: one engine or several
The first decision is whether the estate standardises on one database engine or accepts several on purpose. The honest answer for most enterprises above a certain size is several, and the record should say which workload shapes justify which engine: relational OLTP, document, key-value, wide-column, columnar analytics, vector, time series, graph.
the case for polyglot persistence across five database types is the archive’s post on making that decision deliberately rather than by accretion, and its measurement is the one the record should carry: for each workload, the p99 latency and the cost per million operations on the engine it is on, against the same figures on the engine it would be consolidated onto.
The rejected option is written down too. Consolidating everything onto PostgreSQL is a defensible strategy for an estate under a certain scale, and the record should say what scale, so that the revisit trigger is a number rather than an argument. tuning MariaDB for cloud and containerised environments and Redis performance troubleshooting: a field guide are two of the archive’s engine-specific posts that show what the operational cost of each additional engine looks like once it is in production.
Data strategy decision 2: managed, self-managed or both
The second decision is the shared-responsibility line: which workloads run on a DBaaS, which are self-managed, and on what grounds. The record’s measurement is the ledger from the DBaaS archive: what the provider takes off the plate, priced against what it takes out of the operator’s hands, per workload.
pgvector on RDS: what CTOs should know is the archive’s worked example of a workload where the managed service’s extension allow-list is the deciding fact, and Azure SQL caching strategies for sub-millisecond reads shows a managed relational service being made to hit a latency objective it does not reach by default.
mastering Azure Cosmos DB performance is the fully-managed end of the spectrum, where the operator’s levers are request units and partition keys and nothing else, and the record should state whether that is acceptable for the workload in question and why.
Data strategy decision 3: the analytical home
The third decision is where analytical queries run: a cloud warehouse, a lakehouse, a real-time columnar engine, or the operational database itself. The measurement is the warehouse bill read line by line, which the data warehousing archive covers, and the archive’s data strategy posts on the platforms are the evidence for each option. reducing Snowflake costs and rightsizing Snowflake are the warehouse option’s cost side; Databricks performance bottlenecks and Databricks repartitioning are the lakehouse option’s; Vertica index usage is the on-premises columnar option’s.
The revisit trigger for this decision is usually latency: a dashboard that needs sub-second refresh on live data is not a warehouse workload, and the record should name the threshold at which a real-time engine (ClickHouse, through ChistaDATA, is the one MinervaDB supports) is evaluated. TiDB performance troubleshooting and profiling queries in CockroachDB are the archive’s posts on the distributed SQL option, where the analytical and transactional homes are the same cluster and the measurement is whether the HTAP replicas keep up.
Data strategy decision 4: the semantic layer and who owns the numbers
The fourth decision is where a metric is defined once. Without it, revenue is computed four ways by four teams and the strategy meeting is spent reconciling them.
a metrics layer with dbt: five tests that keep the numbers honest is the archive’s post on the mechanism, and its five tests are the measurement for this record: uniqueness, not-null, accepted values, referential integrity and freshness, run on every build, with the failure count over time as the health signal. The rejected option is the BI tool’s own semantic layer, and the record should say why it was rejected (portability, usually) so that the decision survives a BI tool change.
-- decision records as data: one table, nine rows, queried at every quarterly review (illustrative)
CREATE TABLE data_strategy_decision (
decision_id smallint NOT NULL,
title text NOT NULL,
chosen text NOT NULL,
rejected text NOT NULL,
measured_by text NOT NULL, -- the metric or system view, named exactly
revisit_when text NOT NULL, -- a number or an event, never "as needed"
decided_on date NOT NULL,
last_reviewed date NOT NULL,
CONSTRAINT pk_data_strategy_decision PRIMARY KEY (decision_id),
CONSTRAINT ck_data_strategy_decision_revisit_named
CHECK (revisit_when !~* 'as needed')
);
-- the review query: every record whose revisit trigger has not been checked this quarter
SELECT decision_id,
title,
measured_by,
revisit_when,
last_reviewed
FROM data_strategy_decision
WHERE date_trunc('quarter', CURRENT_DATE) > last_reviewed
ORDER BY decision_id;
Keeping the records as rows rather than as a document is a small trick with a large effect: the review query cannot be skipped by not opening the document, and a record with “as needed” in its revisit column fails a check constraint the moment someone tries to write it.
Data strategy decision 5: the event backbone
The fifth decision is how data moves between the homes chosen in decisions one to three. The usual answer is a log (Kafka or a compatible service), and the record’s measurement is end-to-end lag from a write in the operational store to its visibility in the analytical home, at the p99, under the peak write rate.
Kafka tiered storage and multi-region replication is the archive’s post on the two properties that turn a log into a strategic asset: retention long enough to replay a consumer from scratch, and replication that survives a region.
The rejected option is point-to-point ETL, and the revisit trigger is the number of consumers, because below three a log is overhead and above ten it is the only thing that scales.
Data strategy decision 6: the exit from the incumbent
The sixth decision is what to do about the engine the enterprise is paying the most for, which in a large estate is usually a commercial relational database on proprietary hardware.
Oracle Exadata cost optimisation: migrating to an open-source stack is the archive’s post on that decision, and the record’s measurement is a rehearsed migration of one representative workload with the replication lag, the query-plan differences and the cutover time recorded. The rejected option is renewal, and it should be priced honestly, including the features the incumbent has that the target does not; the decision is only sound when the rehearsal has produced numbers on both sides.
the SQL Server Query Store FAQ is in this archive for a related reason: the Query Store is the measurement tool for an estate deciding whether its SQL Server workloads stay, move to Azure SQL, or migrate, because it records the plan history that a migration rehearsal has to reproduce.
Data strategy decision 7: the vector and AI workloads
The seventh decision is new and the one most often made without a record. Embeddings need a home, retrieval needs an index, and the choice is between an extension in the relational store, a dedicated vector database, or a feature of the warehouse.
troubleshooting Milvus performance is the archive’s post on the dedicated option’s operational cost (index build time, segment compaction, memory per collection), and the pgvector post above is the extension option’s. The measurement is recall at the latency objective for the top-k the application actually uses, on the production embedding dimension, and the revisit trigger is the collection size at which the chosen index type’s build time exceeds the maintenance window.
Data strategy decision 8: the reliability contract
The eighth decision is the SLO for each home and the recovery objectives behind it, written as numbers: availability, p99 latency, RPO and RTO, per workload tier.
The archive’s operational posts are the evidence that the numbers are achievable on the chosen engine: PostgreSQL checkpoint tuning for I/O spikes and recovery time is the RTO side, optimising MongoRocks for dynamic thread handling is a latency-under-load example on a storage engine few teams tune, and the Kafka post above is the RPO side of the backbone.
The rejected option is a single SLO for everything, and the revisit trigger is the error budget: a tier that has not spent its budget in two quarters is over-provisioned, and one that spends it monthly is under-designed.
Data strategy decision 9: who decides, and how often
The ninth decision is the governance of the other eight. Someone owns the records, the review runs on a calendar, and the measurement columns are read from the systems rather than reported by the teams.
the fractional chief data officer and real-time analytics is the archive’s post on one answer to the ownership question for organisations that do not yet have a full-time one, and from chaos to clarity: a case study of a failed platform is the account of what the absence of this decision costs. The rejected option is a committee, and the revisit trigger is a missed review.
The nine data strategy decisions, side by side
| Decision | Governs | Measured by | Revisit when |
|---|---|---|---|
| 1. One engine or several | Which workload shapes get which engine | p99 and cost per million ops, per engine, per workload | A workload’s cost on its engine exceeds its cost consolidated |
| 2. Managed or self-managed | The shared-responsibility line per workload | The DBaaS ledger: taken off the plate vs out of the hands | An extension, parameter or scaling ceiling blocks a requirement |
| 3. Analytical home | Warehouse, lakehouse, real-time engine, or none | The bill line by line; refresh latency the dashboards need | Sub-second live refresh is required, or the bill outgrows the value |
| 4. Semantic layer | Where each metric is defined once | Test failures per build; number of definitions per metric | Two teams report different figures for one metric |
| 5. Event backbone | How data moves between homes | End-to-end p99 lag at peak write rate; replay time | Consumer count crosses the threshold in either direction |
| 6. Incumbent exit | The most expensive engine’s future | A rehearsed migration with lag, plan diffs and cutover time | Renewal date minus the rehearsal’s measured duration |
| 7. Vector and AI | Where embeddings live and are searched | Recall at the latency objective for the real top-k | Index build time exceeds the maintenance window |
| 8. Reliability contract | SLO, RPO, RTO per tier | Error budget spend; restore and failover drill results | Budget unspent for two quarters or spent every month |
| 9. Governance | Owner, calendar, source of the measurements | Records reviewed on time; measurements read from systems | A review is missed |
The measurements in the third column are named so that they can be read from a system table, a billing view or a drill ledger; a record whose measurement can only be reported by the team that owns the decision is not yet a record.

How a data strategy record gets written
A record is written in a working session with the team that runs the workload, not by the strategy owner alone, and it takes the four fields in a fixed order. Chosen comes first and is usually already known. Rejected is second and is where the session spends its time, because the team has to name the alternative it did not take and say why in terms of a measurement rather than a preference.
Measured by is third, and the rule is that the metric must be one that can be read from a system (a catalog view, a billing export, a drill ledger) by someone who was not in the room.
Revisit when is last, and it is the field that turns a decision into a strategy: a number, a date, or a named event. A record whose revisit field cannot be filled is a decision that was never really made, and the session ends by scheduling the measurement that would fill it.
Data strategy in one industry: the retail and CPG records
The archive’s industry posts show the nine records filled in for one sector. retail data strategy: six proven layers for margin is the layered version of decisions one to five for a retailer, from point-of-sale capture through the event backbone to the metrics that the merchandising team acts on.
dynamic assortment planning with Databricks is decision three and seven together, a lakehouse workload with a model-training component. unlocking growth in the CPG industry is the executive statement of what the records are for: a decision about promotion spend or shelf allocation that is made on last week’s data rather than last quarter’s.
What the industry posts have in common is that every one of them names a measurement, and none of them names a vendor as the strategy. That is the vendor-neutral principle applied to strategy rather than to tooling: MinervaDB supports every engine and platform in this archive and recommends against each of them for the workloads where the record’s measurement says so.
Version notes: the platforms named in the archive change their pricing, their managed-service boundaries and their feature sets with each release, and a record’s “measured by” column should always name the version of the system it was measured on.
Every decision above is rehearsed before it is made: a migration on a copy, a consolidation on a replica, a retention change on a snapshot. Keep a restorable copy of every dataset outside the platform that holds it, because decision six and decision two both eventually depend on being able to leave, and the only guaranteed exit is the copy the organisation already has.
Where this data strategy archive sits
This archive is the decision layer above the rest of minervadb.com. The DBaaS archive is decision two in depth, the data warehousing archive is decision three, the Kafka archive is decision five, the monitoring archive is where decision eight’s measurements come from, and the PostgreSQL archive and InnoDB archive are the operational evidence for decision one. The decision-record form itself is described in the architecture decision records documentation, from which the four fields are adapted.
For a data strategy written as records rather than slides, a platform selection or exit rehearsed with numbers, or fractional data leadership that owns the quarterly review, the MinervaDB database consulting practice starts from these nine records, and states for every recommendation which record it changes and what measurement justified the change.