AI & Data Engineering · Vector Data Engineering on PostgreSQL, Milvus, ClickHouse, Redis, Valkey, MariaDB, MongoDB and every major cloud
Vector Data Engineering for Production RAG, Search and Recommendation
Vector Data Engineering at MinervaDB is the discipline of turning text, images, events and rows into embeddings that can be searched at production scale, with the same rigour we apply to any other database workload: measured recall, measured latency, a rollback path for every change and named engineers on call around the clock. We build on the database you already run wherever it holds up under measurement, and we move you to a dedicated vector store only when the numbers say so.
The discipline
What Vector Data Engineering covers, and what a vector database alone does not
A vector database is one component. Vector Data Engineering is the pipeline around it: the chunking that decides what a vector represents, the embedding model that decides what “similar” means, the index that decides how much of the true answer you get back, the filters that keep tenants apart, and the evaluation harness that tells you when any of those degrades. Most failed RAG projects we are called into had a working vector index and no Vector Data Engineering owner for the other four.
The work begins at the source of truth. Product catalogues, support tickets, contracts, clinical notes and telemetry already live in PostgreSQL, MySQL, SQL Server, MongoDB or a lake, and the embeddings derived from them are just that: derived. Our Vector Data Engineering pipelines treat vectors as a rebuildable projection of operational data, keyed idempotently on the document, the chunk and the embedding model version, so a model change, a chunking change or a corrupted index is a rebuild, never a data-loss event.
Between source and index sit the decisions that determine quality. Chunk size and overlap change recall more than any index parameter does. Layout-aware parsing of PDFs and tables changes what a chunk even contains. The embedding model fixes the dimensionality, the distance metric and the memory bill for the life of the index. Each of these is a Vector Data Engineering decision made against a golden query set for your corpus, not from a vendor tutorial.
On the serving side, Vector Data Engineering means a retrieval path that combines approximate nearest-neighbour search with lexical matching and metadata filters, fuses the candidates, reranks the top of the list and hands the result to an application, whether that is a retrieval-augmented generation assistant, a recommendation service or an anomaly detector on transaction streams. Every stage has a latency budget and a quality metric, and both are measured in production, not estimated in a notebook.
The evaluation loop closes the system. Recall at k against an exact scan, nDCG on the golden set, faithfulness of generated answers, p95 latency per stage, index lag versus the source and embedding drift after a model change are the signals that tell us whether to change the chunker, the model, the index parameters or the engine. That loop is what separates Vector Data Engineering from a proof of concept that happened to work on the demo corpus.
Figure 1. The Vector Data Engineering reference pipeline. Sources stay authoritative; embeddings are versioned, derived data; retrieval combines dense, lexical and filtered candidates; the evaluation loop feeds back into chunking, model and index choices.
Services
Vector Data Engineering services
Eight fixed-scope services. Most engagements begin with the readiness assessment and continue as the findings dictate; the 24×7 support retainer keeps the same engineers on the platform once it is live.
Vector Data Engineering readiness assessment
Two to three weeks against your actual corpus and query log. We measure the candidate engines you already run (pgvector on PostgreSQL, Milvus, ClickHouse, Redis or Valkey, MariaDB, MongoDB) for recall at 10 and p95 latency on a golden set, size memory for the embedding dimension and vector count, and hand back a ranked recommendation with the numbers behind it. The output tells you whether a dedicated vector store is needed at all.
Embedding and chunking pipeline design
The core of Vector Data Engineering delivery: layout-aware extraction, chunking strategy selected by measured recall, metadata schema for tenant, ACL and time filters, in-VPC or managed embedding serving with batching and back-pressure, and a model registry that pins model version and dimension on every vector. Delivered as code you own, with idempotent upserts and replayable CDC.
Index engineering and tuning
The Vector Data Engineering work most teams skip. HNSW, IVF, DiskANN-class and quantised indexes chosen and tuned per engine: M, ef_construction, ef_search, nlist, nprobe, half-precision and binary storage, two-stage rerank. Every parameter change is plotted as recall against latency on your golden set before it reaches production.
Hybrid retrieval and RAG serving
Dense plus BM25 or learned-sparse retrieval, Reciprocal Rank Fusion, cross-encoder reranking, metadata pre-filtering that does not collapse recall, semantic caching in Redis or Valkey, citation tracking and guardrails. Built to a per-stage latency budget with telemetry wired before go-live.
Migration between vector platforms
Vector Data Engineering migrations move Pinecone or another managed service into pgvector, Milvus or ClickHouse; pgvector into Milvus at scale; Elasticsearch or OpenSearch kNN into a consolidated store. Shadow index, parity check, quality gate, alias switch, retire, with the old platform kept warm until the new one has a week of clean metrics.
High availability and disaster recovery
Vector Data Engineering for production means Patroni and pgBackRest for PostgreSQL with pgvector, distributed Milvus on Kubernetes with etcd and object-storage replication, Redis or Valkey Cluster with replicas, ClickHouse ReplicatedMergeTree with Keeper. RPO and RTO stated in numbers and proven in quarterly drills.
Security, governance and multi-tenancy
Tenant isolation as a filter inside the engine or as a partition, row-level security on PostgreSQL, RBAC in Milvus and Redis, PII masking before embedding, embedding-inversion risk review, audit logging and the evidence packs for GDPR, DPDP, HIPAA and SOC 2 reviews of AI systems.
24×7 Vector Data Engineering support
Named engineers on the pipeline and the vector store with S1 acknowledgement in 15 minutes, monthly recall and latency reviews, capacity forecasting, upgrade planning for pgvector, Milvus and ClickHouse releases, and runbooks for every failure mode we have seen.
Index internals
Where recall, latency and memory are decided
Approximate nearest-neighbour search trades a little recall for a lot of speed, and every engine exposes that trade through a handful of parameters. Vector Data Engineering that treats those parameters as defaults ships an index that is either slower or less accurate than it needs to be, and usually both.
HNSW: a navigable graph
Hierarchical Navigable Small World indexes keep every vector in a bottom layer with M links to near neighbours, and a shrinking set of vectors in upper layers that act as express lanes. A query enters at the top, greedily descends and then widens its beam at layer zero to ef_search candidates. Higher M and ef_construction build a better graph at the cost of build time and memory; higher ef_search raises recall and latency together. This is the index in pgvector, Milvus, ClickHouse, Redis, Valkey and MariaDB, under slightly different names, and it is where most Vector Data Engineering tuning time goes.
The trap is filtered search. When a tenant or date filter removes most of the graph, a naive HNSW traversal runs out of matching candidates before it reaches k results and returns fewer, worse rows. pgvector 0.8.0 added iterative index scans that keep walking until k rows pass the filter; Milvus pushes the filter into the segment scan; Qdrant and others use filter-aware graph traversal. Which mechanism your engine has, and how it behaves below about 2% selectivity, is the first thing our Vector Data Engineering assessment measures.
IVF: partition then probe
Inverted-file indexes cluster the space into nlist cells around trained centroids and search only the nprobe closest cells. They build faster and use less memory than HNSW, and pair naturally with product or scalar quantisation, which is why Milvus IVF_PQ and the newer IVF_RABITQ are the memory-economic choice at hundreds of millions of vectors. The cost is that the centroids go stale as data shifts: a re-embedding, a new tenant with a different vocabulary, or a large backfill all warrant a retrain, and our Vector Data Engineering runbooks schedule it from a recall alert rather than a calendar.
Quantisation and two-stage search
A 1,536-dimension float32 vector is 6,144 bytes before any graph overhead. Half precision halves it (pgvector halfvec, Milvus FLOAT16_VECTOR), scalar int8 quarters it, and binary or RaBitQ representations bring it to under 200 bytes. Recall drops with each step unless the quantised index is used only to generate candidates and a full-precision rerank of the top 100 to 200 restores the ordering. Milvus reports IVF_RABITQ with SQ8 refinement at roughly 95% recall on a quarter of the memory; that is a vendor figure, and we reproduce it on your corpus before it enters a sizing decision.
Reading the plan as Vector Data Engineering does
On PostgreSQL the question is always whether the planner used the vector index at all. An ORDER BY embedding <=> $1 LIMIT 10 with a selective WHERE clause can flip to a sequential scan with an exact sort, which is correct and slow. We read it the same way we read any other plan.
-- pgvector 0.8.x on PostgreSQL 16+: confirm the HNSW index is used and
-- see how many candidates the iterative scan had to walk.
SET hnsw.ef_search = 120;
SET hnsw.iterative_scan = relaxed_order;
EXPLAIN (ANALYZE, BUFFERS)
SELECT chunk_id,
doc_id,
embedding <=> ${QUERY_VECTOR} AS distance
FROM doc_chunks
WHERE tenant_id = ${TENANT_ID}
ORDER BY embedding <=> ${QUERY_VECTOR}
LIMIT 10;
-- Expect: Index Scan using idx_doc_chunks_embedding_hnsw on doc_chunks
-- Not: Seq Scan ... Sort ... (exact but unbounded at scale)
Figure 2. The three index families and the parameters that govern them. Panel C shows bytes per 1,536-dimension vector at each precision; sizing figures are illustrative and are measured per engine before any capacity decision.
Ingestion
Ingestion pipelines and re-embedding without an outage
The ingestion path is where Vector Data Engineering either earns its name or becomes a nightly batch script nobody wants to touch. We design it for three properties: idempotency, replayability and a model change that is a routine operation rather than a migration project.
In Vector Data Engineering, idempotency comes from the key. Every chunk is addressed as doc_id + chunk_no + model_version, so the same chunk re-processed from a retry, a replay or a backfill is an upsert rather than a duplicate, and a delete on the source tombstones every chunk of that document. Content hashes let the pipeline skip chunks whose text has not changed, which is most of them on a daily run.
Replayability, the second Vector Data Engineering property, comes from change data capture. Debezium on PostgreSQL logical replication, MySQL binlog or MongoDB change streams publishes inserts, updates and deletes to Kafka with the log position attached. A stream processor in Flink, Kafka Streams or plain workers chunks, dedupes and batches 64 to 256 chunks per embedding call, and the consumer offset is the recovery point. Index lag, measured as source updated_at against vector indexed_at, is a service-level indicator with an alert.
Embedding serving runs in your VPC on vLLM, Text Embeddings Inference or Triton when data governance requires it, or against a managed endpoint when it does not; either way it is batched, rate-limited and retried with back-pressure to the stream, so a slow model never becomes a lost event.
Re-embedding is the operation most teams postpone because it looks like downtime. Under our Vector Data Engineering pattern a new model, or a new dimension, is a second consumer on the same stream from the same offset, writing into a second collection or table built in parallel. The application reads through an alias. Promotion is gated on parity (chunk counts per tenant within 0.1%), quality (recall at 10 and nDCG on the golden set at least equal to the live index) and latency (p95 at production concurrency within budget, with ef_search or nprobe retuned for the new dimension).
The alias moves in one metadata operation, the old index stays warm for a week, and rollback is pointing the alias back.
The same mechanism handles a chunking change, a corrupted index, an engine migration and a major-version upgrade of the vector store itself. Nothing is rewritten in place, which is the property that lets us commit to change windows measured in minutes.
Figure 3. Ingestion with change data capture and the shadow-index cutover used for re-embedding, chunking changes and engine migrations. Promotion is a data decision gated on parity, quality and latency; rollback at every step is an alias move.
Retrieval
Hybrid retrieval: dense, lexical and filtered, then reranked
Pure dense retrieval misses exact terms, part numbers, error codes and names; pure lexical retrieval misses paraphrase. Production Vector Data Engineering fuses both, applies filters inside the engine and spends its remaining latency budget on a reranker, because the reranker is where answer quality is most often won.
Three candidate lanes in a Vector Data Engineering retrieval path
The dense lane runs the ANN query with a candidate depth (top 200 is typical) larger than the final k. The lexical lane runs BM25 on a PostgreSQL tsvector, Milvus Sparse-BM25, Redis FT.SEARCH or a learned-sparse model such as SPLADE where the vocabulary is technical. The filter lane is not a lane at all: tenant, ACL, date and document type are predicates pushed into the engine so that ranking never sees a row the user may not read. Post-filtering a top-k list is the single most common cause of both empty results and data leakage we find in review.
Fusion and reranking
Reciprocal Rank Fusion with the conventional constant of 60 merges the two ranked lists without score calibration, which is why Vector Data Engineering at MinervaDB prefers it to weighted sums that drift when a model changes. A cross-encoder or late-interaction reranker then scores the fused top 50 against the query and returns the final 5 to 10 chunks with their citations. On CPU that stage costs 30 to 60 milliseconds at batch size 50; on a small GPU it is under 15. Both figures are planning budgets we confirm on your hardware.
Semantic caching in the Vector Data Engineering stack
A Redis or Valkey cache keyed on the query embedding, with a similarity threshold around 0.95 and a TTL set from the freshness of the underlying corpus, removes a large share of retrieval and generation cost on workloads with repeated intent, such as support assistants. The metrics that matter are hit ratio and false-hit rate, the fraction of cached answers a reviewer judges wrong for the new query, and we tune the threshold from the second, not the first.
Engine-native hybrid search
Some of this fusion now happens inside the store: Milvus hybrid search since 2.5, Redis FT.HYBRID since 8.4, MongoDB $scoreFusion since 8.3, ClickHouse with a pre-filter strategy on its vector index. When the engine can fuse in one round trip we use it and measure whether recall holds; when it cannot, the fusion lives in the retrieval service. Either way the telemetry below is wired before the first user query.
Figure 4. The hybrid retrieval serving path with an illustrative latency budget for a 150 ms p95 target, and the per-stage telemetry we wire before go-live. Figures are planning budgets, not measurements.
Capacity and tuning
Capacity, build performance and query tuning by engine
Vector indexes are memory-resident structures with expensive builds. Sizing them wrong is the second most common reason we are engaged, after recall problems, and the two are related: an index that does not fit in memory is tuned downward until it fits, and recall goes with it.
Sizing the index: Vector Data Engineering arithmetic
The estimate starts from vector count times bytes per vector at the chosen precision, plus graph overhead of roughly M × 2 × 4 bytes per vector for HNSW at layer zero, plus the metadata columns that filters need. Ten million 1,536-dimension float32 vectors is about 61 GB of vectors alone; at half precision it is 31 GB, and at binary with full-precision rerank the index is under 2 GB with the rerank served from disk. We produce this table for your corpus at three precisions and three engines, then load a representative sample and measure, because page layout, segment sealing and index metadata differ enough between engines that arithmetic alone is not a sizing.
Build performance in Vector Data Engineering
HNSW builds are CPU-bound and, on PostgreSQL, memory-bound: pgvector builds the graph in maintenance_work_mem and spills to disk when it does not fit, at which point build time grows by an order of magnitude. Parallel builds with max_parallel_maintenance_workers help on PostgreSQL 15 and later; 0.8.2 fixed a buffer overflow in the parallel path, so we pin that release. On Milvus the build runs on data nodes in 2.6 and its throughput sets your re-embedding window. On ClickHouse the index is built per part at merge time, so index_granularity and merge settings determine both build cost and query pruning.
Query-time knobs, per engine
| Engine | Recall dial | Where to measure |
|---|---|---|
| pgvector | hnsw.ef_search, ivfflat.probes, hnsw.iterative_scan |
EXPLAIN (ANALYZE, BUFFERS), pg_stat_statements, pg_stat_user_indexes |
| Milvus | ef, nprobe, search-level consistency_level |
query-node metrics, search_latency_bucket, segment load state |
| ClickHouse | hnsw_candidate_list_size_for_search, vector_search_filter_strategy |
system.query_log, system.data_skipping_indices, EXPLAIN indexes = 1 |
| Redis / Valkey | EF_RUNTIME, HYBRID_POLICY, BATCH_SIZE |
FT.PROFILE, FT.INFO, INFO memory |
| MariaDB | mhnsw_ef_search, mhnsw_max_cache_size |
EXPLAIN showing the vector index, performance_schema |
Each dial is swept against the golden set in every Vector Data Engineering tuning cycle and plotted as recall against p95, and the operating point is chosen from that curve with headroom for growth. The sweep is repeated after any re-embedding, because the curve moves with the model.
-- Ground truth for recall@10: exact scan on a 1,000-query sample.
-- Run once per model version; store the result as the golden set.
SET LOCAL enable_indexscan = off;
SELECT q.query_id,
ARRAY_AGG(c.chunk_id ORDER BY c.embedding <=> q.embedding) AS exact_top10
FROM eval_queries AS q
CROSS JOIN LATERAL (
SELECT chunk_id, embedding
FROM doc_chunks
WHERE tenant_id = q.tenant_id
ORDER BY embedding <=> q.embedding
LIMIT 10
) AS c
GROUP BY q.query_id;
Operations
High availability, disaster recovery and security for vector workloads
A vector store that serves a customer-facing assistant is a production database and is run as one. Vector Data Engineering at MinervaDB inherits the same HA, DR and security practice we apply to PostgreSQL, ClickHouse and Redis estates, with the additions that vector workloads need.
PostgreSQL with pgvector
A Patroni cluster with etcd consensus, PgBouncer in transaction pooling and a read service that spreads ANN queries across replicas. HNSW indexes are WAL-logged, so replicas receive the graph rather than rebuilding it, at the cost of significant WAL volume during builds that the archive and the replication slots must be sized for. Synchronous replication with remote_apply gives an RPO of zero for vector upserts where the application needs read-your-writes across nodes. pgBackRest to object storage with monthly restore tests; the index restores with the data.
Milvus distributed
Milvus 2.6 on Kubernetes with two proxies, an active-standby MixCoord, streaming nodes owning the write path over the Woodpecker write-ahead log on object storage (removing the Kafka or Pulsar tier that earlier releases required), query nodes in replica groups with two replicas per serving collection, and etcd in a three-member quorum. Segments, indexes and the WAL live in a versioned bucket with cross-region replication; milvus-backup covers metadata. DDL pauses during coordinator failover, which is planned into the runbook.
DR targets and the Vector Data Engineering rebuild path
Because vectors are derived data, DR has a second recovery path that ordinary databases lack: re-embed from the source of truth. We state both in the runbook: replication-based recovery with an RPO of five minutes or better and an RTO of thirty minutes measured in the drill, and a full rebuild whose RTO is set by embedding throughput, so that a lost index is an incident with a known duration rather than a crisis. Both paths are exercised quarterly and the timings go in the revision log.
Security and governance in Vector Data Engineering
Tenant isolation is enforced inside the engine as a filter predicate or a partition, never in the application layer alone, with row-level security on PostgreSQL and RBAC in Milvus and Redis. PII is detected and masked before embedding, because embeddings of sensitive text can be partially inverted and an index is not a redaction. Access is logged, model endpoints stay inside the VPC where regulation requires it, and the design produces the evidence GDPR, DPDP, HIPAA and SOC 2 reviewers ask for about AI systems: what data went in, which model, who can query it and how deletion propagates to every chunk.
Figure 5. Reference HA topologies for PostgreSQL with pgvector and for distributed Milvus 2.6, with the DR region, targets and the re-embedding rebuild path that applies to both.
Evaluation
Evaluation and observability: how we know it works
Every Vector Data Engineering deliverable ships with a golden query set, an evaluation harness and dashboards, because a retrieval system without measurement degrades silently as the corpus, the users and the models change under it.
Retrieval metrics
The first Vector Data Engineering SLIs: recall at k against an exact scan (the ceiling for everything downstream), mean reciprocal rank and nDCG on a golden set of 500 to 2,000 queries with graded relevance, per tenant where tenants differ. Recomputed on every index parameter change and every model version.
Generation metrics
Faithfulness and groundedness of answers to the retrieved chunks, citation precision, refusal correctness on out-of-scope questions, and a weekly sample rated by people who know the domain. Model-graded scores are calibrated against those human ratings before they drive decisions.
Operational metrics
Vector Data Engineering operations track p50, p95 and p99 per stage; QPS; index lag against the source; embedding drift after model changes; cache hit and false-hit rates; memory and build time per index; cost per thousand queries. Alerts on recall, lag and p95, with runbooks behind each.
Illustrative figures on this page (latency budgets, memory per vector, recall at a given precision) are planning values. We never quote a recall or a latency we have not measured on the client’s corpus, and we never repeat a vendor benchmark as a fact.
Engagement
How a Vector Data Engineering engagement runs
Three phases with a decision point at the end of each, so you can stop after the assessment with a usable answer or continue into build and operations with the same engineers.
| Phase | Duration | What we do | What you receive |
|---|---|---|---|
| 1 · Assess | 2–3 weeks | Corpus and query-log analysis, golden set construction, engine measurement on your data (recall at 10, p95, memory), chunking and model experiments, security and governance review | Ranked engine recommendation with measurements, sizing at three precisions, pipeline design, risk register, fixed-price proposal for build |
| 2 · Build | 4–10 weeks | Ingestion and CDC pipeline, embedding service, index engineering, hybrid retrieval service, evaluation harness, HA and DR topology, runbooks, load and failover tests | Production platform in your accounts and repositories, dashboards and alerts, tested runbooks, handover sessions with your engineers |
| 3 · Operate | ongoing | 24×7 support with S1 in 15 minutes, monthly recall and latency reviews, capacity forecasting, model and engine upgrade planning, quarterly DR drills | Named engineers, monthly report with measured SLIs, upgrade and re-embedding windows executed under change control |
Every Vector Data Engineering deliverable is code in your repositories and infrastructure in your cloud accounts. There is no MinervaDB runtime in the path and no dependency on us to keep it running; the support retainer exists because most teams prefer to have the engineers who built it on call.
Pricing
Vector Data Engineering pricing
Transparent Vector Data Engineering rates, invoiced against logged hours or a fixed-price scope agreed after the assessment. Retainers include the 24×7 severity matrix below.
| Severity | Definition | Acknowledgement |
|---|---|---|
| S1 | Retrieval down or returning wrong tenants’ data; index unavailable; pipeline stopped with customer impact | 15 minutes, 24×7 |
| S2 | Recall or p95 outside SLO; index lag beyond target; failover degraded | 12 hours |
| S3 | Non-urgent defects, tuning requests, capacity questions | 24 hours |
| S4 | Advisory, roadmap and upgrade planning | 48 hours |
FAQ
Vector Data Engineering: frequently asked questions
The questions we are asked most often before a Vector Data Engineering engagement starts.
Do we need a dedicated vector database, or is pgvector enough?
For most estates under a few tens of millions of vectors, pgvector on PostgreSQL 16 or later, with HNSW and iterative index scans from pgvector 0.8, meets recall and latency targets while keeping vectors beside the data they describe, with joins, row-level security and one backup regime. We recommend a dedicated store such as Milvus when the measured triggers appear: recall collapsing under selective filters, memory economics at hundreds of millions of vectors, learned-sparse or late-interaction retrieval, or tenant counts that a single table cannot serve. The Vector Data Engineering readiness assessment measures those triggers on your data.
How do you choose chunk size and the embedding model?
Against a golden query set built from your real queries and rated by your domain experts. We run chunking strategies (fixed windows with overlap, heading-aware, parent-child, table-aware) and candidate models through the same harness and pick by recall at 10 and nDCG, then confirm the memory and latency cost of the winning dimension. Model choice is pinned in a registry so every vector records how it was produced; that registry is a standard Vector Data Engineering artefact.
What happens when we want to change the embedding model later?
A second consumer on the same change stream writes into a shadow index built in parallel; the application reads through an alias. We promote only when parity, quality and latency gates pass, keep the old index warm for a week, and roll back by moving the alias. No vectors are rewritten in place and there is no downtime; this is the Vector Data Engineering pattern shown in Figure 3.
How do you keep tenants and permissions separate in a shared index?
Tenant and ACL predicates are pushed into the engine as filters or partitions so ranking never sees rows a user may not read; on PostgreSQL that is row-level security, on Milvus it is partition keys and RBAC. We then measure recall at the filter selectivity your tenants actually have, because naive HNSW traversal degrades under selective filters and the fix (iterative scans, partitioned indexes) depends on the engine.
Can you run this entirely inside our VPC for GDPR, DPDP or HIPAA?
Yes. Embedding models and language models run on vLLM, Text Embeddings Inference or Triton in your accounts, the vector store and pipeline are in your Kubernetes or database estate, and no data leaves the boundary. Vector Data Engineering under those regimes produces the evidence reviewers ask for: data lineage into the index, model versions, access logs and deletion propagation to every chunk.
Which cloud services do you support for vector workloads?
Amazon RDS and Aurora PostgreSQL with pgvector, Amazon ElastiCache for Valkey with built-in search, Azure Database for PostgreSQL, Google Cloud SQL and AlloyDB with ScaNN, MongoDB Atlas Vector Search, ClickHouse Cloud, Zilliz Cloud for Milvus and Redis Cloud. Each managed service diverges from the open-source engine in version lag and configuration limits; we document the divergence and the exit path as part of the Vector Data Engineering design.
What does the 24×7 retainer cover?
Vector Data Engineering support means named engineers on the pipeline and the vector store with S1 acknowledgement in 15 minutes, S2 in 12 hours, S3 in 24 hours and S4 in 48 hours; monthly recall, latency and capacity reviews; upgrade planning for pgvector, Milvus, ClickHouse and Redis or Valkey releases; and quarterly DR drills. Retainers start at US $4,500 per quarter.
How is Vector Data Engineering priced?
US $300 per hour for remote work and US $500 per hour on site plus travel, or a fixed price for the assessment and build phases quoted after a scoping call. The assessment is designed to stand on its own: you can stop after it with a measured recommendation and use it with any team.
Talk to a Principal Architect about your vector workload
Vector Data Engineering starts with what you have: bring the corpus, the query log and the engines you already run. We will tell you what to measure first, and whether you need a new database at all.