Redis support · Valkey support · 24×7 consultative support and remote DBA · Sentinel and Cluster HA · Memory and eviction engineering · Persistence and recovery · ElastiCache, MemoryDB, Memorystore and Azure Managed Redis · Redis to Valkey migration

Redis Support Measured in p99 Latency, Fragmentation Ratio, Replication Offset and Restore Drills, Not in Ticket Counts

MinervaDB Redis support is run by senior engineers who read INFO, SLOWLOG, LATENCY HISTORY and MEMORY STATS before they touch a parameter. The service covers 24×7 monitoring and incident response under a published severity matrix for Redis Open Source and Valkey, keyspace and data-structure engineering, memory and eviction tuning, Sentinel and Cluster high availability, persistence and backup with scheduled restore drills, upgrades across the Redis 7.x to 8.x and Valkey 8.x to 9.x lines, the Redis-or-Valkey licensing decision, and operations on ElastiCache, MemoryDB, Memorystore and Azure Managed Redis. Every change ships as a reviewed runbook with a rollback path.

15 minS1 acknowledgement for every Redis support customer, 24×7×365
900+enterprises supported across every major engine
46cities with on-site delivery presence
15+years of in-memory data store operations
200+years of combined leadership experience

01 · Why MinervaDB for Redis support

Redis fails at the speed it succeeds: one blocking command, one unbounded key or one fork stall takes the whole instance with it

A single command thread means every millisecond of latency is shared by every client. Redis support from MinervaDB is delivered by engineers who read the keyspace, the command mix and the allocator before they read the ticket, on Redis Open Source, Valkey and every managed offering of the two, as part of the wider NoSQL DBA support practice that also covers MongoDB and Cassandra.

Engine-internals depth

The engineers on watch have tuned buffer pools, storage engines and replication on PostgreSQL, MySQL and ClickHouse as well as Redis. Redis support from MinervaDB reads a fragmentation ratio, a slowlog entry or a MOVED storm the way an application engineer reads a stack trace, and fixes the cause rather than the symptom.

Real 24×7 with a published matrix

Follow-the-sun pods with S1 acknowledged within 15 minutes, S2 in 12 hours, S3 in 24 hours and S4 in 48 hours. Every alert rule maps to a runbook, every S1 closes with a written root cause analysis that names the INFO field behind it and the prevention step. 24×7 consultative support →

Vendor-neutral by principle

MinervaDB resells no Redis Cloud, no Redis Enterprise and no cloud capacity. The Redis support recommendation can be Valkey under BSD instead of Redis 8.x under its tri-licence, Sentinel instead of Cluster, no persistence at all for a pure cache, or a workload that belongs in PostgreSQL or ClickHouse rather than in memory.

Measured before and after

Nothing is reported as improved until INFO says so: p99 per command family from LATENCY HISTORY and client probes, used_memory and mem_fragmentation_ratio, evictions and hit ratio, replication offset lag, restore duration per drill. Gains are estimates until the next measurement confirms them.

02 · Redis support operating model

From INFO fields and slowlog entries to an accountable engineer, with a response target on every page

Instrumentation is layered: the INFO sections, SLOWLOG, the latency monitor, MEMORY STATS, client and replication state are exported by redis_exporter, correlated with operating system counters, evaluated by recording rules and routed to the Redis support engineer on watch. Keyspace sampling uses SCAN on a replica, never KEYS on a primary.

Redis support monitoring and escalation pipeline: INFO, SLOWLOG, LATENCY and MEMORY signals, collectors, Prometheus, Alertmanager and the on-call engineer with the S1 to S4 severity matrix

Figure 1. The Redis support monitoring and escalation pipeline: signals, collectors, Prometheus, Alertmanager and the on-call engineer, with the severity matrix and representative alert conditions.

Alerts are actionable by design. Thresholds are set per cluster from its own baseline rather than from a template, so a cache that evicts by design does not page on evicted_keys and a session store that should never evict does. Latency is alerted on at the 99th percentile from client-side probes rather than on server averages, and every rule carries the runbook the Redis support engineer executes when it fires.

The monitoring stack integrates with what you already run: Percona Monitoring and Management, Prometheus and Grafana, Datadog or CloudWatch. An instance is not accepted into 24×7 Redis support until a backup has been restored on an isolated host and a Sentinel or cluster failover has been rehearsed. Access is through named ACL users with least privilege, time-bounded and audited; no shared default user, no FLUSHALL without a confirmation gate.

# Instance health as the on-call Redis support engineer
# reads it: memory, fragmentation, persistence, replication
redis-cli -h ${REDIS_HOST} -p ${REDIS_PORT} \
  --user ${REDIS_USER} --pass ${REDIS_PASSWORD} \
  INFO memory persistence replication stats \
  | grep -E '^(used_memory_human|maxmemory_human|\
mem_fragmentation_ratio|evicted_keys|expired_keys|\
rdb_last_bgsave_status|aof_last_write_status|\
latest_fork_usec|master_link_status|master_repl_offset|\
slave[0-9]+|instantaneous_ops_per_sec|rejected_connections)'

# The slowest commands since the last reset, oldest first
redis-cli ... SLOWLOG GET 25

# Latency events the latency monitor has recorded
redis-cli ... CONFIG SET latency-monitor-threshold 10
redis-cli ... LATENCY LATEST

# Biggest keys by type, sampled with SCAN on a replica
redis-cli -h ${REDIS_REPLICA_HOST} ... --bigkeys -i 0.01

03 · Redis support for execution and memory

One command thread, one allocator and one eviction policy decide latency, and each has an INFO field that says whether it is right

Redis executes every command on a single thread and Valkey has moved further work off it; memory is governed by encodings, jemalloc and maxmemory; persistence costs a fork. Redis support treats these three as a single budget and tunes them from INFO, OBJECT ENCODING and MEMORY USAGE rather than from folklore.

Redis support view of the Redis and Valkey execution and memory model: command thread and I/O threads, encodings and fragmentation, persistence costs and the INFO field that validates each parameter

Figure 2. The Redis and Valkey execution and memory model as operated by Redis support: the command thread and I/O threads, keyspace encodings and fragmentation, persistence costs, and the INFO field that validates each parameter.

Parameter Starting envelope for a dedicated instance What Redis support monitors before changing it
maxmemory / maxmemory-policy Set on every instance, leaving headroom for the replication backlog, client buffers, fork copy-on-write and fragmentation; allkeys-lru for a cache, noeviction for a store; CONFIG SET online used_memory and used_memory_peak against maxmemory, evicted_keys rate, keyspace hit ratio, whether TTLs exist for a volatile policy
io-threads / io-threads-do-reads; Valkey I/O threading Enabled where the main thread is network-bound on a multi-core host; restart Main-thread CPU against I/O thread CPU, instantaneous_ops_per_sec, client count and pipelining depth
hash-max-listpack-entries and the other listpack thresholds Defaults kept unless sampled keys show large objects just over the threshold; online OBJECT ENCODING and MEMORY USAGE on sampled keys, used_memory per key count
appendonly / appendfsync / auto-aof-rewrite-percentage Off for a pure cache; everysec for sessions, queues and rate limits; always only where the recovery point demands it; online aof_delayed_fsync, latest_fork_usec, disk write latency during rewrite, aof_last_write_status
save points Snapshot from a replica on schedule rather than from the primary on write count; online rdb_last_bgsave_time_sec, rdb_changes_since_last_save, copy-on-write memory during BGSAVE
repl-backlog-size / client-output-buffer-limit replica Backlog sized to the longest tolerated replica disconnect at the measured write rate; online master_repl_offset minus replica offsets, partial versus full resync counts, replica disconnects in the log
activedefrag and thresholds On where mem_fragmentation_ratio stays above 1.5; online allocator_frag_ratio, active_defrag_running, main-thread CPU cost while defragmenting
timeout / tcp-keepalive / maxclients Idle timeout for abandoned connections, keepalive on, maxclients from the file descriptor limit; online connected_clients, rejected_connections, blocked_clients, CLIENT LIST output buffer sizes
Operating system: THP, vm.overcommit_memory, somaxconn Transparent huge pages off, overcommit set to 1, backlog raised; restart of the host settings Fork duration in latest_fork_usec, latency spikes correlated with page faults, WARNING lines at startup

The envelope above is a starting point Redis support derives from, not a template to copy; final values come from the measured workload, and every change is applied on a non-production instance first with online-versus-restart stated in the runbook. Exact Redis or Valkey versions are confirmed before any version-sensitive setting is proposed.

04 · Redis support for Sentinel and Cluster high availability

Sentinel while one primary holds the dataset, Cluster when it cannot, and a failover that is rehearsed rather than assumed

Replication with a Sentinel quorum, Redis or Valkey Cluster with 16384 hash slots, and cross-region replicas or managed multi-AZ services. Redis consulting from MinervaDB chooses from dataset size, write throughput and key access pattern, and Redis support operates the result with a quarterly drill.

Redis support high availability topologies: replication with Sentinel, Redis or Valkey Cluster with hash slots, cross-region and managed variants, selection criteria and the quarterly failover drill

Figure 3. Redis and Valkey high availability topologies operated by Redis support: replication with Sentinel, Cluster with hash slots and redirection, cross-region and managed variants, with the selection criteria and the quarterly failover drill.

Topology Failover model Data-loss and recovery characteristics When Redis support recommends it
Replication with Sentinel Three or more Sentinels elect and promote; clients rediscover the primary Asynchronous replication, so writes not yet shipped can be lost; min-replicas-to-write bounds loss during a partition; promotion in seconds Datasets that fit one primary’s memory and write capacity; multi-key commands and Lua across keys; the default below the cluster threshold
Redis or Valkey Cluster Replicas of a failed primary elect among themselves; cluster-aware clients follow MOVED Same asynchronous semantics per shard; slots without a primary stop writes unless full coverage is relaxed; resharding online Datasets or write rates beyond one primary; independent keys or disciplined hash tags; Valkey 9.x for atomic slot migration and per-slot metrics
Cross-region replicas Manual or scripted promotion of the remote replica Asynchronous by nature; recovery point is the replication lag at failure; used for reads and disaster recovery, not active-active Regional DR and read locality; active-active needs a CRDT product such as Redis Enterprise or an application-level design
Managed: ElastiCache, MemoryDB, Memorystore, Azure Managed Redis Provider-managed Provider-published behaviour; MemoryDB adds a durable multi-AZ transaction log; parameter groups, node types, replica counts and backup design still owned by Redis support Teams optimising for operational simplicity, with the version and feature surface of each service checked against the workload

Sentinel and Cluster are configured from the Redis Sentinel documentation and the Valkey cluster tutorial, with down-after-milliseconds, quorum and cluster-node-timeout tuned so a network partition cannot produce two primaries. Primary loss, election, client reconnect and replica re-sync are drilled every quarter with a recorded time.

05 · Redis support for persistence, backup and recovery

A backup that has not been restored is a hope, and an in-memory store that persists by accident is a latency problem

RDB snapshots by fork, the append-only file with an fsync policy and rewrite, replica-driven backups shipped off host, and a scheduled restore drill. Redis support decides the persistence mode per workload, because a pure cache and a session store have opposite requirements.

Redis support persistence, backup and recovery: the durability path, RDB and AOF cost per mode, the restore path and the scheduled restore drill cadence

Figure 4. Redis persistence, backup and recovery as rehearsed by Redis support: the durability path, what each persistence mode costs, the restore path and the drill cadence.

Persistence chosen per workload

A pure cache runs with persistence off and relies on replicas for availability; sessions, queues, rate limits and Streams run AOF with appendfsync everysec plus periodic RDB for fast restart; the rare store that cannot lose a second of writes runs always or moves to MemoryDB-class durability. Redis support writes the choice into the runbook with the loss it accepts.

Forks kept off the primary

BGSAVE and AOF rewrite fork the process and pay copy-on-write memory and time that grow with the dataset. Backups are taken from a replica on schedule, transparent huge pages are off so the fork does not stall, and latest_fork_usec is alerted on so a growing dataset does not quietly turn a snapshot into a latency incident.

The drill is scheduled

RDB and AOF files are shipped encrypted to S3, GCS or Blob and retained in generations and time; the drill loads them on an isolated host, runs redis-check-rdb and redis-check-aof, compares key counts by type with the source, and records the restore duration and achieved recovery point in the runbook for review with you. The mechanics follow the Redis persistence documentation.

06 · Redis support for versions, licensing and the Redis-or-Valkey decision

Redis 8.x is tri-licensed, Valkey is BSD, and the right answer depends on the features in use and the terms your legal team will accept

Redis 8.10 is current with 8.2 as the long-term line supported to September 2030; Redis 7.4 and 7.2 remain under RSALv2 or SSPLv1 to December 2029; 6.2 is the last BSD line, to April 2027. Valkey 9.1 is current under BSD-3 with 9.0, 8.1, 8.0 and 7.2 in maintenance. Redis support tracks all of them and makes the choice from evidence.

Redis support release lifecycle and licensing in September 2026: Redis 8.x tri-licence with 8.2 long-term, Redis 7.x and 6.2, Valkey 7.2 to 9.1 under BSD, cloud service mapping and the Redis or Valkey decision

Figure 5. Redis and Valkey release lifecycle, licensing and cloud service mapping as tracked by Redis support in September 2026, with the Redis-or-Valkey decision and migration mechanics.

Stay on Redis 8.x when the features earn it

Redis 8 folded JSON, search and query, vector similarity, time series and the probabilistic types into core, and 8.2 carries them on a long-term line. Estates that use those types, that run Redis Cloud or Enterprise features such as active-active, or whose legal team has accepted the RSALv2, SSPLv1 or AGPLv3 terms, stay; Redis support keeps them on a supported minor and plans the 7.x exits before December 2029.

Move to Valkey when the licence or the cloud decides

Valkey is protocol-compatible with Redis 7.2, carries async I/O threading, per-slot metrics and atomic slot migration in 9.x, and ships search, JSON and bloom as BSD modules. Where the 7.2-compatible surface is enough and a BSD licence is required, or where the cloud provider has standardised on Valkey engines, Redis support migrates by replication where versions allow or by RDB transfer with a short write freeze. valkey.io →

Upgrades within a defined window

Minor releases are tracked per line and applied replica-first within a defined window; major moves, 7.x to 8.x or Valkey 8.x to 9.x, are rehearsed on a clone with the captured command mix replayed and memory and p99 compared with the source before clients move. Managed services are upgraded on the provider’s calendar, with the parameter group and client behaviour checked before each maintenance window. Redis release notes →

07 · Redis performance engineering method

Baseline, attribute, one reversible change, re-measure; the same loop for a slowlog entry and for a fragmented allocator

Redis performance work is about what the single thread is asked to do and how much memory it is asked to hold. Redis support finds the keys, commands and settings responsible from the instance’s own telemetry, changes one thing, and reports the p99 as measured.

Redis support performance engineering method: baseline, attribute, change, validate, the findings that recur across health checks and the engagement lifecycle

Figure 6. The Redis performance engineering method used by Redis support, the findings that recur across health checks, and the engagement lifecycle from takeover assessment to ongoing operations.

Keyspace and data structures

Unbounded hashes, sets and lists are split or capped; keys get TTLs where the policy assumes them; hash tags are designed before a cluster move rather than after the first CROSSSLOT error; encodings are checked with OBJECT ENCODING so that a structure just over a listpack threshold does not cost ten times its data. Redis performance is designed into the keyspace before it is tuned in the configuration.

Commands and round trips

KEYS, whole-structure reads and large synchronous deletes are replaced with SCAN, ranged reads and UNLINK; pipelining, Lua scripts and Functions cut round trips on hot paths; client-side caching with CLIENT TRACKING removes reads entirely where the data allows it; Streams consumer groups replace list-based queues where acknowledgement matters.

Memory and the allocator

maxmemory is set with headroom for buffers, backlog and fork copy-on-write; fragmentation above 1.5 is defragmented or the instance is restarted from a replica in a window; client output buffers are bounded; the eviction policy matches whether TTLs exist. Every recommendation names the INFO field it moves and is verified on the next scrape.

08 · Redis support scope, environments and engagement

The same NoSQL DBA support runbooks on bare metal, managed cloud services and Kubernetes; the same access model everywhere

On managed services the work shifts from operating system and persistence mechanics to parameter groups, node right-sizing, replica and multi-AZ design, cost control and the keyspace and command-level engineering no provider does for you. Redis support says plainly when a workload does not belong on a managed service, and when it does not belong in memory at all.

On-premises and bare metal

Transparent huge pages off, vm.overcommit_memory set, NUMA placement, network backlog raised, TLS on client and replication connections, named ACL users, a hardened configuration under version control, and security aligned to CIS controls before an instance enters Redis support coverage. Data security →

ElastiCache, MemoryDB, Memorystore, Azure Managed Redis

Engine choice between Redis OSS and Valkey where offered, parameter groups tuned from CloudWatch or Cloud Monitoring and INFO, node types and replica counts from p95 utilisation, backup retention and cross-region design drilled, monthly cost per million operations reported alongside p99. Cloud database optimization and FinOps →

Kubernetes

Sentinel or Cluster with pod anti-affinity across zones, memory requests equal to limits so the kernel never reclaims from Redis, storage classes with predictable fsync semantics where persistence is on, PodDisruptionBudgets that never evict a quorum together, and backups to object storage on the same restore-drill schedule as everywhere else.

Redis support service line What it covers technically
Monitoring and alerting 15-second scrape with redis_exporter, recording rules per cluster, latency probes from the client side, integration with PMM, Prometheus and Grafana, Datadog or CloudWatch
Incident response Severity-based paging, runbook execution, root cause analysis naming the INFO field or slowlog entry behind it, prevention ticket, customer communication on a shared channel
Capacity planning and sizing Memory from sampled key sizes and growth, headroom for buffers and forks, ops per second against main-thread CPU, network bandwidth for replication and client traffic, benchmark with memtier_benchmark on the captured command mix
Redis performance engineering Keyspace and data-structure review, slowlog and latency attribution, command and round-trip optimisation, allocator and eviction tuning, verified p99 before and after
High availability operations Sentinel quorum and timing, Cluster slot layout and resharding, cross-zone replica placement, quarterly failover drills, client library configuration review
Persistence and backup assurance Persistence mode per workload, replica-driven RDB and AOF backups shipped off host, retention in generations and time, scheduled restore drills with a recorded attestation
Upgrades and migrations Minor upgrades within a defined window, 7.x to 8.x and Valkey 8.x to 9.x rehearsed on a clone, Redis to Valkey moves, Sentinel to Cluster moves, self-managed to managed and back
Security and compliance TLS everywhere, ACL users with least privilege, renamed or disabled dangerous commands, protected mode and bind addresses, audit through connection logs and command tracing, evidence for SOC 2, PCI DSS and HIPAA
Change control and knowledge transfer Reviewed runbooks, ticketed changes with an approver, verified rollback paths, monthly health report, documentation and joint reviews

Engagement models

Redis support is delivered as a 24×7 consultative support and remote DBA retainer, as Redis consulting projects with a fixed scope such as a health check, a Cluster migration, a Redis to Valkey move or a cloud move, and as emergency response for estates outside a subscription. Retainers include the takeover assessment, the runbook, monitoring integration, a monthly health report with the measured metrics behind every recommendation, and a quarterly restore and failover drill. Scope and pricing are set from the assessment. Emergency database support →

Access model

VPN or bastion reachability, named ACL users per engineer with least privilege rather than the default user, TLS on every connection, SSH certificate or key authentication with MFA to hosts, and full session and change auditing with every production modification tied to a ticket and an approver. FLUSHALL, FLUSHDB, DEBUG and CONFIG REWRITE sit behind a confirmation gate in every runbook.

Standing caveat: every recommendation on this page is tested on a non-production instance before it is applied to production, changes are staged and reversible by design, and a verified backup and DR posture is confirmed before any persistence, topology, eviction or upgrade change is made.

09 · FAQ

Redis support questions we are asked most

Short answers to what engineering and platform leaders ask before the first call.

What does MinervaDB Redis support include?

24×7 monitoring and incident response for Redis Open Source and Valkey under a published severity matrix with S1 acknowledged within 15 minutes, S2 in 12 hours, S3 in 24 hours and S4 in 48 hours; keyspace and data-structure engineering; memory, eviction and allocator tuning; Sentinel and Cluster operations with quarterly failover drills; persistence and backup with scheduled restore drills; upgrades and Redis to Valkey migrations; security hardening with ACLs and TLS; and operations on ElastiCache, MemoryDB, Memorystore and Azure Managed Redis. Every recommendation names the INFO field behind it and every change carries a rollback path.

Which Redis and Valkey versions do you support?

Redis Open Source 8.x including the 8.2 long-term line, Redis 7.4 and 7.2, and Redis 6.2 until its security support ends; Valkey 9.1, 9.0, 8.1, 8.0 and 7.2; and the managed offerings of both on AWS, Google Cloud and Azure. Estates on releases past their support window get an upgrade plan as the first deliverable. Exact versions are confirmed before any version-sensitive guidance is given.

Should we run Redis or Valkey?

It depends on the features in use and the licence terms your organisation accepts. Redis 8.x carries JSON, search, vector and time series in core under the RSALv2, SSPLv1 or AGPLv3 tri-licence; Valkey is BSD-3, protocol-compatible with Redis 7.2, and adds async I/O threading, per-slot metrics and atomic slot migration in 9.x with search, JSON and bloom as modules. MinervaDB resells neither and makes the call from your command mix, feature inventory and licensing position, then migrates by replication or RDB transfer with the source retained for rollback.

How do you ensure Redis high availability?

Replication with a Sentinel quorum while one primary holds the dataset, Redis or Valkey Cluster when memory or write throughput needs more than one primary, cross-zone replica placement, min-replicas-to-write where loss during a partition must be bounded, and cluster-aware or Sentinel-aware client configuration. Primary loss, election, client reconnect and replica re-sync are drilled every quarter with a recorded time, and the timings that prevent split brain are set from the network, not from defaults.

Do you provide 24×7 Redis support on ElastiCache, MemoryDB or Memorystore?

Yes. On managed services the work shifts to engine choice, parameter groups, node right-sizing, replica and multi-AZ design, backup retention, maintenance-window planning and the keyspace and command-level engineering the provider does not do. The same severity matrix and the same engineers apply, and MinervaDB says plainly when a workload does not belong on a managed service.

Can you take over an estate that already has Sentinel, Cluster or a managed service in place?

Yes. Onboarding starts with a takeover assessment that documents the topology, configuration deltas from default, persistence and backup state, keyspace profile and security posture, and produces the runbook we then operate from. Nothing is changed until the assessment is reviewed with you, and the first restore drill and failover drill happen before the estate enters 24×7 coverage.

How is Redis support priced?

As a 24×7 consultative support and remote DBA retainer scoped from the takeover assessment, as fixed-scope Redis consulting projects such as a health check, a Cluster migration or a Redis to Valkey move, or as emergency response for estates outside a subscription. Retainers include the assessment, the runbook, monitoring integration, a monthly health report and a quarterly restore and failover drill.

Talk to a senior Redis support engineer

Bring INFO ALL from every node, SLOWLOG GET 50, the output of redis-cli --bigkeys from a replica, your configuration, the persistence and backup arrangement and the date of your last restore drill to the first call. We will tell you what the next incident will be and what we would change first.