Skip to content

Valkey High Availability

TruePPM uses Valkey — the BSD-3-Clause, Linux Foundation fork of Redis — for four distinct roles at the same time. A single Valkey outage therefore degrades or disables four subsystems simultaneously. For a production on-prem deployment, running Valkey highly available is effectively mandatory, not optional.

Valkey speaks the Redis wire protocol, so TruePPM’s configuration surface keeps the Redis names: the connection string is REDIS_URL, the scheme is redis:// (or rediss:// for TLS), and the client libraries are redis-py and channels-redis. Any Redis-compatible server works. Valkey is what the chart and Compose files ship.

The dependency surface — one Valkey, four load-bearing roles

Section titled “The dependency surface — one Valkey, four load-bearing roles”

Valkey is not a “nice to have” cache you can shed. It is wired into four independent subsystems, each on its own logical database index:

RoleDatabaseWhat uses itWhat it does
Celery broker/0Async / background workQueues every asynchronous job — CPM recalculation drains, MS Project imports, webhook delivery, retention purges, notification email.
Django Channels layer/1Real-time collaboration, WebSocket fan-outCarries live board/schedule updates and presence between API pods. Every connected client depends on it.
Django cache backend/2Read-path caching, rate limiting, transient stateBacks cached reads, DRF throttle counters, and short-lived OIDC/OAuth login state (the PKCE verifier and nonce for an in-flight SSO login).
Notification throttles/3Mention fan-out limitsCounters bounding notification volume per user.

Because all four point at the same Valkey instance, its availability is a shared fate: a broker outage is also a Channels outage is also a cache outage. Sizing and hardening Valkey is therefore a production concern on par with the database, not an afterthought.

Which HA topologies TruePPM supports today

Section titled “Which HA topologies TruePPM supports today”

This is the part that determines your deployment, so read it before choosing a topology. TruePPM addresses four logical databases, which is what decides whether a given topology can work at all:

TopologySupportedNotes
Replicated Valkey behind one stable endpointYesPrimary with one or more replicas, fronted by a managed service endpoint, Kubernetes Service, or VIP that always resolves to the current primary. The simplest production path.
Managed Valkey / Redis-compatible serviceYesCluster mode must be disabled — see below. The provider handles failover, patching, and backups.
SentinelExperimentalShipped in 0.4 as experimental. Configure it with the TRUEPPM_VALKEY_* settings below; all four databases are wired to follow the primary across a failover with no restart. Not yet verified against a live Sentinel quorum — see the caution below.
Cluster modeNoNot supported, and not planned. A clustered endpoint exposes only database 0, and TruePPM uses four. The Channels layer has no cluster support in any case.

So the goal is reachable two ways: no single Valkey process whose loss takes down real-time, async, and caching together — via replication behind one endpoint that survives failover, or via Sentinel.

For a production deployment you can only rely on today, choose a replicated primary behind one stable endpoint. That path is the one this project exercises; Sentinel is new and carries the caveat below.

Sentinel monitors a primary/replica set and promotes a replica when the primary fails. TruePPM resolves the current primary from the Sentinels on every connection, so a failover is designed to need no restart and no config change.

Set these on the API and every Celery worker. A non-empty TRUEPPM_VALKEY_SENTINELS is what switches Sentinel on; when it is empty (the default), TruePPM uses REDIS_URL exactly as before.

VariableRequiredMeaning
TRUEPPM_VALKEY_SENTINELSYesComma-separated host:port list of the Sentinel nodes. Use three or more — Sentinel needs a quorum to authorize a promotion, so two can never fail over.
TRUEPPM_VALKEY_MASTER_NAMEYesThe name the Sentinels monitor the primary under — the first argument of sentinel monitor in sentinel.conf, commonly mymaster.
TRUEPPM_VALKEY_PASSWORDNoPassword for the data nodes (primary and replicas).
TRUEPPM_VALKEY_SENTINEL_PASSWORDNoPassword for the Sentinel nodes themselves. Separate on purpose: sentinels commonly carry a different password, or none.
TRUEPPM_VALKEY_USE_TLSNotrue to use TLS to the data nodes. Default false.

In Sentinel mode REDIS_URL is ignored. Leave it unset — TruePPM emits a startup warning (trueppm.valkey.W001) if a stale value is left behind, so it is never ambiguous which topology is in effect. A missing TRUEPPM_VALKEY_MASTER_NAME or a malformed sentinel list refuses to boot (trueppm.valkey.E001 / E002) rather than silently degrading.

Disable the bundled pod and fill in the valkey.sentinel block. The chart routes both passwords through its connection Secret, so neither is rendered into a Deployment manifest in plaintext:

valkey:
enabled: false
sentinel:
enabled: true
nodes: "sentinel-0.valkey:26379,sentinel-1.valkey:26379,sentinel-2.valkey:26379"
masterName: "mymaster"
password: "DATA_NODE_PASSWORD"
sentinelPassword: "" # leave empty if the sentinels are unauthenticated
tls: false

With sentinel.enabled: true you do not need to supply env.REDIS_URL — the chart stops requiring it, because there is no single endpoint to name.

Verify a failover in staging before relying on it — this is the step that matters most while support is experimental. The steps below are a manual runbook you run yourself; they are not a substitute for an automated, continuously exercised quorum test, which TruePPM does not yet run (tracked in #2554 and #3404). Completing this runbook once tells you Sentinel works for your topology on that day — it does not make Sentinel support any less experimental for the next reader.

Prerequisites

  • A Sentinel quorum of three or more nodes, watching a primary with at least one replica. Two Sentinels can observe a failure but can never authorize a promotion — you need three to actually exercise one.
  • A staging environment that mirrors your intended production topology (same network boundaries, same auth/TLS settings you plan to run in production).
  • TruePPM’s TRUEPPM_VALKEY_* settings (or the Helm valkey.sentinel block) pointed at that quorum, per Configuring Sentinel above. valkey.enabled: false if you’re using the chart, since Sentinel replaces the bundled pod rather than fronting it.
  • A way to watch all four subsystems while the failover happens: a browser tab open on a board (real-time), a way to enqueue and watch a background job (Celery — e.g. trigger a CPM recalculation), an active login session (cache), and a WebSocket client connected to a project (ticket auth on reconnect).

Steps

  1. Confirm the starting state. Run SENTINEL master <masterName> against any Sentinel node and note which data node it reports as the current primary. Confirm TruePPM is connected and all four roles are functioning normally (real-time updates land, a queued job drains, you’re logged in, WebSocket is connected).
  2. Stop the primary. Kill the data node process (or its pod/container) that Sentinel currently reports as primary — do not fail over manually via SENTINEL failover; the point is to exercise Sentinel’s own failure detection, not to skip it.
  3. Watch Sentinel promote a replica. Poll SENTINEL master <masterName> on the surviving Sentinels until the reported primary address changes. Note how long the promotion took.
  4. Confirm TruePPM follows without a restart. TruePPM resolves the current primary from the Sentinels on every connection, so no pod restart or config change should be required. Check, without restarting anything:
    • Real-time (Channels) — updates on the open board tab resume.
    • Async (Celery) — the in-flight or a freshly queued job drains.
    • Cache — your existing login session stays valid; a fresh login succeeds (this exercises the cache-backed SSO PKCE/nonce state).
    • WebSocket reconnect — disconnect and reconnect the WebSocket client; ticket auth succeeds against the new primary.
  5. Bring the old primary back as a replica (or leave it down, depending on what you’re testing) and confirm the quorum re-stabilizes cleanly.
  6. Record the result — see the log below — and report it on #2554, whether it passed or found a problem. Real-world results are what move Sentinel from experimental to supported; a silent pass helps nobody but you.

If any of the four roles stays broken after promotion, that is a bug worth reporting on #2554 — do not treat it as expected experimental behavior.

Results log — copy this table into your own runbook and fill in a row each time you exercise a failover:

DateEnvironmentTruePPM versionSentinel versionPromotion timeResultNotes
—e.g. staging, 3 Sentinels + 1 primary + 2 replicas, Kubernetes———pass / fail—

Licensing and cost — you do not need a commercial Redis

Section titled “Licensing and cost — you do not need a commercial Redis”

Self-hosters reasonably ask whether HA means paying for Redis. It does not. Self-hosting is free either way; the difference is license terms, not money.

  • Valkey is BSD-3-Clause under the Linux Foundation, forked from Redis 7.2.4. It is unambiguously open source, and it is what TruePPM ships (valkey/valkey:8-alpine). There is no license conversation to have.
  • Redis 7.4 and later moved to RSALv2 / SSPLv1 — source-available, not OSI-approved open source. You may still self-host it at no cost; the restriction targets offering it as a managed service to third parties.
  • Redis 8.0 and later added AGPLv3 as a third option, making it OSI-open again. Running it alongside TruePPM does not affect TruePPM’s Apache 2.0 licensing — it is a separate process reached over a network protocol, not linked code. But AGPL is blocked outright by many enterprise legal teams, which is a practical obstacle even though it is not a technical one.

Where cost actually appears is managed services. Valkey-based offerings are generally priced below their Redis-OSS equivalents:

ProviderValkey option
AWSElastiCache for Valkey, MemoryDB for Valkey — priced below the ElastiCache for Redis OSS equivalents
Google CloudMemorystore for Valkey
DigitalOceanManaged Caching (Valkey)
AivenAiven for Valkey
AzureNo first-party Valkey. Azure Managed Redis is a commercial Redis Enterprise SKU. On Azure, the license-free path is self-hosting Valkey on AKS.

If you want HA with no vendor bill at all, run a replicated Valkey StatefulSet in your own cluster with a Kubernetes Service fronting the primary, and give it a PersistentVolume so the Celery broker survives a pod restart.

Helm guidance — point TruePPM at external Valkey

Section titled “Helm guidance — point TruePPM at external Valkey”

The bundled Valkey subchart is single-node. For production, disable it and point TruePPM at an external, highly available endpoint.

  1. Disable the bundled Valkey pod in your values override:

    valkey:
    enabled: false
  2. Set REDIS_URL to your external endpoint. When the bundled Valkey is disabled, the chart no longer builds REDIS_URL for you, so you must provide it under env (or via an override):

    env:
    # A managed Valkey endpoint, or your own replicated Valkey behind a Service.
    # Use rediss:// for TLS-terminated managed services.
    # Cluster mode must be DISABLED — TruePPM needs databases 0, 1, 2, and 3.
    REDIS_URL: "rediss://:PASSWORD@my-valkey.example.internal:6379"

    Do not append a database index yourself — TruePPM appends /0, /1, /2, and /3 to this base URL for the four roles.

  3. Keep the password out of plaintext. As with DATABASE_URL, prefer sourcing REDIS_URL (or just its password) from a Kubernetes Secret via secretKeyRef rather than committing it into a values file.

  4. Verify the endpoint survives failover. The single most important property is that the hostname in REDIS_URL keeps resolving to a writable primary after a failover, without a TruePPM restart. Managed services do this for you. If you assemble it yourself, test it by killing the primary and confirming that WebSocket updates and Celery jobs resume on their own — or use Sentinel, which removes the need for a stable endpoint entirely.

All four roles read REDIS_URL, so one correct external endpoint moves all four onto your HA Valkey at once.

Failure-mode matrix — what happens when Valkey is down

Section titled “Failure-mode matrix — what happens when Valkey is down”

If Valkey becomes unavailable, the impact on a Kubernetes deployment is total, not partial. The API does go dark, and the reason is worth understanding precisely, because it is a deliberate design choice rather than an accident:

SubsystemBehavior when Valkey is unavailable
API / REST reads and writesUnreachable in the chart’s default topology. /api/v1/readyz performs a set-then-get round-trip against the cache (_probe_cache), and probes.api.readinessPath points at it, so every API pod is marked NotReady, the Service loses its endpoints, and the ingress returns 503. Django itself would happily serve database-backed reads — the process stays up and /api/v1/health/ still answers 200 — but nothing routes to it.
Real-time (Channels / WebSockets)Disrupted. WebSocket clients disconnect and live updates stop. This is the reason readiness gates on the cache at all: the same Valkey carries the Channels layer, so a pod that cannot reach it cannot serve real-time collaboration and should not claim to be ready.
Async (Celery broker)Halted. New tasks cannot be enqueued and queued work is not processed; drains (CPM recalculation, imports, webhooks, notification email) stall until Valkey returns.
CacheCold on recovery. Throttle counters reset, and in-flight SSO logins fail and must be retried — the PKCE verifier and nonce live in the cache.

Queued work is delayed, not lost. The Celery queue is not the record of what needs doing — PostgreSQL is. Fourteen outbox drains run every 30 seconds and re-dispatch pending and orphaned rows, and every task carries acks_late=True with reject_on_worker_lost=True. Within about 30 seconds of Valkey returning, outbox-backed work resumes on its own. What has no outbox row behind it — a fire-and-forget dispatch whose only record was the queue entry — is the part that can be lost. See Why losing the broker does not lose the work.

The bundled Valkey does persist, which is a separate question from availability. charts/valkey runs valkey-server --appendonly yes against a 2Gi PVC, so AOF at the default appendfsync everysec puts roughly one second of appended commands at risk on an unclean stop. docker-compose.prod.yml is the opposite: /data is a tmpfs with no AOF, so a container restart drops the queue entirely. The per-artifact table states all three shapes.

The takeaway: an outage does not corrupt committed data, and it does not lose outbox-backed work — but on Kubernetes it takes the entire application offline, not just its real-time and async layers. For production, that is why HA Valkey is treated as mandatory.

  • Deployment Sizing — resource guidance for the API, worker, and cache tiers.
  • Durability & Redundancy — what is authoritative, what is reconstructible, and the step-up ladder that ends at an HA broker.
  • Beat Liveness — how a dead Beat process is detected.
  • Troubleshooting — diagnosing a live 503 and telling a cache outage from a database one.
  • System Health — the in-app health surface for operators.