Valkey High Availability
TruePPM uses Valkey — the BSD-3-Clause, Linux Foundation fork of Redis — for four distinct roles at the same time. A single Valkey outage therefore degrades or disables four subsystems simultaneously. For a production on-prem deployment, running Valkey highly available is effectively mandatory, not optional.
Valkey speaks the Redis wire protocol, so TruePPM’s configuration surface keeps
the Redis names: the connection string is REDIS_URL, the scheme is redis://
(or rediss:// for TLS), and the client libraries are redis-py and
channels-redis. Any Redis-compatible server works. Valkey is what the chart and
Compose files ship.
The dependency surface — one Valkey, four load-bearing roles
Section titled “The dependency surface — one Valkey, four load-bearing roles”Valkey is not a “nice to have” cache you can shed. It is wired into four independent subsystems, each on its own logical database index:
| Role | Database | What uses it | What it does |
|---|---|---|---|
| Celery broker | /0 | Async / background work | Queues every asynchronous job — CPM recalculation drains, MS Project imports, webhook delivery, retention purges, notification email. |
| Django Channels layer | /1 | Real-time collaboration, WebSocket fan-out | Carries live board/schedule updates and presence between API pods. Every connected client depends on it. |
| Django cache backend | /2 | Read-path caching, rate limiting, transient state | Backs cached reads, DRF throttle counters, and short-lived OIDC/OAuth login state (the PKCE verifier and nonce for an in-flight SSO login). |
| Notification throttles | /3 | Mention fan-out limits | Counters bounding notification volume per user. |
Because all four point at the same Valkey instance, its availability is a shared fate: a broker outage is also a Channels outage is also a cache outage. Sizing and hardening Valkey is therefore a production concern on par with the database, not an afterthought.
Which HA topologies TruePPM supports today
Section titled “Which HA topologies TruePPM supports today”This is the part that determines your deployment, so read it before choosing a topology. TruePPM addresses four logical databases, which is what decides whether a given topology can work at all:
| Topology | Supported | Notes |
|---|---|---|
| Replicated Valkey behind one stable endpoint | Yes | Primary with one or more replicas, fronted by a managed service endpoint, Kubernetes Service, or VIP that always resolves to the current primary. The simplest production path. |
| Managed Valkey / Redis-compatible service | Yes | Cluster mode must be disabled — see below. The provider handles failover, patching, and backups. |
| Sentinel | Experimental | Shipped in 0.4 as experimental. Configure it with the TRUEPPM_VALKEY_* settings below; all four databases are wired to follow the primary across a failover with no restart. Not yet verified against a live Sentinel quorum — see the caution below. |
| Cluster mode | No | Not supported, and not planned. A clustered endpoint exposes only database 0, and TruePPM uses four. The Channels layer has no cluster support in any case. |
So the goal is reachable two ways: no single Valkey process whose loss takes down real-time, async, and caching together — via replication behind one endpoint that survives failover, or via Sentinel.
For a production deployment you can only rely on today, choose a replicated primary behind one stable endpoint. That path is the one this project exercises; Sentinel is new and carries the caveat below.
Configuring Sentinel
Section titled “Configuring Sentinel”Sentinel monitors a primary/replica set and promotes a replica when the primary fails. TruePPM resolves the current primary from the Sentinels on every connection, so a failover is designed to need no restart and no config change.
Set these on the API and every Celery worker. A non-empty
TRUEPPM_VALKEY_SENTINELS is what switches Sentinel on; when it is empty (the
default), TruePPM uses REDIS_URL exactly as before.
| Variable | Required | Meaning |
|---|---|---|
TRUEPPM_VALKEY_SENTINELS | Yes | Comma-separated host:port list of the Sentinel nodes. Use three or more — Sentinel needs a quorum to authorize a promotion, so two can never fail over. |
TRUEPPM_VALKEY_MASTER_NAME | Yes | The name the Sentinels monitor the primary under — the first argument of sentinel monitor in sentinel.conf, commonly mymaster. |
TRUEPPM_VALKEY_PASSWORD | No | Password for the data nodes (primary and replicas). |
TRUEPPM_VALKEY_SENTINEL_PASSWORD | No | Password for the Sentinel nodes themselves. Separate on purpose: sentinels commonly carry a different password, or none. |
TRUEPPM_VALKEY_USE_TLS | No | true to use TLS to the data nodes. Default false. |
In Sentinel mode REDIS_URL is ignored. Leave it unset — TruePPM emits a
startup warning (trueppm.valkey.W001) if a stale value is left behind, so it is
never ambiguous which topology is in effect. A missing
TRUEPPM_VALKEY_MASTER_NAME or a malformed sentinel list refuses to boot
(trueppm.valkey.E001 / E002) rather than silently degrading.
With Helm
Section titled “With Helm”Disable the bundled pod and fill in the valkey.sentinel block. The chart routes
both passwords through its connection Secret, so neither is rendered into a
Deployment manifest in plaintext:
valkey: enabled: false sentinel: enabled: true nodes: "sentinel-0.valkey:26379,sentinel-1.valkey:26379,sentinel-2.valkey:26379" masterName: "mymaster" password: "DATA_NODE_PASSWORD" sentinelPassword: "" # leave empty if the sentinels are unauthenticated tls: falseWith sentinel.enabled: true you do not need to supply env.REDIS_URL — the
chart stops requiring it, because there is no single endpoint to name.
Verify a failover in staging before relying on it — this is the step that matters most while support is experimental. The steps below are a manual runbook you run yourself; they are not a substitute for an automated, continuously exercised quorum test, which TruePPM does not yet run (tracked in #2554 and #3404). Completing this runbook once tells you Sentinel works for your topology on that day — it does not make Sentinel support any less experimental for the next reader.
Manual failover verification runbook
Section titled “Manual failover verification runbook”Prerequisites
- A Sentinel quorum of three or more nodes, watching a primary with at least one replica. Two Sentinels can observe a failure but can never authorize a promotion — you need three to actually exercise one.
- A staging environment that mirrors your intended production topology (same network boundaries, same auth/TLS settings you plan to run in production).
- TruePPM’s
TRUEPPM_VALKEY_*settings (or the Helmvalkey.sentinelblock) pointed at that quorum, per Configuring Sentinel above.valkey.enabled: falseif you’re using the chart, since Sentinel replaces the bundled pod rather than fronting it. - A way to watch all four subsystems while the failover happens: a browser tab open on a board (real-time), a way to enqueue and watch a background job (Celery — e.g. trigger a CPM recalculation), an active login session (cache), and a WebSocket client connected to a project (ticket auth on reconnect).
Steps
- Confirm the starting state. Run
SENTINEL master <masterName>against any Sentinel node and note which data node it reports as the current primary. Confirm TruePPM is connected and all four roles are functioning normally (real-time updates land, a queued job drains, you’re logged in, WebSocket is connected). - Stop the primary. Kill the data node process (or its pod/container) that
Sentinel currently reports as primary — do not fail over manually via
SENTINEL failover; the point is to exercise Sentinel’s own failure detection, not to skip it. - Watch Sentinel promote a replica. Poll
SENTINEL master <masterName>on the surviving Sentinels until the reported primary address changes. Note how long the promotion took. - Confirm TruePPM follows without a restart. TruePPM resolves the current
primary from the Sentinels on every connection, so no pod restart or config
change should be required. Check, without restarting anything:
- Real-time (Channels) — updates on the open board tab resume.
- Async (Celery) — the in-flight or a freshly queued job drains.
- Cache — your existing login session stays valid; a fresh login succeeds (this exercises the cache-backed SSO PKCE/nonce state).
- WebSocket reconnect — disconnect and reconnect the WebSocket client; ticket auth succeeds against the new primary.
- Bring the old primary back as a replica (or leave it down, depending on what you’re testing) and confirm the quorum re-stabilizes cleanly.
- Record the result — see the log below — and report it on #2554, whether it passed or found a problem. Real-world results are what move Sentinel from experimental to supported; a silent pass helps nobody but you.
If any of the four roles stays broken after promotion, that is a bug worth reporting on #2554 — do not treat it as expected experimental behavior.
Results log — copy this table into your own runbook and fill in a row each time you exercise a failover:
| Date | Environment | TruePPM version | Sentinel version | Promotion time | Result | Notes |
|---|---|---|---|---|---|---|
| — | e.g. staging, 3 Sentinels + 1 primary + 2 replicas, Kubernetes | — | — | — | pass / fail | — |
Licensing and cost — you do not need a commercial Redis
Section titled “Licensing and cost — you do not need a commercial Redis”Self-hosters reasonably ask whether HA means paying for Redis. It does not. Self-hosting is free either way; the difference is license terms, not money.
- Valkey is BSD-3-Clause under the Linux Foundation, forked from Redis 7.2.4.
It is unambiguously open source, and it is what TruePPM ships
(
valkey/valkey:8-alpine). There is no license conversation to have. - Redis 7.4 and later moved to RSALv2 / SSPLv1 — source-available, not OSI-approved open source. You may still self-host it at no cost; the restriction targets offering it as a managed service to third parties.
- Redis 8.0 and later added AGPLv3 as a third option, making it OSI-open again. Running it alongside TruePPM does not affect TruePPM’s Apache 2.0 licensing — it is a separate process reached over a network protocol, not linked code. But AGPL is blocked outright by many enterprise legal teams, which is a practical obstacle even though it is not a technical one.
Where cost actually appears is managed services. Valkey-based offerings are generally priced below their Redis-OSS equivalents:
| Provider | Valkey option |
|---|---|
| AWS | ElastiCache for Valkey, MemoryDB for Valkey — priced below the ElastiCache for Redis OSS equivalents |
| Google Cloud | Memorystore for Valkey |
| DigitalOcean | Managed Caching (Valkey) |
| Aiven | Aiven for Valkey |
| Azure | No first-party Valkey. Azure Managed Redis is a commercial Redis Enterprise SKU. On Azure, the license-free path is self-hosting Valkey on AKS. |
If you want HA with no vendor bill at all, run a replicated Valkey StatefulSet in your own cluster with a Kubernetes Service fronting the primary, and give it a PersistentVolume so the Celery broker survives a pod restart.
Helm guidance — point TruePPM at external Valkey
Section titled “Helm guidance — point TruePPM at external Valkey”The bundled Valkey subchart is single-node. For production, disable it and point TruePPM at an external, highly available endpoint.
-
Disable the bundled Valkey pod in your values override:
valkey:enabled: false -
Set
REDIS_URLto your external endpoint. When the bundled Valkey is disabled, the chart no longer buildsREDIS_URLfor you, so you must provide it underenv(or via an override):env:# A managed Valkey endpoint, or your own replicated Valkey behind a Service.# Use rediss:// for TLS-terminated managed services.# Cluster mode must be DISABLED — TruePPM needs databases 0, 1, 2, and 3.REDIS_URL: "rediss://:PASSWORD@my-valkey.example.internal:6379"Do not append a database index yourself — TruePPM appends
/0,/1,/2, and/3to this base URL for the four roles. -
Keep the password out of plaintext. As with
DATABASE_URL, prefer sourcingREDIS_URL(or just its password) from a Kubernetes Secret viasecretKeyRefrather than committing it into a values file. -
Verify the endpoint survives failover. The single most important property is that the hostname in
REDIS_URLkeeps resolving to a writable primary after a failover, without a TruePPM restart. Managed services do this for you. If you assemble it yourself, test it by killing the primary and confirming that WebSocket updates and Celery jobs resume on their own — or use Sentinel, which removes the need for a stable endpoint entirely.
All four roles read REDIS_URL, so one correct external endpoint moves all four
onto your HA Valkey at once.
Failure-mode matrix — what happens when Valkey is down
Section titled “Failure-mode matrix — what happens when Valkey is down”If Valkey becomes unavailable, the impact on a Kubernetes deployment is total, not partial. The API does go dark, and the reason is worth understanding precisely, because it is a deliberate design choice rather than an accident:
| Subsystem | Behavior when Valkey is unavailable |
|---|---|
| API / REST reads and writes | Unreachable in the chart’s default topology. /api/v1/readyz performs a set-then-get round-trip against the cache (_probe_cache), and probes.api.readinessPath points at it, so every API pod is marked NotReady, the Service loses its endpoints, and the ingress returns 503. Django itself would happily serve database-backed reads — the process stays up and /api/v1/health/ still answers 200 — but nothing routes to it. |
| Real-time (Channels / WebSockets) | Disrupted. WebSocket clients disconnect and live updates stop. This is the reason readiness gates on the cache at all: the same Valkey carries the Channels layer, so a pod that cannot reach it cannot serve real-time collaboration and should not claim to be ready. |
| Async (Celery broker) | Halted. New tasks cannot be enqueued and queued work is not processed; drains (CPM recalculation, imports, webhooks, notification email) stall until Valkey returns. |
| Cache | Cold on recovery. Throttle counters reset, and in-flight SSO logins fail and must be retried — the PKCE verifier and nonce live in the cache. |
Queued work is delayed, not lost. The Celery queue is not the record of what
needs doing — PostgreSQL is. Fourteen outbox drains run every 30 seconds and
re-dispatch pending and orphaned rows, and every task carries acks_late=True
with reject_on_worker_lost=True. Within about 30 seconds of Valkey returning,
outbox-backed work resumes on its own. What has no outbox row behind it — a
fire-and-forget dispatch whose only record was the queue entry — is the part that
can be lost. See Why losing the broker does not lose the
work.
The bundled Valkey does persist, which is a separate question from
availability. charts/valkey runs valkey-server --appendonly yes against a 2Gi
PVC, so AOF at the default appendfsync everysec puts roughly one second of
appended commands at risk on an unclean stop. docker-compose.prod.yml is the
opposite: /data is a tmpfs with no AOF, so a container restart drops the queue
entirely. The per-artifact
table states all
three shapes.
The takeaway: an outage does not corrupt committed data, and it does not lose outbox-backed work — but on Kubernetes it takes the entire application offline, not just its real-time and async layers. For production, that is why HA Valkey is treated as mandatory.
Related pages
Section titled “Related pages”- Deployment Sizing — resource guidance for the API, worker, and cache tiers.
- Durability & Redundancy — what is authoritative, what is reconstructible, and the step-up ladder that ends at an HA broker.
- Beat Liveness — how a dead Beat process is detected.
- Troubleshooting — diagnosing a live 503 and telling a cache outage from a database one.
- System Health — the in-app health surface for operators.