Upgrading
This page walks through moving an existing TruePPM instance to a newer version, safely, for whichever way you run it — Docker Compose, a single server, or Helm/Kubernetes. If you are standing up TruePPM for the first time rather than upgrading one, see Installation instead.
Before you upgrade
Section titled “Before you upgrade”-
Read the changelog for the target version — check
CHANGELOG.mdor the release notes for breaking changes and migration notes. -
Back up PostgreSQL with
scripts/backup.sh. Valkey state is ephemeral (broker + cache); PostgreSQL is the only stateful service. The exact invocation differs by how you run TruePPM — a bare./scripts/backup.sh --output-dir ./backupsonly works whenDATABASE_URLis already exported andpg_dumpis on the host’sPATH, which is true for neither the production Compose stack (no host database port, nopg_dumpin the application images) nor a fresh shell against Helm. See Backup & Restore for the command for your stack — Compose (development), Compose (production), or Kubernetes/Helm.Use this rather than a hand-rolled
pg_dump > file.sql. The runbook’s restore tooling reads--format=customarchives, so a plain-SQL dump is an artifactscripts/restore.shcannot consume — you would discover that during a rollback, which is the worst moment to discover it. The custom format also preservesCREATE EXTENSIONordering, which theltree/pg_trgm/btree_gistobjects depend on; see Backup & restore. -
Note your current version before starting.
Terminal window docker inspect registry.gitlab.com/trueppm/trueppm/api:latest --format '{{.Config.Labels}}'# Or: helm list -n trueppm
Per-release operational change notes
Section titled “Per-release operational change notes”Every release carries a short operational change note answering one question: what does an operator have to check or change before and after this upgrade? This is distinct from the changelog (user-facing changes) — it is the operator’s pre-flight. Each release’s note appears in a versioned section on this page (see Upgrading to 0.3 below for the shape); the template used to write one is:
## Upgrading to <version>
**Migration behavior:** <additive-only / includes destructive ops / data backfill>.Downtime: <none beyond the migrate run / brief write pause / maintenance window>.
**New or changed env vars / Helm values:**- `NEW_VAR` — <what it does, default, whether action is required>- `changed.helm.value` — <old → new default, action required?>
**Breaking config:** <none / describe what an existing config must change>.
**New migrations operators will see:**- `<app>.<NNNN_name>` — <one line: what schema it adds/changes>
**Pre-upgrade action:** <back up (always) / rotate a credential / set a new value>.**Post-upgrade verification:** <what "green" looks like — see the checklist below>.**Rollback notes:** <forward-only? safe to roll back? migration-reversibility caveat>.Fill this in from the release’s changelog fragments, the diff of
packages/helm/values.yaml, and the new files under
packages/api/**/migrations/. Even an all-additive release gets a note so the
operator has a complete picture rather than inferring “nothing changed.”
Upgrading to 0.4
Section titled “Upgrading to 0.4”Migration behavior: includes destructive ops (see below). Downtime: a
maintenance window sized to your task count, because projects.0148 blocks all
reads and writes on the task table while it builds a constraint (see the
projects.0148 caution below), plus the rest of the migrate run and the
transient-500 rollout windows described per-migration below on multi-replica
installs.
0.4 (tagged v0.4.0-beta.1 on 2026-09-15) carries five migrations with a
RemoveField, DeleteModel, or raw DROP TABLE — four of them are destructive
in a way an operator needs to plan around, not just a schema-hygiene detail. If
you are upgrading from 0.3, read all four before you run migrate, in
particular the workshop-data warning:
Known transient-500 windows on multi-replica installs. Three of the five
migrations drop a column or table outright, with no intervening
null=True-then-remove deprecation release, so an old pod’s code still names
what the new schema no longer has. If you run the API at replicaCount >= 2
(the posture values-prod.yaml sets by
default, and the posture 0.4’s basic-HA carve-out encourages), Kubernetes’
default RollingUpdate keeps an old pod serving while the new pod’s migration
applies — for the duration of that rollout, each of the following reads or
writes will get a transient 500 from any request landing on an old pod:
profiles.0008_remove_userprofile_schedule_in_deliver— drops the retiredschedule_in_delivercolumn outright (ADR-0942 §3, #3137). Affects the old pod’sget_profile_prefs()read, used by/auth/me/(the web shell’s bootstrap call). See #3782 for the full analysis.webhooks.0009_encrypt_webhook_secret_and_failure_counters— encryptsWebhook.secretinto a newsecret_ciphertextcolumn and drops the plaintextsecretcolumn in the same migration (#2885). On the old pod,webhooks/models.pystill declaressecretas a concrete, non-deferredCharField, so Django selects it on everyWebhookquery — every webhook read and every outbound delivery attempt from an old pod fails withProgrammingError: column webhooks_webhook.secret does not existfor the duration of the window. This silently stops the entire outbound-webhook subsystem with no operator-facing warning; if you depend on webhook deliveries during the rollout, scale to a single replica for the upgrade (see the mitigation below).projects.0123_remove_historicalproject_agile_features_and_more— dropsProject.agile_features/HistoricalProject.agile_featuresoutright (#2025), with no deprecation window and no migration docstring.ProjectandHistoricalProjectreads happen on nearly every authenticated request path, so this has the broadest blast radius of the three — but no data-loss concern: the field is re-exposed as a derived, read-onlySerializerMethodField(ProjectSerializer.get_agile_features), so this is dead-column removal, not a feature regression. What is missing is purely the HA-safety disclosure this entry now provides.
This is a known, bounded, and accepted trade-off for each of the three, not a
regression to report: no crash-loop, no data loss beyond what each entry
states, and every window closes as soon as the rollout finishes.
Single-replica installs never see any of this. If you are on
replicaCount >= 2 and want to avoid the windows entirely, scale to one
replica (kubectl scale deployment/trueppm-api -n trueppm --replicas=1)
before the upgrade and scale back up once the rollout completes.
Already running v0.4.0-beta.1 or later? These migrations already ran —
there is nothing to re-run and no remediation for the schema changes
themselves. The one exception is workshop data: if you upgraded from 0.3
without exporting it first, it is already gone (see the danger box above) —
this section exists so the next self-hoster upgrading from 0.3 does not lose
theirs the same way.
New migrations operators will see:
profiles.0008_remove_userprofile_schedule_in_deliver— drops the retiredschedule_in_deliverboolean preference. See above.webhooks.0009_encrypt_webhook_secret_and_failure_counters— encrypts webhook signing secrets and adds failure-tracking columns. See above.projects.0149_drop_workshops_tables— drops the two tables backing the removed Board Workshop Mode. Destructive, irreversible. See above.projects.0123_remove_historicalproject_agile_features_and_more— drops the deadagile_featurescolumn (now derived). See above.sso.0002_remove_oidcprovider_workspace_ssoproviderpolicy_and_more— migrates anyOIDCProvider/OIDCIdentityrows toallauth’sSocialApp/SocialAccount(ADR-0517), then drops the bespokeOIDCProviderandOIDCIdentitymodels. Benign: SSO never shipped in a tagged 0.3 release, so no installation has rows to migrate or lose — no disclosure needed beyond this line.
Breaking API changes: 0.4 also removes several /api/v1/ paths that
appeared in the published v0.3.0-alpha.3 schema. These are documented as
deliberate deprecation-window exceptions in API Stability & Deprecation
Policy — check there for the full
reasoning and migration guidance for each. If you have an integration calling
/api/v1/projects/{id}/workshop/*, /api/v1/tasks/{id}/scope/,
/api/v1/teams/{id}/, or /api/v1/projects/{id}/history/summary/, read that
section before upgrading.
Pre-upgrade action: back up PostgreSQL (always — see Before you
upgrade); if you have workshop data, export or preserve
it before running migrate (see the danger box above); if you cannot tolerate
a transient-500 window on webhook deliveries or /auth/me//project reads
during rollout, scale to one replica first.
Post-upgrade verification: see the checklist below.
Rollback notes: all five migrations are destructive or table-dropping —
per migration reversibility, a
clean rollback means restoring the pre-upgrade backup, not a migrate reverse.
projects.0149 in particular cannot be reversed by any means other than
restore — its reverse_sql is a no-op by design.
Helm: the api and web pods get a default-deny ingress NetworkPolicy
Section titled “Helm: the api and web pods get a default-deny ingress NetworkPolicy”Upgrading the Helm chart from 0.4.0-beta.3 or earlier adds default-deny
ingress NetworkPolicies to the api, web, celery-worker, and celery-beat
pods. The api and web policies admit traffic only from the ingress controller
named by networkPolicy.ingressControllerSelector. By default, that is a
namespace called ingress-nginx.
If your controller runs anywhere else and your CNI enforces NetworkPolicy, those
policies would cut off all traffic to the site, and every pod would still report
Ready. k3s (Traefik) and RKE2 (rke2-ingress-nginx) both run their controller in
kube-system and enforce policies out of the box — and a tunnel such as
Cloudflare Tunnel is not an ingress controller at all, so it never matches the
default no matter which namespace it runs in. To prevent a silent outage, the
first upgrade that adds these policies refuses to render until you choose one
of the following:
# Your controller really is in the ingress-nginx namespace:helm upgrade trueppm ... --set networkPolicy.ingressControllerConfirmed=true
# It runs elsewhere, for example kube-system (or wherever your tunnel client runs):helm upgrade trueppm ... --set-json \ 'networkPolicy.ingressControllerSelector={"namespaceSelector":{"matchLabels":{"kubernetes.io/metadata.name":"kube-system"}},"podSelector":{}}'
# Defer the policies (you can enable them later):helm upgrade trueppm ... --set networkPolicy.enabled=falseThe check above runs only on the upgrade that introduces the policies — fresh
installs and later upgrades skip it, because it compares against a policy that
must already exist for there to be a “transition” at all. A separate check
fires on a fresh install too, for the two paths that bypass any in-cluster
ingress controller entirely: demo.enabled (the documented Cloudflare Tunnel
exposure) and web.service.type: LoadBalancer/NodePort. Neither has a prior
policy to compare against, so the render simply refuses while the selector is
still the default — see Public read-only demo
mode. A cloud load
balancer that sends traffic from outside the cluster needs an ipBlock peer
instead; see
networkPolicy.ingressControllerSelector.
To upgrade a release you installed with --set flags without retyping them,
start from what Helm recorded at install time:
helm -n <namespace> get values <release> -o yaml > current-values.yamlhelm upgrade <release> oci://ghcr.io/trueppm/charts/trueppm --version <version> \ -n <namespace> -f current-values.yamlThis keeps your own settings and picks up the new chart’s defaults.
--reuse-values does not: it also freezes the previous chart’s defaults.
Three more things to check for this upgrade:
- In-cluster monitoring. If Prometheus or a Blackbox exporter in another
namespace scrapes the api Service directly (the health endpoints described in
Observability), set
networkPolicy.monitoringSelectorto that namespace. Otherwise the new policy drops those scrapes, and the dead-letter and beat-staleness alerts go quiet instead of firing. Scrapes that go through your public hostname pass through the ingress controller and are unaffected. - Previews. A client-side
helm upgrade --dry-runcannot see the cluster, so it always reports this refusal. Use--dry-run=serverto preview the upgrade accurately. - GitOps. Tools that render the chart with
helm templateand apply the result themselves, such as Argo CD, never run this check. Set the selector before you sync.
Milestone date shift after recalculation
Section titled “Milestone date shift after recalculation”Upgrading from 0.4.0-beta.4 or earlier changes how the scheduler places
zero-duration milestones (#4079). No migration runs and no stored date changes
during the upgrade itself; the new dates appear the next time each project’s
schedule recalculates, which happens on the next edit that affects the schedule.
Earlier versions gave every milestone a working day of its own, so each milestone on a path pushed everything after it back by one working day. The scheduler now follows the MS Project and Primavera P6 convention: a milestone is a point in time. A milestone that follows work sits on the day that work finishes, and the next task starts on the following working day, exactly as if the milestone were not there. A milestone with no predecessor, such as a project-start milestone, sits at the start of its day and does not delay the work after it.
What you will see after recalculation:
- Dates move earlier. Downstream tasks, forecasts, and the project finish move earlier by one working day for every milestone on their path. A milestone that follows work is now shown on its predecessor’s finish day rather than the day after it.
- Float and the critical path can change. A milestone no longer adds a day to the paths it sits on, so float on parallel paths is recomputed against the shorter path.
- Monte Carlo forecasts move with the schedule. P50/P80/P95 use the same convention, so a plan with fixed durations still simulates to its CPM finish.
- Imported MS Project plans now match their source file. Recalculated dates
for an imported plan no longer drift one day per milestone from the dates in
the
.xmlfile.
Baselines captured before the upgrade keep the old dates, so the first comparison after recalculation can show milestone-driven variance that reflects the convention change, not a schedule change. To compare like with like, capture a new baseline after the project recalculates.
Project and program keys (projects.0154)
Section titled “Project and program keys (projects.0154)”Upgrading from 0.4.0-beta.4 or earlier makes every project and program key
(the code field; see Project and program keys)
unique across the workspace, compared case-insensitively. Migration
projects.0154_object_keys repairs existing data before it adds the uniqueness
constraint:
- A project or program with no code gets one derived from its name. Platform
Migration becomes
PM, and a program called Office Move becomesoffice-move. - When several projects (or several programs) share a code, the oldest keeps
it and each of the others gets a numeric suffix:
PLAT2,PLAT3for projects,atlas-2for programs. - Every other code is kept exactly as it is, including an existing hyphenated
project code such as
GA-SEC. New project keys can’t contain hyphens, but existing ones are never rewritten.
Each code the migration changed is logged at WARNING during migrate with the
prefix key repair (ADR-1237). The log is optional: the database records every
rewrite too. To list them later, run this in python manage.py shell:
from trueppm_api.apps.projects.models import ObjectKey
for row in ObjectKey.objects.filter(source="backfill").select_related("project", "program"): owner = row.project or row.program print(row.kind, row.key, owner.name if owner else "(deleted)")If a rewritten key isn’t what your team wants, rename it on the project’s or program’s General settings page. The old key keeps working in links and now opens the renamed project.
Rolling upgrades. The constraint excludes blank codes for this release. While an older API pod is still running during a rolling update, it can create a project with an empty code without failing. The new version gives that project a key the next time it is edited. The exclusion is planned to be removed in 0.5.
Dead-letter and observability scrape credentials now require superuser
Section titled “Dead-letter and observability scrape credentials now require superuser”Upgrading from 0.4.0-beta.4 or earlier changes who can reach these endpoints.
Through 0.4.0-beta.4, FailedTaskViewSet (dead-letter requeue/drop/list) and
the observability app’s System Health, Prometheus metrics
(/api/v1/health/beat/, /api/v1/health/dead-letter/, /api/v1/health/email/),
retention policy, and telemetry-export endpoints were gated with Django’s
IsAdminUser, which passes for any is_staff=True account — including one
that is not a superuser. They now use IsWorkspaceOperator, which checks
is_superuser directly (#4009). Superusers are unaffected; a WorkspaceRole.ADMIN
membership was never sufficient for these endpoints and still is not.
This is a least-privilege regression for anyone who provisioned a staff-only
service account for scraping or automation against these endpoints — it now
gets 403 Forbidden. Before upgrading:
- Identify any Prometheus scrape job, alerting script, or automation that calls
/api/v1/health/{beat,dead-letter,email}/, the retention endpoints, or the dead-letter queue admin API. - Confirm the JWT it uses belongs to a superuser account (
is_superuser=True), not merely a Django-admin (is_staff=True) account.create_adminproduces a superuser and is unaffected. - If it does not, mint a new token from a superuser account before or immediately after the upgrade, or the scrape/automation starts failing silently (a gauge that stops updating, not a loud error) the moment the new code is live.
See Dead-letter Alerting and Beat Liveness for the affected scrape configs.
Upgrading to 0.3
Section titled “Upgrading to 0.3”0.3 adds new database tables and columns for the agile-team feature set. A
migration is a script that changes TruePPM’s database structure to match a
new version of the code (adding a table or column, for example) — TruePPM ships
them, so you never write or edit one yourself. All of the migrations in 0.3 are
additive (new tables and nullable columns — no destructive operations), so
the upgrade is a standard migrate run with no manual data steps and no downtime
beyond that run. Every deploy path runs that step for you on startup — the
Compose api / api-init service, or the Helm chart’s migrate init container
— and each section below shows how to watch it complete. The new schema:
- Forecast snapshots (
scheduling.0007_projectforecastsnapshot) — a newProjectForecastSnapshottable that persists each project’s P50/P80/P95 Monte Carlo forecast over time, so the Schedule view can show a forecast history. Retention is bounded byMC_HISTORY_CAP(see configuration). - Sprint outcomes (
projects.0064_sprinttaskoutcome,projects.0065_historicalsprint_goal_outcome_sprint_goal_outcome) — a newSprintTaskOutcometable plus agoal_outcomecolumn onSprint(MET / PARTIAL / MISSED), capturing the sprint close-out snapshot. - Scope-change audit (
projects.0054_sprintscopechange_goal_impact_and_more) — agoal_impactcolumn onSprintScopeChange, recording whether a post-activation scope change affected the sprint goal.
If you maintain a fork, note that 0.3 also collapses each app’s migration history
into a 0001_squashed_… migration via Django’s replaces= (issue #1286). Because
the original migrations remain on disk and applyable, an existing database records
the squashed migration as already-applied and upgrades as a no-op — there is no
drop, recreate, or data step.
Docker Compose (development)
Section titled “Docker Compose (development)”git pull origin maindocker compose builddocker compose up -dbuild, not pull (#3189). The dev stack’s api and web services are
build:-based, and celery references image: trueppm-api:local with no
build: of its own — so docker compose pull has nothing to fetch for the
services that changed, and reports success. Building is what actually picks up
the new code.
Migrations run automatically when the api container starts.
Single-server with Docker Compose
Section titled “Single-server with Docker Compose”cd /opt/trueppmgit pull origin main
# Update the target version in .env:# APP_VERSION=0.4.0-beta.4
docker compose -f docker-compose.prod.yml pulldocker compose -f docker-compose.prod.yml up -dThe api-init service runs migrate --noinput before the API starts. Watch it complete:
docker compose -f docker-compose.prod.yml logs -f api-init# Should end with: "0 unapplied migration(s)." or a list of applied migrations.Helm / Kubernetes
Section titled “Helm / Kubernetes”helm upgrade trueppm oci://ghcr.io/trueppm/charts/trueppm \ --version <version> \ --namespace trueppm \ -f my-values.yaml<version> is the release version without a leading v, for example
0.4.0-beta.4. Always pass it: Helm skips pre-release chart versions unless you
name one, so while 0.4 is in beta a bare helm upgrade fails with could not locate a version matching provided version string. --devel selects the newest beta.
Migrations run in a migrate init container of the api Deployment before the new pods start serving. Check its logs:
kubectl logs -n trueppm deployment/trueppm-api -c migrateUpgrading to the hardened Helm chart
Section titled “Upgrading to the hardened Helm chart”The Helm chart now installs secure by default: it generates the PostgreSQL and
Valkey passwords and stores them in a chart-owned connection Secret
(<release>-trueppm-connection, annotated helm.sh/resource-policy: keep),
injects DATABASE_URL / REDIS_URL via secretKeyRef, enables Valkey auth by
default, and applies restricted container security contexts. A few notes when
upgrading from a pre-hardening release:
- Rotate the old default password. Earlier chart versions shipped a default
database/cache password of
trueppm. If you ran with that default, rotate it. The simplest path is to set explicit, strong passwords on the upgrade so the chart writes them into the connection Secret and the bundled datastores pick them up:Prefer supplying these through an external Secret overTerminal window helm upgrade trueppm oci://ghcr.io/trueppm/charts/trueppm \--namespace trueppm \-f my-values.yaml \--set postgresql.auth.password="<new-strong-password>" \--set valkey.auth.password="<new-strong-password>"--set. After the rollout settles you can clear the explicit values and let the chart manage the password from the connection Secret going forward. - Leave the passwords blank to keep the generated ones. On an upgrade where a
connection Secret already exists, leaving
postgresql.auth.password/valkey.auth.passwordempty makes the chart read the existing password back rather than minting a new one — so re-runninghelm upgradenever churns the credential or orphans the database PVC. - The connection Secret survives
helm uninstall. Theresource-policy: keepannotation means an accidental uninstall/reinstall reuses the same password and keeps the existing data reachable. If you intend a clean wipe, delete the Secret and the PersistentVolumeClaims explicitly. - Using managed datastores? When
postgresql.enabled/valkey.enabledarefalse,env.DATABASE_URLandenv.REDIS_URLare now required — the chart fails the render if either is missing. Add them (ideally via an external Secret) before upgrading. - App-side auth/CSP defaults. The refresh token now rides an httpOnly Secure cookie and a strict CSP header is sent on every response. A standard deploy needs no changes: TruePPM is served from a single origin, which is what these defaults assume. Splitting the SPA and API across hostnames is not supported — see Split-origin deploys.
Rollback
Section titled “Rollback”Migration reversibility — read this first
Section titled “Migration reversibility — read this first”The safe rollback path depends entirely on what the upgrade’s migrations did, so classify them before you touch anything (the release’s operational change note states this):
- Additive-only (new tables, new nullable columns, new indexes — the common case, and every 0.3 migration). The new schema is a superset of the old, so the previous image runs against it unchanged. Roll back the image/chart revision only — do not reverse the migrations and do not restore the database. The extra tables/columns sit unused until you roll forward again. The readiness probe cannot verify “additive-only” for you, so it holds the rolled-back pods out of the Service until you confirm that classification — see the caution below for the one-line opt-out.
- Destructive or transforming (a column drop/rename, a type change, or a data
backfill that rewrites rows). The old code cannot run against the new schema,
and reversing the migration loses the data the new schema captured. Here a
clean rollback means restore the pre-upgrade backup — a
migratereverse is not a substitute, because Django’s reverse operations recreate structure but cannot recover dropped or transformed data. This is why the pre-upgrade backup is mandatory, not optional.
Docker Compose rollback
Section titled “Docker Compose rollback”# Restore the previous APP_VERSION in .env, then:docker compose -f docker-compose.prod.yml pulldocker compose -f docker-compose.prod.yml up -dIf the migration applied schema changes, restore from the pre-upgrade backup:
docker compose -f docker-compose.prod.yml downdocker compose -f docker-compose.prod.yml up db -d./scripts/restore.sh --artifact ./backups/trueppm-backup-<timestamp>.tar.gz --yes# Then bring up the full stack at the previous version.Concurrent migrations at replicaCount >= 2
Section titled “Concurrent migrations at replicaCount >= 2”migrate runs as a per-pod init container, not as a pre-upgrade hook Job,
so at two or more API replicas every pod runs it concurrently against one
database. Django has no concurrency control of its own: each process reads
django_migrations, decides the same migration is unapplied, and both run it.
From 0.4 the chart runs manage.py migrate_locked, which serializes them behind
a PostgreSQL advisory lock. The losers block rather than fail, and by the time
they acquire the lock the winner has finished, so their own migrate is a
no-op. Advisory locks release automatically when the holder’s connection dies,
so a killed init container cannot wedge the next rollout.
Nothing to configure. If a migration legitimately runs longer than ten minutes,
raise the wait with --lock-timeout in the init container’s command, or scale
the API to one replica for that upgrade.
Helm rollback
Section titled “Helm rollback”helm rollback trueppm -n trueppmThis restores the previous chart revision. If the migration applied schema changes, restore from backup and trigger a fresh migrate run.
Post-upgrade verification
Section titled “Post-upgrade verification”# Check all containers are healthydocker compose -f docker-compose.prod.yml ps# orkubectl get pods -n trueppm
# Hit the health endpointcurl https://trueppm.example.com/api/v1/health/# → {"status": "ok"}
# Confirm the expected version is running — check the deployed image tagkubectl get deployment -n trueppm trueppm-api \ -o jsonpath='{.spec.template.spec.containers[0].image}'# or: helm list -n trueppmCommon issues
Section titled “Common issues”Migrations fail on startup
Check that DATABASE_URL is correct and the database is reachable. Run migrations manually to see the full traceback:
docker compose -f docker-compose.prod.yml exec api python manage.py migrate --noinputStatic files not updating
Trigger a collectstatic run:
docker compose -f docker-compose.prod.yml exec api python manage.py collectstatic --noinput --clearWebSocket connections drop after upgrade
Expected — clients reconnect automatically within a few seconds. The Channels layer (Valkey) is not drained between upgrades; in-flight messages are lost but clients recover via the reconnect loop.