Seed data schema
This page is for developers authoring or modifying a bundled sample, or building an importer that targets the canonical format. If you just want to load or export demo data, see Sample projects & JSON import/export.
The format
Section titled “The format”One seed document describes one program and all of its projects. The JSON Schema is the contract:
- v2 —
packages/api/src/trueppm_api/apps/projects/schemas/seed_v2.json(ADR-0114). - v1 —
seed_v1.json(ADR-0109); still loads. v2 is an additive superset.
validate_seed() checks a document against the schema for its major version and
then runs a referential-integrity pass (no dangling slug or task references)
that JSON Schema cannot express. Every error is anchored to a JSON path.
External references, blockers, and risk narrative
Section titled “External references, blockers, and risk narrative”Three capabilities the product models had no seed surface for until 0.4, so no bundled sample could demonstrate them.
tasks[].attachments[] is URL-only: {external_url, external_title?, is_pinned?, uploaded_by?}. TaskAttachment enforces file XOR external_url at
the database level, and a seed is a text document with no bytes to carry, so the
file half is deliberately inexpressible rather than half-supported. Note that
links is a different key entirely — it is the task-to-task TaskRelation
graph (relates to / blocks / duplicates), not external references.
tasks[].blocked is the explicit human blocker flag (ADR-0124), not the
derived “has incomplete predecessors” signal the board card owns:
{reason?, since?, type?, blocking_task?, by?}. A present cluster means blocked,
and reason is optional: an export includes it only for the task’s assignee or an
@-mentioned user (ADR-0124), and the importer stamps (private) when it is absent. blocking_task is a
soft “waiting on” link that never enters CPM; a scheduling constraint is a dependencies[] edge.
since matters more than it looks. Task.save() stamps blocked_since with
timezone.now(), but the importer inserts through bulk_create_tasks, so
save() never runs and nothing stamps it — an unstamped blocker renders no age,
and age is the entire triage signal (“3 tasks blocked more than a week”). The
importer therefore sets it explicitly, falling back to the task’s planned start
and then the project’s, so the value is never null while the flag is raised.
task.block / task.unblock events give a blocked span a real duration on
the timeline. task.block writes the flag and backdates blocked_since to the
beat; task.unblock empties reason only and lets Task.save() run its own
cascade (clearing blocked_since, blocker_type, blocking_task and
blocked_by), so “unblocked” has exactly one definition.
risk.note events append a RiskComment, giving a risk.status flip its
reason. Without it a risk walks OPEN → MITIGATING → RESOLVED with no artifact
of the work — the register records that a risk was mitigated and never how.
Unlike risk.status, notes are reconstructed on export: a comment is
append-only and does not depend on the risk’s current status, so replaying it
reproduces the rows it came from.
The schema is a two-way contract
Section titled “The schema is a two-way contract”additionalProperties: false is set on every definition, so an unknown key
is a validation error. The reverse — a key the schema declares that no reader
implements — is the more dangerous direction, because the schema tells an author
it works. Four keys shipped that way and were only found by audit:
task.dor, project.board_columns, project.agile_features, and
baseline.captured_at.
test_declared_keys_are_implemented now closes that direction: every key
declared in either schema must be mentioned by importer.py, replay.py, or
forecast_backfill.py. A key that is genuinely read some other way needs an
entry in that test’s _EXEMPT map with the reason — an unexplained
exemption would let the gate be silenced by adding a line, which is the failure
it exists to prevent.
Three consequences worth knowing when authoring:
baseline.captured_atis honored. It setsBaseline.created_at, which is otherwiseauto_now_add. The interval between two baselines is what planned-vs-actual is measured over, so a rebaseline authored 75 days after its contract baseline must land 75 days after it.board_columnsmirrorsBoardColumnConfig, not a list of labels. Each entry is{status, label, visible, color?, wip_limit?, age_threshold_days?, lanes?}and all five canonical statuses must appear exactly once —statusis what every downstream reader (burndown, WIP, velocity) keys on, and a label list cannot carry it. Omit the key entirely to leave a project on the API’s defaults; emitting them would turn “uses the defaults” into “pinned today’s defaults”, which is a different claim on re-import.project.healthis a PM override, not a computed value. Omit it (or setAUTO) to leave the project’s chip to the rollup. The explicit values —ON_TRACK,AT_RISK,CRITICAL— say a human made a judgment, so a pack that sets one on every project turns the chip into decoration. Atlas sets exactly one, on Migration Tooling, and leaves the other two onAUTOso the difference is legible.
Who can see a project: accounts[].role vs projects[].members[]
Section titled “Who can see a project: accounts[].role vs projects[].members[]”A seed carries two different membership levels, and confusing them produces a pack whose personas cannot open anything:
accounts[].rolegrants a program membership. That reaches the program rail and no project.projects[].members[]grants project memberships —{account, role}— and project access is scoped by these. This is the one a persona needs in order to see a project at all.
Omit members and every account is granted its program-level role on that
project, so a pack written before the key existed still works. Declare it and it
replaces the fallback for that project: only the accounts listed become
members. That replacement is the point — it is what lets one person hold
different roles on different projects, which is the only way a seed can
demonstrate project-scoped RBAC:
{ "slug": "platform-core", "members": [ { "account": "priya", "role": "ADMIN" }, { "account": "jordan", "role": "MEMBER" } ]}In atlas-platform-launch.json, priya is ADMIN on Platform Core, MEMBER on
Migration Tooling, and absent from GTM Readiness.
Two rules the importer enforces regardless of what a seed says. The importing
user is always granted OWNER and can never be demoted by a members entry
naming them. And an account that resolved to None — a pre-existing real user on
an untrusted import — is skipped, so a crafted seed cannot pull a stranger into a
program. A members entry naming no accounts[] slug is a validation error
rather than a silent skip, because the silent version reproduces exactly the
blindness the key exists to fix.
Forecast trend and Monte Carlo run history: forecast_history + mc_history_*
Section titled “Forecast trend and Monte Carlo run history: forecast_history + mc_history_*”A freshly-loaded project has no history for the forecast-trend chart or the
Monte Carlo run panel — both are populated only after a real recompute or a
real “Run” click, which never happens on a demo that nobody has touched yet. A
project’s forecast_history block backfills that history at import time
instead of hand-authoring dozens of rows:
{ "forecast_history": { "days": 60, "commitment_finish": "A+40", "cpm_start": "A+55", "cpm_end": "A+46", "p50_start": "A+58", "p50_end": "A+50", "p80_start": "A+62", "p80_end": "A+54", "p95_start": "A+68", "p95_end": "A+60", "mc_iterations": 2500, "completion_ratio": 0.6 }}One ProjectForecastSnapshot is synthesized per day across the window, with
the deterministic CPM finish and each Monte Carlo percentile interpolated
linearly from its *_start to its *_end date while commitment_finish
stays fixed. The percentiles are all-or-none — author the full p50/p80/p95
start+end set, or omit them entirely for a CPM-only history (the MC lines stay
null until someone runs Monte Carlo for real).
MonteCarloRun history (the run-attribution panel and the top-drivers
tornado read this, not the snapshot table) is derived from that same backfill
— never recomputed independently, so the two views of one drift can’t
disagree. Roughly one run is persisted per week across the window (a real
Scheduler+ user does not click “Run” daily), attributed to the highest-ranking
Scheduler+ persona the seed casts on that project (SCHEDULER, then ADMIN, then
OWNER, project scope before program scope). A project whose block omits the MC
percentiles gets no run history either — MonteCarloRun.n_simulations is
required, and there is nothing to derive it from.
The MC history panel itself is gated per program by two more keys, both
null-means-inherit-the-workspace-default the same way public_sharing and
allow_guests do:
mc_history_enabled(boolean) — turns the run-history panel on for every project in the program.mc_history_attribution_audience("admin_owner"/"scheduler_plus"/"none") — who may see the run-author name on a history row.
A program with at least one project carrying forecast_history should set
mc_history_enabled: true, or the backfilled runs exist but the panel that
reads them stays off.
Worked examples: the bundled fixtures
Section titled “Worked examples: the bundled fixtures”The prose below explains the format’s shape. The five bundled fixtures are the format — correct, non-trivial, and validated on every CI run. Reading one beats inferring from a description, so they are downloadable from any running instance rather than only from a repository checkout:
| Sample | Download | What it demonstrates |
|---|---|---|
| Atlas Platform Launch | GET /api/v1/programs/samples/atlas-platform-launch/download/ | The largest surface — a hybrid multi-project program, cross-project dependencies, all four dependency types plus a lead, three-point estimates, two baselines on one project (a kickoff capture and the re-plan that superseded it), a calendar exception that moves the program finish, a manual health override, and a populated risk register with mitigation arcs. |
| Aurora Mobile App | .../aurora-mobile-app/download/ | Pure agile — an epic-grouped backlog and a sprint history that goes wrong and recovers: an epic descoped mid-sprint after beta feedback, the velocity dip that causes and the sprint that climbs back out, a cancelled sprint whose scope folds forward, capacity that moves with leave and holidays, and a sprint-zero baseline to measure the pivot against. No CPM. |
| Bayside Civic Center | .../bayside-civic-center/download/ | Pure waterfall under constraint — all four dependency types, calendar-aware lag, a contract baseline plus a change-order rebaseline captured months apart, and a site calendar whose stand-downs actually bite: a crane window that stretches the framing tail and is absorbed by its float, and a contract weather allowance that pushes the certificate of occupancy. Its risk register carries triggers and contingencies, one realized risk whose mitigation failed and whose contingency shows up as baseline variance, and the only TRANSFER response in any pack, with its terms and its limits stated. |
| 1.0 GA Launch | .../ga-launch/download/ | Program coordination — four workstreams (platform hardening, SOC 2 readiness, security remediation, launch) shipping one outcome. Cross-project dependencies form a critical path that runs across projects, shared people over-allocate in overlapping windows, and every project carries the full 5-role RBAC matrix. The security workstream runs a WIP-limited remediation board. |
| Helios CRM Replacement | .../helios-crm-replacement/download/ | The entry-level hybrid — a completed waterfall phase feeding an agile build phase across one cross-phase dependency, sprints that state the goal their outcome is judged against, a mid-sprint injection the team rejects, and a mitigation arc that costs something: a dry-run harness scheduled against the migration risk, paid for by displacing another story out of the sprint. |
In the UI the same list lives at Settings → System → Demo data, with each
file’s size, entity counts and SHA-256. The catalog endpoint
(GET /api/v1/programs/samples/) returns the same metadata as JSON.
The counts shown beside each file come from the same inspect_seed() that backs
the dry run (POST /api/v1/programs/import/validate/,
ADR-0651),
so the catalog and the validator cannot disagree about a document.
Why the format looks the way it does
Section titled “Why the format looks the way it does”ltree WBS paths. Tasks are identified within a project by an ltree path
("1.2.3") rather than a UUID, so a seed file carries stable, human-readable,
per-project task identity. Cross-project references use "<project-slug>:<wbs>".
File-local stable slugs. Seed files carry no UUIDs — they would collide
across instances and re-imports. Instead, accounts, calendars, resources, and
sprints use kebab-case slugs that are a file-local symbol table: the
importer resolves them to freshly-minted UUIDs at import time. The one slug that
persists is the program slug, which is written into Program.code as the
program’s natural key. That is what makes re-import idempotent — a program with
a matching code is replaced, not duplicated. The seed format carries no stable
entity ids until 0.5 (#1959),
so there is nothing for a field-level merge to key on: replace-then-rebuild is
the correct idempotency model for this format. Because it is destructive, the
REST import refuses a collision until the caller confirms it, and the
replacement is a soft delete — the replaced program’s projects move to
project Trash, where each can be restored individually as a standalone project.
The program shell itself is not recoverable, and a restored project does not
return to it. Only the disposable demo path (is_sample) still hard-deletes.
Keys carry over on both the synchronous and the queued import: the rebuilt
program takes the replaced program’s key, and each rebuilt project takes the key
of the replaced project with the same name, so existing /projects/PLAT/… links
open the rebuild. On the queued path the keys move only once the rebuild
succeeds; if it fails, the projects in Trash keep their keys. A project restored
from Trash after its key moved gets a new derived key. See ADR-0726.
Three-point estimates as an all-or-none sub-object. A task’s PERT estimate
is an estimate: { optimistic, most_likely, pessimistic } sub-object. Modelling
it as a single object makes the all-or-none invariant
(ADR-0093)
structurally enforceable: a task has all three points or none. Imported
estimates are written as accepted, bypassing estimation governance.
Anchor-relative dates + an events timeline (v2). A v1 seed pins absolute
dates, so a bundled demo ages. v2 instead authors dates as offsets from an
import-day anchor ("A-120", "A+15"), weekend-snapped to a working day
via the project calendar — so the demo always reads as current. On top of that,
an ordered events array is replayed with backdated history: each beat
writes a history row dated to the event, so a completed task shows dated
transitions by named people, closed sprints accumulate real burndown snapshots,
and velocity is actual history. A deterministic synthesizer fills the unauthored
“boring middle” — any task whose final column implies it passed through earlier
ones gets synthetic transitions, seeded reproducibly per program and task so
re-import is stable.
The implemented v2.0 action set covers status, assignment, estimate, points,
comment, AC-met, block/unblock, sprint activate/close, scope inject/resolve,
baseline capture, risk status and notes, and the retrospective pair
retro.action / retro.promote — an action item on the target sprint’s retro,
and its promotion (matched by body) into a BACKLOG task with no sprint,
exactly what the live promote endpoint produces.
v2.1 adds the collaboration layer (schema_version: "2.1", same
seed_v2.json; every 2.0 document is still valid). New events: task.note (a
dated task note, decision/pinned optional — decisions feed the Decisions
view), task.react and task.ack (a reaction or acknowledgement on a comment,
addressed as comment:<slug>), and time.log (a TimeEntry of minutes,
dated to the beat and rejected if forward-dated). task.comment gains slug
and reply_to — replies are one level deep and stay on their thread’s task —
and task.ac_met gains an optional criterion index. time.log,
task.react and task.ack must name their actor: an hour or a reaction by
nobody in particular is never re-attributed to the importing user. New
sections: program.backlog_items (a pulled item names the task it became in
pulled_to), program.ceremonies (program-level only — standup,
sprint review, retrospective and the other sprint events are rejected, as
in the API), and tasks[].acceptance_criteria.
v2.1 also carries the team and task-shape layer. projects[].team names the
project’s default team and sets its Scrum Master and Product Owner facets — at
most one of each, and only on a member of the project, because team membership
itself mirrors project membership. resources[].skills puts skills on the
workspace catalog (matched case-insensitively, so two samples share one “Python”)
with a beginner / intermediate / expert proficiency, and
tasks[].skill_requirements states what a task needs, which the assignment fit
compares. tasks[].recurrence makes a task a recurrence template — only the rule
is seeded, and occurrences spawn on the generator’s horizon — with the same
conditional rules as the API (weekly needs weekdays, monthly a day of month, one
end at most). tasks[].is_subtask marks a drawer subtask: one level under a leaf
task that holds no structural children, and never a parent itself. task.type
adds tech_debt.
Two further v2.1 sections are sample-only: program.agent_actions and
projects[].share_links. Only the bundled-sample loader honors them; every
other import path rejects the file. Seeded agent actions go through the real
hash-chained audit log, marked as sample data in hashed fields
(actor_token_prefix: "sample", a [Sample data] summary prefix), and
share-link tokens are generated at load, never read from the file. On the
sample path the importer also synthesizes logged time on completed and
in-flight work and submits each fully elapsed week; a generic import writes
only the time.log beats the file authors.
The v2 exporter round-trips backlog items, ceremonies, criteria, notes,
threads, reactions and acknowledgements. It never exports agent actions,
share links or time.log: the first two are evidence and credentials, and
per-person hours would reach anyone who can run an export.
baseline.capture snapshots the project’s live task state at the beat’s
time, dated and attributed — the events-timeline counterpart to a declared
baselines[] row, and the only way to give a rebaseline the actor + reason a
static row cannot carry (#3495). Pair it with a task.comment or risk.note
naming the reason, on the same or an adjacent beat. is_active (default
false) collapses the live app’s separate capture-then-activate steps into one
beat: when true, any baseline already active for the project is deactivated
first, so an authored rebaseline can supersede the one it replaces without a
second, unauthored write. Because replay runs before the post-commit CPM
recalc, the snapshot’s start/finish come from planned_start, not
early_start — a project whose tasks rely on the engine (rather than an
authored planned_start) for their dates will still show has_cpm_dates: false on a beat-captured baseline, same as it would on the equivalent
declared row.
retro.action’s optional notes sets the retro’s own summary, not the
action item. SprintRetro.notes is a plain free-text field with no
dedicated beat of its own — the retro is only ever created lazily, the first
time a retro.action targets its sprint (#3497). Put notes on any one of
that sprint’s retro.action beats — the last one is the natural place, once
the meeting has something to summarize — and it sets/updates the parent
retro’s notes; every other retro.action on the same sprint still only
carries its own action item’s body/assignee/points. A retro with notes
but zero action items has no way to reach the database through replay: there
is no standalone “open retro” beat.
Authoring a new sample
Section titled “Authoring a new sample”The bundled samples are generated by developer scripts, then committed as schema-validated fixtures — never hand-edited as raw JSON:
scripts/seeds/build_atlas_seed.py— Atlas (hybrid-large).scripts/seeds/build_samples.py— Aurora, Bayside, Helios.scripts/seeds/build_ga_launch.py— 1.0 GA Launch.
Each script builds the document in Python, validates it against the schema, and
writes the fixture under
packages/api/src/trueppm_api/apps/projects/fixtures/seeds/. To add a sample:
- Add a builder to one of the scripts (or a new one), emitting
schema_version"2.0"and anchor-relative dates. - Re-run the script to regenerate and validate the fixture.
- Register the sample’s key and filename in
apps/projects/seed/samples.pyso the loader and picker surface it.
Seed files are self-contained by design (ADR-0109) — a document carries
everything it needs, with no references to external files. The shared demo cast
(consistent people, roles, and capacity profiles reused across samples) is
therefore a shared authoring convention in the build scripts, not a separate
sample-resources.json the importer would have to dereference.