Observability, Incidents, and Governance
Operate, restore, govern, and retire a generated-application fleet, ending with Workboard export and deletion, a separate revenue-app connector-revocation reference, retention expiry, and residual-access proof.
The enterprise problem and today’s slice
Enterprise problem: One healthy launch does not show that a customer can control a fleet through failures, cost pressure, legal obligations, restoration, and retirement; hidden routes, grants, or backups can survive after an app appears deleted.
Whole-course context: Day 5 produced digest-bound v1/v2 releases, compatible migration and canary evidence, a rollback trace, and a corrected production deployment; today operates those artifacts across their remaining lifecycle.
Today’s slice: Customer and HelixWorks administrators correlate cross-plane telemetry and audit, enforce SLOs, quotas, and cost controls, rehearse an incident and restore, govern residency and retention, then export and retire one app completely.
End-of-day evidence: An assurance dossier joins incident, break-glass, restore, export, decommission, grant-revocation, deletion, backup-expiry, and residual-access observations through immutable IDs.
Still unsolved: Only organization-specific risk acceptance and ongoing reassessment remain; the course claims the tested archetype envelope, not universal correctness.
Customer use cases
Operations guidance can celebrate dashboards while leaving restoration, retirement, and surviving external grants untested, so the customer's complete fleet jobs must terminate in evidence owned by the correct plane. These use cases keep Workboard retirement separate from the revenue application's connector lifecycle.
| Use case ID | Actor | Customer job | Success outcome | Denial or recovery evidence |
|---|---|---|---|---|
| D06-UC-01 | Customer fleet administrator | Monitor each app's SLOs, quotas, costs, release, and lifecycle state | A non-authoritative fleet projection links current owner records and alerts without becoming a new permission source | A stale projection is visibly marked and cannot authorize a release, support action, or data access |
| D06-UC-02 | Incident commander | Contain a Workboard incident with bounded break-glass support | Plane-specific grants expire, commands are audited, risky paths stop, and customer communication records scope | An overbroad, unapproved, or expired grant is denied and independently audited |
| D06-UC-03 | Recovery operator and customer approver | Restore Workboard into isolation and decide whether to return traffic | Checksum, RPO, RTO, schema, tenant, role, object, queue, and audit continuity evidence support an explicit traffic decision | Failed integrity or isolation checks hold the recovery environment and leave production routing unchanged |
| D06-UC-04 | Workboard app owner | Export and retire the complete Workboard resource graph | Routes, workloads, identities, app data, project records, and caches reach scoped tombstones while app-revenue connector grants remain positively active | Residual probes expose any surviving Workboard path; no sf-pipeline or snow-bookings revocation is emitted |
| D06-UC-05 | Compliance officer | Observe backup and retained-copy expiry after Workboard retirement | Policy time, holds, vault catalog, deletion checkpoints, and residual probes prove expiry or authorized retention | An active hold prevents purge, remains narrowly accessible, and is reported rather than mislabeled deleted |
| D06-UC-06 | Revenue app owner | Retire app-revenue separately and revoke its connector grants | Only the revenue retirement emits revocations for sf-pipeline, snow-bookings, and their workload grants while source-native data and audit remain source-owned | Revocation denial or timeout leaves the revenue retirement incomplete; Salesforce and Snowflake source records are never claimed deleted |
Actor-centred user stories
Fleet controls that lack actor-centred acceptance can merge responsibilities and hide whether a customer, operator, or compliance owner can complete the job, so every operational use case needs an observable story.
| Story ID | Use case IDs | User story | Observable acceptance conditions |
|---|---|---|---|
| D06-US-01 | D06-UC-01 | As a customer fleet administrator, I want a current inventory view with service and cost signals, so that I can act before one app harms its users or budget. | Projection freshness, source versions, release digest, SLO burn, quota, cost, owner, and lifecycle state are visible; stale state cannot authorize. |
| D06-US-02 | D06-UC-02 | As an incident commander, I want separate expiring support grants, so that responders contain harm without standing cross-plane access. | Ticket, approver, plane, commands, expiry, revocation, and denial probes are correlated through immutable IDs in separate audit stores. |
| D06-US-03 | D06-UC-03 | As a recovery operator, I want to restore Workboard in isolation, so that the customer can verify data and tenant integrity before traffic changes. | Restore checksum, RPO/RTO, schema, equal-local-ID isolation, roles, objects, queues, audits, and traffic decision are recorded. |
| D06-US-04 | D06-UC-04 | As a Workboard app owner, I want export and retirement to cover the whole Workboard graph, so that no controllable route or credential survives while unrelated revenue connectors keep working. | Export receipt, quiescence, tombstones, Workboard denial probes, and positive revenue connector probes are all observable. |
| D06-US-05 | D06-UC-05 | As a compliance officer, I want retention expiry and holds reported honestly, so that deletion claims match every remaining copy and obligation. | Vault expiry, hold scope, tombstone checkpoints, retained access controls, and residual probes support retired-and-expired or a named retained state. |
| D06-US-06 | D06-UC-06 | As a revenue app owner, I want connector revocation tied only to revenue retirement, so that Workboard deletion cannot disrupt Salesforce or Snowflake access used by another app. | Revenue retirement references Day 4 connector IDs, revokes their grants, preserves source-native records, and records connector and source signals separately. |
End-to-end product flows
An operational checklist is not end-to-end evidence when it starts inside infrastructure or ends at “done,” so these flows start with a customer action and terminate in observed service, recovery, retirement, or denial state.
| Flow ID | Use case IDs | Path | Trigger | Numbered steps | Terminal evidence |
|---|---|---|---|---|---|
| D06-FLOW-01 | D06-UC-01 | Happy | Fleet administrator opens the organization inventory | 1. Projection reads versioned owner records; 2. join authorized provider, runtime, app, connector, and source signals; 3. calculate SLO, quota, and cost status; 4. show freshness and drill-down links | View names projection time, source versions, app owner, digest, alerts, quota, budget, and lifecycle without granting new authority |
| D06-FLOW-02 | D06-UC-02 | Failure | Customer reports Workboard mutation failures and the SLO alert fires | 1. Declare incident; 2. approve narrow runtime break-glass; 3. contain release and risky paths; 4. revoke sessions and support grant; 5. correlate separate audit receipts | Incident timeline binds alert, commands, containment, grant expiry, denial probes, customer notice, and preserved evidence |
| D06-FLOW-03 | D06-UC-03 | Recovery | Incident commander selects Restore to isolated environment | 1. Select eligible backup; 2. restore under new environment identity; 3. verify checksum and schema; 4. run tenant, role, object, queue, and audit tests; 5. approve or hold traffic | Restore job records measured RPO/RTO and a customer-approved traffic decision; any failed invariant leaves routing unchanged |
| D06-FLOW-04 | D06-UC-04 | Happy | Workboard owner selects Export and retire application | 1. Freeze Workboard changes and drain work; 2. create encrypted export and custody receipt; 3. disable routes, deployment, runtime identity, sessions, shares, and project access; 4. delete scoped data through checkpoints; 5. probe Workboard denial and revenue connector health | Workboard tombstones and denial matrix are complete, while positive probes show sf-pipeline and snow-bookings still active for app-revenue |
| D06-FLOW-05 | D06-UC-05 | Recovery | Retention scheduler reaches Workboard backup expiry | 1. Evaluate policy and legal holds; 2. cryptographically erase or expire eligible backups; 3. retain only hold-authorized records; 4. verify catalog and former paths; 5. update lifecycle decision | Evidence says retired-and-expired or retired-with-policy-retained-records and identifies every authorized remainder |
| D06-FLOW-06 | D06-UC-06 | Happy | Revenue owner selects Retire revenue application | 1. Quiesce app-revenue; 2. revoke workload grants for sf-pipeline and snow-bookings; 3. disable connector routes and derived caches; 4. preserve source-native data and audits; 5. run broker and source probes | Separate revenue retirement receipt shows connector denials, cache expiry, source-owned records untouched, and immutable connector/source evidence |
System design derived from the flows
A generic operations architecture can turn a fleet index into an authority or collapse independent audits into one privileged store, so each use case maps to services and state derived from its actual flow.
| Use case ID | Entry point | Responsible services | Authoritative store | Failure evidence |
|---|---|---|---|---|
| D06-UC-01 | Organization fleet inventory | Fleet projection builder, telemetry correlation service, SLO engine, quota and cost services | Owning project, release, deployment, SLO, quota, and cost stores remain authoritative; inventory is a versioned projection only | Staleness marker, source-version mismatch, missing signal, and denied unauthorized drill-down |
| D06-UC-02 | Declare incident and Request support actions | Alert and incident service, support-access broker, plane-specific policy enforcers, revocation orchestrator, customer communication service | Incident store and separate provider, runtime, and Workboard audit stores | Grant denial, command denial, expiry, revocation timeout, and immutable incident timeline |
| D06-UC-03 | Restore to isolated environment | Backup catalog and vault, key service, restore orchestrator, environment controller, verification runner, traffic decision service | Backup, restore-job, environment, and generated-app stores retain separate ownership | Checksum, RPO/RTO, schema, tenant, authorization, object, queue, or audit-continuity failure holds traffic |
| D06-UC-04 | Export and retire Workboard | Export service, lifecycle orchestrator, route and deployment controllers, identity and sharing revokers, deletion workers, probe runner | Workboard app stores, provider project stores, runtime controllers, export receipts, deletion jobs, and tombstones | Incomplete checkpoint or surviving Workboard probe blocks retirement; revenue connector health failure flags out-of-scope damage |
| D06-UC-05 | Confirm retention expiry | Retention engine, legal-hold service, backup vault, key-erasure service, residual-probe runner | Retention policy, backup catalog, hold, deletion checkpoint, and tombstone stores | Active hold prevents purge; missing expiry or catalog evidence prevents an expired claim |
| D06-UC-06 | Retire revenue application | Revenue lifecycle orchestrator, connector broker, grant revoker, private-route controller, cache deletion worker, source probe adapters | Day 4 connector and workload-grant stores plus separate connector lineage and source-native audit stores | Connector denial timeout, surviving derived cache, source mutation, or missing source-native observation blocks completion |
Data model and ownership
Deletion claims become false when inventory copies or tombstones are treated as original authority, and incident evidence becomes overprivileged when plane-specific audits are merged. This model keeps projections, owner records, audit stores, connector signals, and source-native observations distinct.
Generated-application database: Required in this slice — the Workboard data service owns tenant, board, and todo state that restoration verifies, export transfers, and the Workboard deletion job removes under retention and hold policy.
| Record or entity | Store and owner | Primary key | Foreign key or opaque reference | Tenant key | Material constraint | Lifecycle and deletion | Use case IDs |
|---|---|---|---|---|---|---|---|
| ApplicationInventoryProjection | Fleet projection store owned by inventory service | Compound organization_id, application_id, projection_version | Opaque versioned references to owning project, release, deployment, SLO, quota, and cost records | organization_id partitions customer views | Projection is read-only, freshness-marked, and never authorizes actions | Rebuilt from owner stores; stale versions expire and are deleted without deleting source records | D06-UC-01 |
| ServiceLevelObjective | Reliability store owned by SLO service | slo_id | application_id and signal-definition references | organization_id partitions objectives | Indicator, target, window, owner, and security invariants are versioned | Retained with decision history; obsolete versions archive then expire by policy | D06-UC-01, D06-UC-02 |
| Incident | Incident store owned by incident service | incident_id | Opaque references to alert, application, release, trace, and audit events | organization_id partitions incidents | Severity, roles, scope, timeline, customer notices, and decisions are append-only | Retained under incident policy and legal holds, then archived or expired | D06-UC-02, D06-UC-03 |
| BreakGlassGrant | Support-access store owned by access broker | break_glass_grant_id | incident_id plus plane-specific resource reference | organization_id partitions customer approval | One plane, narrow actions, approvers, strong authentication, short expiry, and independent revocation | Expires automatically; metadata and denial proof retained, credential destroyed | D06-UC-02 |
| ProviderAuditEvent | Provider audit store owned by HelixWorks security | audit_event_id | Opaque reference to provider actor, project, incident, or lifecycle event | organization_id partitions provider customers | Append-only integrity metadata; does not grant runtime or app access | Retained and expired under provider audit policy or legal hold; payload is never copied into tombstones | D06-UC-01, D06-UC-02, D06-UC-04 |
| WorkboardAuditEvent | Generated-app audit store owned by Workboard customer policy | audit_event_id | Opaque reference to app_subject_id, app resource, incident, export, or deletion event | app_tenant_id partitions application audit | Append-only and independently authorized from provider audit | Retained or expired under app policy and holds; deleted payload is not embedded in tombstones | D06-UC-02, D06-UC-03, D06-UC-04, D06-UC-05 |
| ConnectorLineageSignal | Connector lineage store owned by connector control plane | lineage_event_id | Day 4 connector_id, workload_grant_id, and opaque source request reference | organization_id partitions connector operations | Signal proves broker decision and minimized movement, not source-row ownership | Retained separately from app audit; expires by connector evidence policy | D06-UC-01, D06-UC-04, D06-UC-06 |
| SourceNativeAuditReference | Reference index owned by source-audit adapter | source_audit_ref_id | Opaque source-owned event reference; no copied source payload | organization_id scopes customer connector configuration | Salesforce or Snowflake remains authoritative and access requires source policy | Reference expires by contract; source-native record retention and deletion remain source-owned | D06-UC-01, D06-UC-04, D06-UC-06 |
| Backup | Backup catalog and vault owned by recovery service | backup_id | application_id, environment_id, key, and parent-backup references | organization_id partitions backup policy | Checksum, region, recovery point, encryption key, hold, and expiry are explicit | Restored only into isolation; expires or is cryptographically erased unless a valid hold retains it | D06-UC-03, D06-UC-05 |
| RestoreJob | Recovery store owned by restore orchestrator | restore_job_id | backup_id and new recovery environment_id | organization_id partitions restore authority | Idempotent job records checksum, RPO/RTO, verification matrix, and traffic decision | Recovery environment is promoted or destroyed; job evidence retained then expires by policy | D06-UC-03 |
| ExportJob | Export store owned by Workboard export service | export_job_id | application_id, requester, approver, and destination receipt references | organization_id and app tenant scope bound exported data | Scope, version, count, checksum, encryption, destination, and custody are immutable | Export staging copy is securely deleted after receipt; customer-delivered copy follows transferred custody | D06-UC-04 |
| DeletionJob | Lifecycle store owned by deletion orchestrator | deletion_job_id | application_id plus dependency and policy references | organization_id partitions lifecycle authority | Directed checkpoints enforce quiesce, export, revoke, decommission, delete, expire, and probe ordering | Job remains until every scoped checkpoint terminates; then retained as minimal evidence and expired by policy | D06-UC-04, D06-UC-05, D06-UC-06 |
| DeletionTombstone | Append-only tombstone store owned by lifecycle evidence service | deletion_tombstone_id | Opaque reference to deleted resource type and identifier hash | organization_id partitions evidence | Contains action, method, time, result, policy, and proof only; never stores deleted customer payload | Retained under minimal evidence policy and legal hold, then expires without resurrecting deleted data | D06-UC-04, D06-UC-05, D06-UC-06 |
| ConnectorRevocation | Connector control-plane store owned by revocation service | connector_revocation_id | Day 4 connector_id, version, and workload_grant_id references | organization_id partitions connector grants | Created only for separate app-revenue retirement, never for Workboard retirement | Retained with revenue retirement evidence; credential and derived cache are destroyed after denial proof | D06-UC-06 |
| ResidualAccessProbe | Evidence ledger owned by verification service | residual_access_probe_id | Opaque reference to lifecycle job, resource, endpoint, credential, or source probe | organization_id partitions evidence | Records expected and observed result, actor, scope, environment, time, and immutable run ID | Retained with assurance dossier, then expired under evidence policy | D06-UC-04, D06-UC-05, D06-UC-06 |
Operate a fleet across three planes
A fleet dashboard that merges every event into one authority can let support staff cross customer boundaries and can hide which owner must respond. HelixWorks correlates provider control plane, hosted runtime, and generated-application signals while preserving separate authorization, audit stores, retention, and customer ownership.
Shared join fields include customer organization, app, environment, release ID, artifact digest, deployment, policy version, trace ID, and timestamp. Generated-app tenant and end-user fields remain access-controlled and are not promoted into broad metric labels. A customer admin sees its fleet; HelixWorks staff see only operational scope granted by policy or a recorded support grant.
Correlate telemetry without confusing it with audit
Metrics alone cannot explain one request, while debug logs alone cannot prove an access decision after they expire. Telemetry provides operational signals—traces, metrics, and redacted logs—whereas audit records are durable, decision-grade events about access, policy, deployment, support, export, and deletion.
Propagate trace context across the gateway, app, authorization service, data store, queue, connector broker, and source proxy. Attach release digest and environment to traces. Keep high-cardinality tenant, user, resource, prompt, and content values out of global metric labels; put authorized investigation detail in protected traces or audit events. Test redaction with synthetic secrets and personal data markers before telemetry export.
Provider and app audit streams stay distinct but correlatable. Each append-only event includes event ID, actor, resource, scope, decision, policy version, trace ID, source plane, environment, occurrence time, observation time, and integrity metadata. Writers cannot rewrite history; retention changes, reads, exports, support access, and verification failures are themselves audited.
Define fleet SLOs, quotas, and cost controls
A platform-wide average can look healthy while one customer journey fails or one generated app consumes the fleet, so operators need per-service objectives and enforceable resource boundaries. A service-level indicator (SLI) is a measured user outcome; a service-level objective (SLO) is its target over a window.
| Control | Example | Required decision |
|---|---|---|
| Availability SLO | 99.9% successful authorized Workboard mutations over 28 days | Error-budget burn can slow release pace |
| Latency SLO | 99% of interactive requests below 750 ms | Scale or investigate before breach |
| Revocation SLO | 99.9% of controllable paths deny within 5 minutes | Escalate stale caches, sessions, streams, or jobs |
| Security invariant | 100% cross-tenant probes denied | Immediate incident; no error-budget trade |
| Quota | Per-app CPU, memory, requests, jobs, connector calls, storage, model tokens | Throttle or reject at the owning boundary |
| Cost budget | Customer/app/environment budget with forecast and alerts | Notify, constrain optional work, require approval for increase |
Use multi-window burn alerts for rapid and sustained SLO consumption. Quotas protect shared capacity but do not replace authorization; a request under quota can still be forbidden. Attribute cost by organization, app, environment, release, workload class, model/tool, connector, storage, and egress without exposing tenant content. A budget breach may pause agent generation or batch work, but it must not silently disable required security logging or backups.
Prepare incidents and bounded break-glass
An improvised emergency response can destroy evidence or turn provider support into standing customer-data access. The incident plan assigns command, operations, security, communications, customer liaison, and evidence roles before a page occurs, and break-glass remains narrow, expiring, approved, and audited.
A break-glass request names ticket, incident, human actor, customer, plane, resource, actions, reason, approvers, start, expiry, and recording requirements. It uses strong authentication and just-in-time credentials, pages the customer where policy requires, and cannot bypass audit. Control-plane access, runtime shell access, and generated-app data access are separate grants. Closeout revokes each credential and proves denial; a ticket marked closed is not evidence of revocation.
Back up and restore the right state
A backup job can be green while its contents are incomplete, cross-region policy is wrong, or restoration cannot meet the customer’s recovery objective. Recovery point objective (RPO) bounds acceptable data loss; recovery time objective (RTO) bounds acceptable restoration time, and both are defined per state class.
Inventory generated-app database, object storage, configuration needed for recovery, encryption-key dependencies, audit data, and evidence indexes. Preview data, production app data, source-system data, and the immutable production artifact have different owners and backup needs. Do not copy Salesforce or Snowflake source data into backups unless the customer contract and minimization policy explicitly require it.
Encrypt backups with separately controlled keys, restrict and audit access, validate residency, make retention and legal holds explicit, and protect catalogs from deletion with the workload. A restore creates a new isolated recovery environment first. Verify checksums, schema and migration state, tenant counts, equal-local-ID isolation fixtures, object references, application authorization, app-scoped connector state, and audit continuity before switching traffic. Workboard has no Day 4 revenue connector grant, so its restore must positively prove those separate grants remain unchanged. Artifact rollback and database restore remain separate approvals.
Govern residency, retention, export, and deletion
Deleting a project record does not prove that deployments, app data, backups, connector grants, or exported copies disappeared. HelixWorks maintains a lifecycle inventory for every data class and executes customer-approved policy at each owner boundary.
| Data class | Owner and location | Retention/deletion control |
|---|---|---|
| Project source and agent history | Provider control plane, selected region | Project policy, holds, export, deletion tombstone |
| Preview artifact and data | Preview environment | Short TTL, explicit promotion rules, environment deletion |
| Production artifact | Artifact store | Immutable retention for rollback/evidence, then policy expiry |
| Production app data and objects | Generated-app plane | Customer retention, export, tenant/app deletion workflow |
| Source-system data | Salesforce organization or Snowflake account | Source owner’s policy; revoke HelixWorks grants and minimize copies |
| Audit and evidence | Separate provider/app stores | Longer governed retention, legal holds, integrity and expiry proof |
| Backups | Declared regions and vaults | RPO/RTO policy, holds, scheduled expiry and catalog proof |
Residency controls placement of processing and storage; it does not by itself authorize access. Retention expiry does not override a valid legal hold, and a hold does not create support access. Export records scope, schema/version, time range, object count, checksum, encryption, destination, requester, approver, and immutable receipt. Revocation cannot recall an export after customer delivery, so the receipt transfers custody explicitly.
Primary lab: incident, restore, then retire Workboard
An operations course that stops after recovery leaves the hardest lifecycle claim untested, so one executable workflow carries Workboard from incident detection through verified restoration and final retirement. The lab ends only when residual probes show no controllable access and backup/retention obligations have reached their declared terminal state.
Execute the workflow with test tenants Alpha and Beta:
- Inject a bounded Workboard failure that burns the mutation SLO and plants a synthetic integrity marker. Declare the incident, preserve trace/audit IDs, stop risky release expansion, and communicate scope.
- Request one runtime break-glass grant. Confirm it cannot access HelixWorks organization policy or generated-app tenant data outside the approved command set; record every command; expire and revoke it.
- Restore the latest eligible production backup into an isolated recovery environment. Measure RPO/RTO, verify checksum and schema, exercise Alpha/Beta equal-local-ID isolation, application roles, queues, objects, and audit continuity, then make an explicit traffic decision.
- After service acceptance, begin retirement: announce the cutoff, block new Workboard memberships and sharing grants, freeze writes, drain jobs, and create the final customer export with manifest, checksum, encryption, and receipt.
- Disable custom and platform routes, remove DNS/TLS bindings, scale workloads to zero, revoke Workboard runtime identity, secrets, webhooks, API keys, sessions, preview links, application grants, and project collaborators as scoped by the retirement order. Do not emit a connector revocation for
sf-pipelineorsnow-bookings, which belong toapp-revenue. - Delete production and preview app data, objects, caches, indexes, queues, artifacts after their approved evidence window, and provider project records; write separate tombstones because project deletion is not transitive proof.
- Advance simulated policy time or use an approved accelerated test class to observe backup, audit, and retained-copy expiry. Preserve only hold-authorized records and prove the hold’s scope and access controls.
- Probe former Workboard URLs, custom domain, APIs, object URLs, queues, workload credentials, app sessions, preview links, project APIs, backup catalog, and search/telemetry indexes for denial or policy-authorized tombstone access. Separately execute positive controls through the connector broker to prove
sf-pipelineandsnow-bookingsstill serveapp-revenue. - In a separate revenue-retirement reference flow, quiesce
app-revenue, revoke only its workload and delegated connector grants, purge derived caches, and prove broker denial while Salesforce and Snowflake source data and source-native audits remain owned and retained by those sources.
Assemble the assurance dossier
Scattered dashboards cannot support a defensible customer decision, so the final dossier joins claims to immutable observations without merging authorization domains. It reports successes, failures, residual custody, holds, exceptions, and the tested envelope.
Each evidence row contains actor, resource, scope, precondition, expected, observed, environment, observedAt, and an immutable traceId, runId, artifactDigest, eventId, or receipt ID. The dossier includes:
- fleet inventory with app, owner, region, release digest, SLO, quota, budget, and lifecycle state;
- correlated but separately authorized provider, runtime, connector, and generated-app telemetry/audit;
- alert, incident declaration, containment, revocation, communication, and break-glass command evidence;
- backup identity, checksum, key dependency, restore run, RPO/RTO, and tenant/application verification;
- final export manifest, checksum, encryption and custody receipt;
- disabled route, DNS/TLS, workload and deployment observations;
- Workboard workload, session, link, membership, API-key and support-grant revocation observations, plus separate revenue-retirement connector revocation observations;
- data/object/index/queue/project deletion tombstones, legal-hold exceptions, and retention/backup expiry events;
- residual-access probe matrix showing every former path denied or intentionally retained under policy.
The final customer decision states operating, restored-with-limitation, retired-with-policy-retained-records, or retired-and-expired. It never says “deleted” when a backup or legal hold remains.
Apply fleet governance beyond one tracer bullet
One restored Workboard instance is not platform reliability evidence, so governance aggregates repeated fleet outcomes while preserving per-app accountability. Shared invariants run against Workboard, the revenue dashboard, and the public intake app; domain-specific SLOs and data policy remain with each application owner.
The model registry records model/revision, intended use, evaluations, data terms, region, and deprecation. The tool and connector registries record schema, owner, permissions, side effects, approvals, rate limits, kill switches, and exit plans. Policies and exceptions are versioned, independently approved according to risk, time-bounded, and linked to deployments and evidence. Humans remain accountable; an agent supplies work and observations but never accepts residual risk.
Fleet statistics can support a platform claim only for the measured sample, period, workload, and controls. Passing all three archetypes exercises the intended envelope; it does not establish arbitrary application correctness.
Further reading
Operations and deletion controls copied from product marketing can overstate recovery or erasure, so use primary specifications and government guidance. These sources were accessed on 2026-07-28.
- OpenTelemetry Specification — trace, metric, log, resource, and context models for cross-service telemetry.
- NIST SP 800-34 Rev. 1 — contingency planning, recovery strategies, testing, and plan maintenance.
- NIST SP 800-88 Rev. 2 — media sanitization and evidence appropriate to the selected sanitization method.
- NIST SP 800-53 Rev. 5 — primary control catalog covering audit, contingency, access, incident, and system integrity controls.
- NIST AI RMF 1.0 — govern, map, measure, and manage functions for AI risk.
Key takeaways
Operating a generated-app platform means controlling the entire fleet lifecycle, not celebrating a release. Recovery, governance, and retirement are credible only when every boundary produces scoped evidence and retained data is named honestly.
- Correlate provider, runtime, connector, and app signals without merging their permissions or retention.
- Use customer-journey SLOs, hard security invariants, quotas, and cost budgets together.
- Make break-glass narrow, expiring, recorded, independently revoked, and customer-visible where required.
- Restore into isolation and prove tenant, authorization, data, and audit integrity before traffic.
- Finish retirement with export, decommission, grant revocation, deletion tombstones, retention/backup expiry, and residual-access probes.
Checklist
Retirement is incomplete while any route, workload, grant, data copy, or backup remains unaccounted for, so every checked item must cite a current evidence row. A project-deleted event cannot satisfy the rest of this list.
- [ ] Provider, runtime, connector, and app telemetry/audit are correlatable but separately authorized and retained.
- [ ] Fleet SLOs, hard security invariants, quotas, costs, owners, alerts, and runbooks are active.
- [ ] Incident containment, communication, break-glass expiry, and independent revocation were observed.
- [ ] Restore met or measured RPO/RTO and passed checksum, schema, tenant, role, object, queue, and audit tests.
- [ ] Final export has scope, manifest, count, checksum, encryption, destination, approval, and custody receipt.
- [ ] Platform/custom routes, DNS/TLS bindings, workloads, jobs, and deployments are disabled.
- [ ] Workboard workload, secret, webhook, API-key, session, link, membership, project, and support grants are revoked; revenue connector grants remain active until the separately evidenced
app-revenueretirement. - [ ] Production/preview data, objects, caches, indexes, queues, artifacts, and project records have separate deletion tombstones.
- [ ] Backup, audit, evidence, and held-copy expiry or authorized retention is observed and named.
- [ ] Residual probes cover former URLs, APIs, objects, queues, identities, connector sources, catalogs, and indexes.
- [ ] Every evidence row carries actor, resource, scope, precondition, expected and observed result, environment, timestamp, and immutable identifier.