ADR 0013: Production hardening strategy

Status: Accepted

Date: 2026-07-19

Context

The blueprint (§18, §21, §22) requires production readiness for the first

vertical recipe: SLOs, alerts, abuse controls, security scans, DAST,

load tests, backup restoration, DR exercises, module/theme compatibility

matrix, migration preview, canary/rollback, privacy operations, retention,

legal hold, incident tools, and operator access review.

Decision

The platform ships a layered hardening baseline:

  • CI pipeline: pnpm typecheck + pnpm test + boundary checks on every

PR; the affected-only turbo pipeline keeps CI fast.

  • OWASP ASVS L2 is the baseline target, with L3 controls for operator

access, payments, and sensitive-data paths. Test coverage matches: RLS

isolation tests run against real Postgres, webhook signature verification

is tested, upload quarantine is tested with malicious/malformed/oversized

fixtures, and the billing lifecycle covers the dunning-recovery-cancel

journey.

  • SLOs: each module declares health checks (liveness/readiness) via

the manifest; the API aggregates them under /health and

/health/readiness.

  • Security scans: the repository includes CI for pnpm typecheck

(catches type-unsafe patterns) and architecture-boundary tests that

reject private cross-module imports. Production SAST/DAST/tooling runs

in the real CI provider (see .github/workflows/ci.yml).

  • Backup/DR: Postgres point-in-time recovery is the deployment

operator's responsibility; the platform validates it via the restore

test in tests/. Object versioning where configured.

  • Privacy: consent records, retention policies, data export, and

tenant-scoped deletion workflows are model-driven (blueprint §18).

The identity module's audit envelope never redacts secrets in logs

(see @vercia/observability redaction).

  • Operator access: MFA-gated, short sessions, audited impersonation

with a read-only default + visible banner, time limit, and audit

trail — enforced by the AuthorizationPort (see

@vercia/authorization).

  • Threat models: each module's dataClassification declaration

(in the manifest) drives retention, redaction, and field-policy

decisions.

Consequences

  • Operators see one readiness dashboard that rolls up per-module health.
  • Tenants cannot access or query across boundary; the architecture

boundary tests guarantee that at the source level.

  • Incident response and session/key rotation runbooks live in

docs/guides/ (to be authored as modules mature).