Scale is rarely a surprise—it is the accumulation of early schema choices, queue semantics, and tenancy boundaries. These notes capture habits that keep SaaS platforms coherent as load and surface area grow.
Tenant model first
Decide whether you are row-level scoping, schema-per-tenant, or hybrid before feature velocity locks you in. Document how migrations roll across tenants and how backups restore per customer.
Bake tenant identifiers into every persistence boundary—jobs, caches, and analytics pipelines—not only the web tier.
Data lifecycle and retention
Ship deletion and export paths alongside onboarding. Regulators and enterprise procurement teams will ask before you are ready.
Partition large tables early where natural time or tenant keys exist; retrofitting hot partitions under traffic is expensive.
Queues, idempotency, and retries
Treat webhooks and integrations as at-least-once delivery. Idempotency keys and dedupe stores prevent duplicate invoices or emails.
Codify retry budgets with jitter so cascading failures do not thunder against fragile partners.
Operability as a feature
Structured logs with tenant and trace identifiers turn vague “slow today” reports into actionable timelines. Pair metrics with SLOs customers actually feel—checkout latency, invite acceptance, sync freshness.
