Data lifecycle
Data exports for a user or an organization as NDJSON files in Storage, and organization deletion with a grace period, a cancel and a purge job.
The data-lifecycle block covers two requests every SaaS gets: "send me my
data" and "delete our account". Exports collect a user's or an
organization's rows from every installed module and your own tables into
one NDJSON file per table in a private bucket. Deleting an organization
disables the tenant right away and purges its data after a grace period,
so an owner can still cancel.
better-supabase sql add data-lifecycle # adds tenant and access as well| Table | Holds |
|---|---|
data_exports | subject (user or organization), status, the object paths in files, and expires_at |
organization_deletions | requested_by, requested_at, purge_after, cancelled_at and purged_at |
| Permission | Lets a member | Default roles |
|---|---|---|
organization.export | export the organization's data | owner, admin |
organization.delete | request the organization's deletion | owner |
Every signed-in user can export their own data.
Tables
The module reads the tables of every other installed module from the module
registry: each module declares which of its tables hold a user's rows, which
hold a tenant's rows, and which the purge keeps. That covers memberships and
permission overrides, profiles, settings, comments and activity,
attachments, API keys, usage counters, events, quotas and history,
notifications with their deliveries, subscriptions and preferences,
onboarding progress, flag overrides, invitations, billing customers, SSO
domains, providers and SCIM users and members, support sessions, invite
codes, platform role assignments, outbox events, webhook endpoints,
deliveries and secrets, incoming webhook endpoints, inbox tables,
announcement dismissals, waitlist entries and redemptions, AI provider
keys, connectors and their grants, workflow credentials, chat installations
and the audit trail. A module added in a later release joins exports and purges
when you run sql sync, and a table you map to another name in
sql.modules.<name>.tables is read under that name. The module's own export
and deletion records stay out, API keys and webhook secrets are never
exported, and incoming webhook endpoints and invitations are exported without
their token hashes and secrets. Audit events are exported without the row
snapshots (old_record, new_record) and the impersonation details, which
describe other people's data and the support staff. Columns that hold a
credential_ref are never exported, and the purge revokes them first (see
Purging). An actor's outbox events are left out of user
exports, because their payloads describe other members too; the purge
deletes them with the tenant. Add your own tables with their user column (for user
exports) and tenant column (for organization exports and the purge):
export default defineConfig({
sql: {
modules: {
"data-lifecycle": {
options: {
tables: {
projects: { tenant: "organization_id", user: "owner_id" },
"billing.invoices": { tenant: "organization_id", purge: false },
},
grace: "30 days",
exportTtl: "7 days",
bucket: "data-exports",
},
},
},
},
});purge: false keeps a table's rows in the purge, for records you must
retain. omit: ["api_token", "password_hash"] leaves columns out of every
export file of that table, for credentials the export must not hand out;
the purge still deletes the rows. export: false keeps the whole table out
of exports, for tables that hold only credentials, while the purge still
deletes its rows. An entry for a table a module already contributes, such as
"better_supabase.invitations", replaces that module's entry, so you can
leave out more columns or keep it out of exports. The audit trail is never purged; its
retention removes old events.
A schema with many tenant tables doesn't need a list. autoTables adds every
table in its schemas that has the tenant column (and, with user, the user
column), read from the catalog when an export or a purge runs, so a new table
is covered without a config change. tables entries still win for the tables
they name, and exclude takes schema.table names or * globs:
options: {
autoTables: {
schemas: ["public"], // default
tenant: "organization_id", // default
user: "created_by",
exclude: ["public.*_archive", "public.invoices"],
purge: true,
},
tables: { "public.invoices": { tenant: "organization_id", purge: false } },
},tables: "auto" is short for autoTables: {}.
Exports
createDataExporter({ format: "csv" }) writes one CSV file per table instead
of NDJSON ({id}/{schema.table}.csv), with a header row, nested values as
JSON and text that starts like a formula prefixed with ', for customers who
open exports in a spreadsheet.
"use server";
import {
createDataLifecycle,
rpcTransport,
} from "better-supabase/blocks/data-lifecycle";
const supabase = await createServerClient();
const lifecycle = createDataLifecycle({
transport: rpcTransport(supabase),
storage: supabase.storage,
});
await lifecycle.requestExport().orThrow();
await lifecycle.requestExport({ organizationId }).orThrow();A request returns the open export for the same subject instead of starting
a second one, and writes data_export.requested to the
outbox. The exporter runs it with a service-role
client: it reads each table in pages, writes {id}/{schema.table}.ndjson
and marks the export ready.
import {
createDataExporter,
sqlTransport,
} from "better-supabase/blocks/data-lifecycle";
const exporter = createDataExporter({
transport: sqlTransport(postgres.admin),
storage: supabaseAdmin.storage,
});
await outbox.relay("data-exports", exporter.sink());The exporter writes concurrency tables at a time (4 by default), each in
pages of pageSize rows (1,000), and the file list keeps the table order.
A tenant with hundreds of tables then exports in a fraction of the time it
takes one table after another. With a postgres.admin pool each table
reads on its own connection, so keep concurrency below the pool size;
concurrency: 1 restores the sequential order. The first failing table
stops the others from starting, and the export is marked failed.
exporter.job is a jobs handler for an { exportId }
payload when you prefer a queue. A failure marks the export failed with
the error, and running it again starts over. An event or job for an export
that is no longer pending or failed (cancelled, deleted, already running or
ready) is done: the sink skips it and the job completes, so it never blocks
the events behind it in the relay. exporter.run() still returns not_found
with the hint DATA_EXPORT_NOT_FOUND for it.
data_export.completed carries the requester and the files, so a
notifications consumer can tell them it's
done. lifecycle.download(id) then returns a signed URL per file. The
bucket's policy serves the files only to the requester (or members with
organization.export for an organization export) until expires_at.
Expired exports stay in Storage until something removes them. Run
purger.purgeExports() from the same cron job as the purge (a
service-role purger with storage): it removes the files of exports past
expires_at (expired_data_exports), then their rows
(forget_data_exports), and returns how many it removed.
Deleting an organization
const deletion = await lifecycle
.requestOrganizationDeletion(organizationId, { grace: "14 days" })
.orThrow();
await lifecycle.cancelOrganizationDeletion(organizationId).orThrow();The request needs organization.delete, and disables the tenant through
the access contract: the organization's disabled_at with the managed
organizations module, or the sql.modules.access.disabled.tenant column.
Disabled tenants get no permissions and no membership claims, so members
lose access at once. grace takes an interval or a Temporal.Duration
and defaults to the grace option.
With sql.modules.data-lifecycle.permissions.deletePlatform set to a
platform permission (is_platform), platform staff can request and cancel a
deletion for any organization without the service role, such as from an
admin console.
The requester or an owner can cancel until the purge runs; cancelling
enables the tenant again, unless it was disabled before the request.
organizationDeletion(organizationId) returns the pending deletion to the
organization's members, for a banner with the purge date.
Purging
The purger runs the deletions past their grace period with a service-role client, from a cron route or a jobs queue:
import {
createOrganizationPurger,
sqlTransport,
} from "better-supabase/blocks/data-lifecycle";
const purger = createOrganizationPurger({
transport: sqlTransport(postgres.admin),
storage: supabaseAdmin.storage,
buckets: ["attachments"],
billing,
});
await purger.purgeDue({ limit: 20 }).orThrow();For each organization it checks that the deletion is due, cancels the
Stripe subscription (billing.cancelSubscription, when you pass
billing), revokes the organization's credentials, removes every object
under the organization's prefix in buckets, then calls
purge_organization. That function runs your
public.on_organization_purge(tenant) hook when it exists, deletes the
tenant's rows from each purged table (your tables first), deletes the
tenant's own row last and writes organization.purged. It returns the
deleted row count per table. A table whose rows another table still
references through a restricting foreign key, or whose delete makes an
on delete set null break a check on another table, waits for a later pass,
so the order of the tables doesn't matter and your foreign keys can stay;
when no pass makes progress, the purge stops with ORGANIZATION_PURGE_BLOCKED
and names the tables.
The tenant row is the organizations module's (managed or adopted), or the
row of the table sql.modules.access.disabled.tenant names, or the one in
options.tenantRow ("schema.table.column", false to keep it), so an
adopted tenant table needs no on_organization_purge hook to go.
A bucket name in buckets clears {organizationId}/. For buckets laid out
another way, pass { bucket, path } with a template, or a function that
returns the prefixes. Each prefix must contain the organization id, so a
purge never clears another tenant's objects:
buckets: [
"attachments",
{ bucket: "files", path: "orgs/{organizationId}/files" },
{ bucket: "media", path: (id) => [`public/${id}`, `private/${id}`] },
],Credentials
Pass credentials with your credential provider
and the purger revokes the credential_ref of every row the purge deletes,
such as AI provider keys, connector servers, connector grants (for the
grant's user) and workflow credentials, before it deletes the rows:
const purger = createOrganizationPurger({
transport: sqlTransport(postgres.admin),
credentials: vaultCredentials({
transport: credentialsTransport(postgres.admin),
}),
});
const [purged] = await purger.purgeDue().orThrow();
purged?.credentials.revoked; // 3
purged?.credentials.unrevoked; // [{ table, column, ref, subject, reason }]better_supabase.organization_credential_refs(tenant) lists the refs for
the purger and refuses a deletion that isn't due. The purger revokes a ref
only when it carries the organization in its tenant, the same check that
guards credential writes, so a purge never revokes another tenant's or the
app's own secret. The refs it leaves come back in credentials.unrevoked
with a reason, for you to revoke or keep:
| Reason | The ref |
|---|---|
foreign | carries no tenant or another tenant |
no_provider | is in the tenant, but the purger has no credentials provider |
not_revocable | belongs to a provider that can't revoke (env) |
A failed revoke stops the purge before it deletes anything, so running it
again retries. Chat installations keep their rows in the purge (purge: false): call uninstallTenant from the
Chat SDK channels before the purge, which revokes
their credentials.
Anonymizing
Some records must lose their personal data a while after a process ends
without being deleted, such as a candidate's details 180 days after the
application closed. options.anonymize lists those rules, one per table:
"data-lifecycle": {
options: {
anonymize: [
{
table: "public.candidates",
after: "180 days",
from: "process_ended_at",
unless: "{row}.talent_pool_consent_until > now()",
set: {
first_name: "Anonymized",
last_name: null,
email: null,
phone_hash: { sql: "md5({row}.phone)" },
},
markedBy: "anonymized_at",
},
],
},
},| Field | Meaning |
|---|---|
table | table or schema.table |
after | an interval: a row is due once from is older than this |
from | the timestamp column the interval counts from; a row with no value is never due |
unless | a condition on the row {row} that keeps a due row as it is while it holds |
set | the new value per column: a string, number, boolean, null, or { sql } with an expression on {row} |
markedBy | a nullable timestamp column the rule sets, so each row is anonymized once |
better_supabase.anonymize_due(max_rows) applies every rule to at most
max_rows (1000) due rows each and returns the rows changed per table. The
service role calls it through the purger, and createOrganizationPurger({ anonymize: true }) also runs it in job, after the due purges:
const anonymized = await purger.anonymizeDue({ limit: 500 }).orThrow();
// { "public.candidates": 12 }The rule changes columns only. Files in Storage that belong to the row,
such as a CV, stay where they are; remove them from an after update
trigger on markedBy or in your own job.
Events
With the outbox installed, the module writes these events with the tenant as
the partition key, except organization.purged, which has none because the
tenant is gone by then (its id stays in organizationId):
| Event | When | Data |
|---|---|---|
data_export.requested | an export is requested | exportId, subject, organizationId, userId, requestedBy |
data_export.completed | its files are written | the same, with files and expiresAt |
data_export.failed | the exporter failed | the same, with error |
organization.deletion_requested | a deletion is requested | organizationId, userId, purgeAfter |
organization.deletion_cancelled | it is cancelled | the same |
organization.purged | the purge deleted the organization | organizationId |
Last updated on
Attachments
Files linked to records in a tenant, uploaded through signed URLs to a private bucket and served only after a malware scan.
SSO
Verified email domains with auto-join, SAML providers per organization, SSO enforcement in the access token hook, and a SCIM 2.0 endpoint that provisions memberships.