Cross-Jurisdictional Global

Data Retention Requirements Worldwide: Limits, Mandates, and the Schedule

How retention law works across jurisdictions: GDPR storage limitation, the maximum-limits regimes, the sectoral minimum-retention mandates that conflict with them, deletion engineering, and building a defensible retention schedule.

Regulation

GDPR Article 5(1)(e), CPRA's retention disclosure duty, the 2025 COPPA retention limits, PIPL Article 19, LGPD Article 15-16, Quebec Law 25's destruction duty, plus sectoral minimums (tax, employment, AML, HIPAA, FINRA)

Max Penalty

Storage-limitation violations sit in the GDPR's 20 million EUR / 4% tier; retention failures are aggravators in most breach fines because data that should have been deleted enlarges every incident

Enforcing Authority

DPAs enforce storage limitation directly (the Danish DPA's Taxa 4x35 case, CNIL's retention findings in most fine decisions); the FTC's orders now routinely mandate retention schedules and deletion

Official Source

www.edpb.europa.eu

Executive Summary

  • Privacy law caps retention (keep no longer than the purpose requires) while sectoral law floors it (tax, employment, AML, and securities rules mandate keeping records for years), and a retention schedule exists to reconcile the two.
  • The GDPR's storage limitation principle is enforced directly: regulators have fined companies simply for holding data too long, and excessive retention aggravates nearly every breach penalty.
  • The modern statutes are converging on explicit duties: CPRA requires disclosing retention periods per category, the 2025 COPPA amendments ban indefinite retention of children's data, and Quebec mandates destruction or anonymization at purpose-end.
  • Anonymization is the lawful alternative to deletion everywhere, but the bar is high: GDPR-grade anonymization must survive reidentification analysis, not just drop the name column.
  • Retention fails in the backups, logs, data warehouses, and vendor copies long after the production database is clean; deletion has to be engineered as a data-lifecycle capability, not a policy PDF.

Retention is where privacy programs meet entropy: data accumulates by default, every system is a pack rat, and the law demands the opposite, a justified lifespan per category, honored by actual deletion, reconciled against the sectoral statutes that mandate keeping some records for years. The regimes have converged on the principle from every direction, storage limitation in Europe, disclosed periods in California, shortest-necessary in China, destroy-or-anonymize in Quebec, purpose-end deletion in Brazil, and enforcement has made the stakes concrete: companies fined for archives that could not delete, breaches made catastrophic by decade-old hoards, ‘anonymized’ datasets reclassified as retained personal data. What separates compliant programs is not the schedule document, everyone has one, but whether deletion is engineered: tagged data, event-anchored clocks, automated jobs with logs, backup and vendor strategies, and sampling audits proving the data is actually gone. A schedule that runs is a control; a schedule that doesn’t is a confession with a timestamp.

The capKeep no longer than the purpose requires (GDPR 5(1)(e), PIPL “shortest period”, Quebec destroy-or-anonymize)
The floorsTax 6-10y, AML/KYC 5y, securities 3-7y WORM, HIPAA 6y docs, employment per country, litigation holds
The disclosuresCPRA: retention period per category, at collection; COPPA 2025: published policy, no indefinite holding
The escapeTrue anonymization (survives singling-out/linkability/inference analysis), not identifier-dropping
Doctrine hubEDPB

Building it

Start from the map. Data mapping and inventory supplies the categories and stores the schedule governs.

Wire it to rights. DSAR automation covers deletion-request mechanics and the exception lanes.

Mind the vendors. Vendor risk assessments puts deletion certification into the processor chain.

Shrink the blast radius. Privacy breach incident response shows where hoarded data prices every incident.

Old tags and forgotten scripts hoard data too: see what your site still collects with a free scan.

Frequently Asked Questions

What does the storage limitation principle require, and how is it enforced?

GDPR Article 5(1)(e): personal data kept in identifiable form no longer than necessary for the purposes it was collected for, with archiving/research exceptions under safeguards, one of the seven foundational principles, backed by Article 13/14 duties to disclose retention periods or criteria in the notice, and by Article 17's erasure right operating alongside. What it demands operationally: a purpose-anchored period (or criteria) per data category and processing purpose, justified rather than habitual, 'we might need it' is precisely the reasoning the principle prohibits, and deletion or anonymization when the period ends. The enforcement record proving it is policed directly, not just rhetorically: the Danish DPA's Taxa 4x35 matter (2019) recommended a fine because the taxi company 'anonymized' rides by deleting names while retaining linked phone numbers for years beyond need, the canonical fake-anonymization case; CNIL's decisions routinely include storage-limitation counts (its Google, Carrefour, and numerous employer decisions cite retention beyond declared periods, and its guidance publishes reference periods, 3 years post-relationship for prospect data being the famous convention); the Berlin DPA's 14.5 million EUR Deutsche Wohnen fine (2019, later procedurally litigated to the CJEU on entity-liability questions) targeted an archive system that could not delete tenants' financial data at all, establishing that architecture incapable of deletion is itself the violation; the Irish DPC and peers cite excess retention as an aggravator in breach cases because every incident's blast radius includes the data that should have been gone; and the FTC reaches the same place through Section 5 and its orders, the Everalbum and Cambridge-adjacent matters mandating deletion (including model deletion), the 2023 amended orders and the 2025 COPPA rule imposing written retention limits, and its blog guidance repeating that indefinite retention is unreasonable. The composite standard a regulator applies: show me the schedule, show me the notice matches it, show me deletion actually ran, three artifacts, all checkable.

How do the other privacy regimes handle retention?

Convergent on the principle, increasingly explicit on the mechanics. The US state family: CPRA added a distinctive disclosure duty, at or before collection, state the retention period (or the criteria) for each category of personal information and sensitive personal information, and prohibits retaining longer than reasonably necessary for the disclosed purpose, which quietly forces Californian-facing businesses to actually build the schedule the disclosure summarizes; the Virginia/Colorado family embeds retention in data-minimization duties, and Colorado's rules require retention review in data protection assessments; the 2025 COPPA amendments prohibit indefinite retention of children's data and require a published retention policy, the FTC's clearest retention rulemaking. PIPL Article 19: retention for the shortest period necessary to achieve the purpose, with laws providing otherwise respected, the 'shortest period' formulation being notably stricter language than the GDPR's, and the PIS specification elaborating deletion-or-anonymization at purpose end plus account-deletion mechanics apps must honor. LGPD Articles 15-16: processing terminates when the purpose is achieved, the period ends, or consent is withdrawn, with deletion following termination subject to enumerated retention exceptions (legal obligation, research with anonymization, transfer to third parties per the law, exclusive controller use anonymized); ANPD enforcement has cited retention in its early sanction practice. Quebec Law 25: at purpose fulfillment, destroy or anonymize (per regulation and criteria), with the anonymization regulation (2024) defining the standard, one of the few statutes making anonymization an explicit lawful terminus. Canada federally: PIPEDA Principle 4.5's retention limits with the OPC's guidance, plus a distinctive wrinkle, information used to make a decision about a person must be kept long enough for the person to access it. APPI: purpose-bound retention through the utilization-purpose regime and a duty to endeavor to delete when use ends (Article 22); Australia APP 11.2: destroy or de-identify when no longer needed, the OAIC's post-Optus/Medibank enforcement posture treating hoarded legacy identity documents as the central failure of both breaches, Australia's loudest retention lesson. India's DPDP Act: erase when purpose is served or consent withdrawn unless law requires retention, with the rules adding time-bound erasure for large fiduciaries. The direction of travel is unmistakable: from principle to enforceable schedule, with notices disclosing numbers and regulators checking them.

How do minimum-retention mandates interact with privacy's maximums?

By statute trumping principle, and the schedule's job is to record which statute wins per row. The floor-setting regimes commercial programs hit constantly: tax and accounting (6-10 years for books, invoices, and supporting records in most jurisdictions, 7 years as a common convention, Germany's HGB/AO 8-10 years, US IRS practice 3-7 years by situation); employment (payroll, safety, and personnel-file rules per country, ranging 2-30+ years, exposure records under OSHA running 30 years); anti-money-laundering and KYC (5 years post-relationship in the EU's AMLD regime and the US BSA, some regimes longer); securities and financial-services books-and-records (SEC 17a-4 and FINRA's 3-6 year WORM-storage rules, MiFID II's 5-7 year communications retention, the CFTC's swaps records); healthcare (HIPAA's 6-year documentation floor for compliance records, with state medical-record retention laws running 7-10+ years and longer for minors); telecoms data-retention laws where they survive judicial review (the CJEU's invalidation line from Digital Rights Ireland through La Quadrature limits blanket mandates in the EU); and litigation holds, which suspend every schedule for potentially relevant data the moment litigation is reasonably anticipated, the common-law overlay US counsel administers. The reconciliation mechanics: the schedule resolves each data category to the longest applicable mandatory floor or the shortest defensible purpose-based cap, whichever governs, with the legal citation in the row; conflicts resolve by scope-narrowing rather than blanket extension, the AML statute requires keeping the KYC file, not the marketing profile attached to the same customer, so mandated-retention data is segregated (archived, access-restricted, excluded from operational processing) rather than left live, which is also what GDPR purpose limitation requires, keeping-for-tax is a different purpose with different access rules than keeping-for-marketing; and DSAR deletion requests meet the floors through the exception lanes every statute provides (Article 17(3)(b)'s legal-obligation exception, CCPA 1798.105(d)'s legal-compliance ground), with the refusal citing the specific mandate, scoped to the mandated records only. The failure pattern regulators flag: using one long sectoral floor as cover for retaining everything, the tax exception does not keep the clickstream.

What does anonymization have to achieve to end retention obligations?

Legal escape velocity, and the bar is higher than most implementations. The GDPR standard: recital 26 places truly anonymous data outside the regulation, with the test being whether any party can identify the individual by all means reasonably likely to be used, singling out, linkability, and inference all defeated, as elaborated by the WP29's Opinion 05/2014 and tightened in practice; pseudonymization (Article 4(5)) explicitly does not qualify, key-coded, hashed, or tokenized data remains personal data, and the Taxa 4x35 lesson generalizes: deleting the direct identifier while retaining quasi-identifiers (phone number, device ID, precise location, rare attribute combinations) is retitled storage, not anonymization; the CJEU's case law (Breyer's relative-identifiability reasoning, and the 2023-2025 SRB/EDPS litigation refining whose 'means' count for recipient-held data) keeps the analysis contextual, which cuts both ways but never blesses casual de-identification. The regime variants a global schedule must respect: Quebec's anonymization regulation requires defined criteria, documented process, and periodic reidentification-risk review for data anonymized 'for serious and legitimate purposes'; LGPD accepts anonymization as a processing terminus with reasonable-means analysis; PIPL's anonymization definition (irreversible, non-restorable) is strict on its face; CPRA's deidentification demands reasonable measures plus public commitment and contractual flow-downs against reidentification, a compliance-infrastructure requirement, not just a data transformation; HIPAA offers the two codified routes (Safe Harbor's 18 identifiers or expert determination), the most mechanical standard and the reason 'HIPAA-deidentified' does not automatically mean GDPR-anonymous. Engineering the real thing: aggregate where possible (cohorts, not rows), apply formal techniques where rows must survive (k-anonymity-family generalization, noise addition, differential privacy for query and telemetry systems), destroy the mapping keys, document the residual-risk analysis against the motivated-intruder test, and re-review as auxiliary data and techniques evolve; treat anonymization claims as assertions requiring evidence, because regulators do, and because the alternative reading, that the data was personal all along, means every downstream use inherited the violation.

How is deletion actually engineered so the schedule is real?

As a lifecycle capability with the same seriousness as backup, because a schedule without execution machinery is evidence against you (it proves you knew the deadline you missed). The mechanics stack: classification and tagging at ingestion, data mapped to schedule categories with retention metadata (category, basis, clock-start event, deadline) carried in or derivable for every store, since deletion jobs can only honor tags that exist; clock-start discipline, periods anchor on events (account closure, contract end, last activity, collection date) and the event must be observable in the data, 'X years after relationship ends' fails silently if nothing records the end; automated deletion jobs per store, scheduled, idempotent, logging what they deleted and, as importantly, what they skipped and why (legal hold, active exception), with the logs retained as the compliance evidence; and hard-case handling designed rather than ignored, backups (the accepted pattern: delete from live systems on schedule, let backup rotation age the copies out on a defined cycle, document the gap, and ensure restore procedures re-apply deletions so restored data does not resurrect the deleted), append-only logs and warehouses (retention-partitioned tables so dropping a partition is the deletion; pseudonymize-then-break-the-key where partitioning fails), ML training data and models (the FTC's model-deletion orders make 'the data is baked into weights' a compliance question to answer deliberately, via dataset lineage and retraining strategy), and vendor copies (DPA deletion clauses with certification, deletion instructions propagated through the processor chain on schedule and on DSAR, verified in vendor audits). Governance that keeps it real: the schedule owned jointly by legal (the law per row) and data engineering (the execution per row), versioned, mapped to the notice disclosures (CPRA makes mismatches visible), litigation-hold integration that suspends jobs surgically and lifts holds affirmatively, deletion-job monitoring with failures alarmed like any pipeline failure, and a periodic sampling audit, pick categories past deadline, prove absence across production, warehouse, and logs, which is the exact test a DPA, the FTC, or a breach post-mortem will run. The strategic frame worth repeating to budget owners: retained data is liability inventory, every breach, DSAR, and discovery request prices it, and the cheapest gigabyte in any incident is the one deleted on schedule the year before.

Regulatory Crosswalk

GDPR Article 5(1)(e)CPRAPIPL Article 19LGPDsectoral retention mandates

Organizations subject to this regulation often operate under these overlapping frameworks. BD Emerson maps controls across frameworks to reduce duplicated compliance effort.

Evaluate your compliance posture now

BD Emerson's automated scanner audits your public-facing properties against your applicable regulations in minutes, not weeks.