Where exactly do AI systems create privacy obligations that generic processing does not?
Four structural points. Training ingestion: assembling datasets from customer records, scraped sources, or logs is processing that needs a lawful basis independent of the original collection purpose, purpose compatibility under GDPR Article 6(4) is the gating analysis for repurposed internal data, and scraped data triggers Article 14 notice duties plus the special-category minefield of Article 9 when health, orientation, or political signals ride along. Model memorization: trained models can reproduce personal data from training sets; whether a model itself is personal data is case-by-case (the EDPB's opinion on AI models declines a blanket answer), which means extraction-resistance testing and output filtering are compliance controls, not just quality controls. Inference generation: AI creates new personal data, scores, classifications, inferred attributes, that inherits GDPR duties (accuracy under 5(1)(d), access rights, retention limits) even though nobody 'collected' it; inferred special-category data (predicted health status) is the sharpest edge. Automated decisions: Article 22 prohibits solely automated decisions with legal or similarly significant effects absent specific bases, and requires meaningful information about the logic plus human-intervention rights, deployment architecture (where the human sits and whether their review is real) becomes a legal control. 42001's impact-assessment and lifecycle controls give these four points a management-system home; the legal analysis at each stays with counsel.
What does EDPB Opinion 28/2024 mean for AI training and deployment?
The December 2024 opinion answered three questions national DPAs posed about AI models under GDPR. Model anonymity: models trained on personal data are not automatically anonymous; anonymity must be assessed case by case, considering the likelihood of extracting or regurgitating personal data with reasonable means, so 'the model is just weights' is not a defense, and documentation of extraction resistance becomes the evidentiary burden. Legitimate interest: Article 6(1)(f) can ground both development and deployment, evaluated through the standard three-step test (legitimate interest, necessity, balancing), with the opinion cataloguing balancing factors: data-subject expectations, the scale and nature of scraped data, mitigations like pseudonymization, output filtering, opt-out honoring, and transparency beyond the legal minimum. Unlawful training's downstream effect: where a model was developed with unlawfully processed data, deployment by the same or another controller is not automatically unlawful, but deployers must assess the model's provenance, effectively importing diligence duties into model procurement. Operational consequences: keep training-data provenance records granular enough to answer basis questions per source; run and document the three-step test before training, not after inquiries; build extraction-resistance evidence; and put provenance representations into model-vendor contracts, because your deployment analysis now legally depends on their training conduct.
Do we need both a DPIA and a 42001 AI impact assessment, and how do they compose?
Frequently both obligations apply to the same system, high-impact AI processing personal data triggers GDPR Article 35 (profiling, large-scale special categories, systematic monitoring all appear in the DPIA trigger lists) and falls squarely in the AIMS assessment scope. They differ in lens: the DPIA assesses risks to data subjects' rights and freedoms from the processing; the 42001 assessment adds group-level effects (cohort bias, differential accuracy) and societal effects, and covers AI-behavior risks (drift, misuse, overreliance) beyond data protection. Compose rather than duplicate: one assessment artifact per AI system with a common core (system description, data flows, stakeholders, risks, mitigations) and regime-tagged sections, the Article 35 mandatory content (systematic description, necessity and proportionality, risks, measures) explicitly labeled so a DPA request can be answered by extraction; the group/societal tiers satisfying the AIMS; and, for EU AI Act high-risk deployers, the fundamental-rights impact assessment content layered in the same document. Governance mechanics: shared triggers (design, material change, new context), one owner, joint review by privacy and AI governance functions, and version control that shows regulators the assessment preceded the deployment. The anti-pattern, separate DPIA and AI-assessment documents drifting on different revision cycles, reliably produces the contradiction an investigator quotes back to you.
How should an AIMS and a PIMS divide the work in one organization?
By object, not by team. The PIMS (27701) owns personal data as such: inventory and mapping (including training sets and inference outputs as PII stores), lawful-basis registers, rights handling (access and deletion requests that now reach into training pipelines), retention, transfers, processor contracts. The AIMS (42001) owns AI systems as such: the system inventory, three-tier impact assessments, lifecycle gates, model and data documentation, transparency about AI use, human-oversight criteria, third-party AI diligence. The seams need explicit joints: training-data governance is shared (PIMS answers 'may we use this data'; AIMS answers 'is this data fit and documented for this model'); DSAR deletion meets model reality at the seam (deleting a subject from a training set does not delete them from a trained model, so the joint policy must define what deletion means, retraining cadence, suppression, or documented impossibility with mitigations); automated-decision controls sit in both (Article 22 machinery in the PIMS, human-oversight design in the AIMS). Mechanically: shared management reviews or cross-attendance, one risk register with privacy and AI risk types, one evidence store, and integrated internal audits, the harmonized ISO structure makes this cheap. Companies running the two as rival programs discover the seams during incidents, which is the expensive discovery method.
What are the highest-value privacy controls to implement inside the AIMS?
Six, ranked by incident-prevention value. Provenance records per training dataset: source, collection basis, license or consent posture, special-category screening, and the purpose-compatibility conclusion, the artifact every downstream question (EDPB diligence, vendor representations, DPA inquiries) resolves against. Purpose-compatibility gate: a documented Article 6(4)-style analysis required before internal data is repurposed for training, with sign-off, this is where 'we had the data anyway' goes to die. Output controls for memorization: extraction-resistance testing before release, PII filters on generative outputs, and regurgitation monitoring in operation, both a privacy control and the anonymity-argument evidence the EDPB expects. Inferred-data governance: inference outputs about people enter the PII inventory with accuracy, retention, and access-right handling; inferred special-category data gets a deliberate legal decision, not a default. Human-oversight design for consequential decisions: defined decision categories, reviewer authority to override, review evidence, and appeal channels, built so Article 22's 'meaningful human involvement' is demonstrable rather than nominal. Model-vendor diligence: provenance and training-conduct representations, deployment-restriction flow-downs, and update-notification duties in AI procurement contracts. Each control produces an artifact; the artifacts are what turn 'we take AI privacy seriously' into something an auditor or regulator can verify.