Privacy & Data Protection

Privacy Meets the EU AI Act: Data Obligations for Organizations Building or Deploying AI

Teams building or deploying AI systems that touch EU users now have two overlapping regimes to satisfy — existing data protection law, and the EU AI Act’s own data-related obligations, which aren’t identical even where they look like they should be.

01

Two regimes, not one

GDPR (and equivalent regimes elsewhere) governs how personal data is collected, processed, and…

02

What’s actually new under the AI Act

Data governance requirements specifically for training, validation, and testing datasets used in…

03

Where the two regimes genuinely conflict in practice

Data minimization under GDPR pushes toward collecting and retaining the least data necessary;…

Two regimes, not one

GDPR (and equivalent regimes elsewhere) governs how personal data is collected, processed, and protected, full stop — it doesn’t distinguish between data used to train a model and data used for any other processing purpose. The EU AI Act layers additional, AI-specific obligations on top: data governance requirements for training data quality and representativeness, transparency obligations about AI system use, and risk-tiered requirements that scale with how consequential the AI system’s decisions are. Satisfying one doesn’t automatically satisfy the other.

Existing lawGDPR / DPDPACollection, processing, protectionData minimization by defaultApplies to all personal dataAdded layerEU AI ActTraining data governanceAI-use transparency obligationsRisk-tiered by consequence
Satisfying one doesn’t automatically satisfy the other — demonstrating training-data representativeness can even pull against minimization, which is why this needs deliberate design, not a single compliance checklist.

What’s actually new under the AI Act

  • Data governance requirements specifically for training, validation, and testing datasets used in higher-risk AI systems — relevance, representativeness, and error examination, not just lawful collection.
  • Transparency obligations requiring people be informed they’re interacting with an AI system in specified contexts, which is a disclosure requirement distinct from privacy notice requirements under GDPR.
  • Risk tiering that determines obligation intensity — a high-risk AI system (used in employment decisions, credit scoring, and similar consequential contexts) faces materially heavier documentation and oversight requirements than a low-risk one.

Where the two regimes genuinely conflict in practice

Data minimization under GDPR pushes toward collecting and retaining the least data necessary; demonstrating training data quality and representativeness under the AI Act can push toward retaining more data, or more diverse data, to prove the dataset isn’t skewed. Neither requirement is wrong, but satisfying both at once requires deliberate design rather than defaulting to either principle alone — retaining a well-justified, representative sample for documentation purposes while still minimizing what’s used in live production processing is one common resolution.

A DPIA isn’t automatically an AI Act risk assessment, and vice versa. They ask overlapping but not identical questions. Organizations that already run DPIAs for AI features (as described in our piece on privacy by design and DPIAs) have a head start on AI Act documentation, but shouldn’t assume the existing DPIA fully covers the AI-specific risk assessment requirements.

A practical starting point

Inventory AI systems by what they actually do and who they affect, not by how sophisticated the underlying model is — a simple rules-based system making consequential decisions about people can carry more AI Act obligation than a sophisticated model used for internal, low-stakes purposes. That inventory, mapped against the AI Act’s risk tiers, is what determines which systems need the heavier documentation and governance work first.

Frequently asked questions

Does the EU AI Act apply to us if we’re not based in the EU?

Yes, in the same extraterritorial pattern as GDPR — the AI Act applies based on whether your AI system’s outputs are used within the EU, regardless of where your organization is headquartered or where the system was built.

Is a chatbot or basic generative AI tool automatically high-risk under the AI Act?

Not automatically — risk tiering depends on the use case and context, not the underlying technology category. A general-purpose chatbot used for customer support carries different obligations than the same underlying model deployed to make employment or credit decisions.

Do we need a separate legal basis for using personal data to train AI models?

Generally yes, under GDPR’s existing legal basis requirements — the AI Act doesn’t replace or waive the underlying lawful basis requirement for processing personal data in training, it adds obligations on top of it. This is an area worth confirming with privacy counsel given how actively the guidance is still developing.