EU AI Act Compliance

EU AI Act Synthetic Data: When Generation Triggers Compliance Obligations

Generating synthetic data with AI tools is not a compliance-free shortcut. Here is exactly when EU AI Act obligations kick in, which articles apply, and what your organisation needs to do before your next dataset build.

· 6 min read · By

EU-law graduate (Maastricht University) · MSc International Tax Law (AI & technology). Builds AI systems and advises SMEs on EU AI Act compliance.

Synthetic Data Is Not a Compliance-Free Zone

The ai act synthetic data question lands on compliance desks constantly: if we generate fake data instead of using real personal data, are we off the hook? The answer is no, not automatically. The EU AI Act imposes obligations on deployers based on what the AI system does, not on whether the output looks like real data. Generating synthetic images, text, audio, or video with a generative AI tool can trigger Article 50 disclosure requirements, and in some cases pulls your organisation into the high-risk category under Annex III.

This post walks through the three most common synthetic data scenarios for SMEs, maps each to the relevant AI Act articles, and gives you a plain-language checklist for each.


What the AI Act Actually Covers

The EU AI Act, which entered into force on 1 August 2024, applies to deployers: organisations that use an AI system in a professional context within the EU. Article 3 defines a deployer as any natural or legal person who uses an AI system under their own authority, other than for personal non-professional purposes.

Synthetic data generation tools, including large language models used to write customer scripts, image generators used to produce marketing visuals, and audio cloning tools used to create training samples, all qualify as AI systems under Article 3(1). The moment your organisation uses one of these tools for a business purpose, you are a deployer with obligations attached.

The specific obligations depend on the risk tier. Most synthetic data use cases sit in the general-purpose or minimal-risk category, but three obligations apply regardless of tier: Article 50 transparency requirements, Article 4 AI literacy duties, and Article 26 general deployer obligations.


Scenario 1: Marketing Creative and Synthetic Imagery

Your marketing team uses an AI image generator to produce lifestyle photos for a campaign. The models in the images are fully synthetic, no real person photographed.

Article 50(4) is the article to know here. It requires deployers that use AI to generate or manipulate images, audio, or video content to label that content as artificially generated or manipulated. The obligation applies even when no real person is depicted, because the regulation targets the nature of the content, not just deepfakes of real individuals.

Practically, this means:

  • Add a visible disclosure label to AI-generated campaign images ("Created with AI" is sufficient in most cases).
  • Retain records showing which assets were AI-generated, as part of your Article 26 documentation obligation.
  • If your campaign will appear on social platforms, check that platform-level disclosure rules align with Article 50 (most major platforms now require this independently).

The exemption in Article 50(4) covers content that is "obviously artistic or creative" where disclosure would be disproportionate, but courts and national authorities interpret this narrowly. Default to disclosing.


Scenario 2: Building Training Datasets with Generative AI

Your data science team uses a large language model to generate thousands of synthetic customer support queries. The goal is to train an internal classification model without using real customer data.

This scenario layers two AI systems: the generative model used to produce the dataset, and the downstream model trained on it. Your organisation is a deployer of the first system and likely also a deployer of the second.

Here the key question is whether the downstream model qualifies as high-risk under Annex III. If the classification model is used to make or assist decisions about natural persons in areas like employment, credit, or benefits, it almost certainly does. Training it on synthetic data does not reduce its risk classification. Risk is assessed by the model's function and context of deployment, not by the provenance of its training data.

For the generative step itself, Article 26 of the EU AI Act requires deployers to:

  • Use AI systems only in accordance with the provider's instructions.
  • Assign human oversight to monitor outputs.
  • Maintain records of use sufficient to demonstrate compliance.

If your organisation generates synthetic training data at scale, build a short internal log that captures: the tool used, the date range, the volume of outputs, the intended downstream use, and who reviewed the outputs for quality and bias. This log is your Article 26 paper trail.

AVG (GDPR) interaction is also relevant here. Even synthetic data derived from real customer interactions may retain statistical patterns that constitute personal data under Dutch and Belgian supervisory authority guidance. The Autoriteit Persoonsgegevens has signalled it will assess synthetic data on a case-by-case basis. Check whether a DPIA is required before you start generation.


Scenario 3: AI-Generated Customer Support Scripts

Your operations team uses a generative AI tool to draft scripts that human agents read aloud during customer calls. The voice is human, but the words are AI-written.

This is where Article 50 and Article 26 intersect in a way that surprises many compliance officers.

Article 50(1) requires that when an AI system interacts directly with natural persons, those persons must be informed that they are interacting with an AI. The exception: when it is obvious from context. A human agent reading an AI script is not an AI interaction in the Article 50(1) sense, so that specific disclosure is not triggered.

However, Article 50(2) covers AI-generated audio content used to influence persons. If scripts are converted to synthetic voice and delivered to customers, disclosure is required. Many organisations run pilot programmes with synthetic voice agents without realising this.

Practical steps for this scenario:

  • Audit your customer contact tooling. Identify every touchpoint where AI generates text or voice that reaches a customer.
  • For AI-scripted human calls: no Article 50(1) disclosure required, but document the process under Article 26.
  • For synthetic voice agents: add an opening statement identifying the caller as an AI system. This is not optional from 2 August 2026 when enforcement provisions fully apply.
  • Train customer-facing staff on what counts as AI interaction under the Act. This overlaps with your Article 4 AI literacy obligation, which applies from 2 February 2025.

The Article 4 Obligation You Cannot Skip

Article 4 requires all deployers to ensure their staff have sufficient AI literacy to use AI systems appropriately. This is not aspirational. It is a legal requirement that applied from 2 February 2025.

For synthetic data workflows, AI literacy means your team understands:

  • Which tools are in use and what risk tier they occupy.
  • What disclosure obligations attach to the outputs.
  • How to escalate when an AI output is uncertain or potentially harmful.

A one-page process guide per tool, combined with a short recorded briefing, satisfies the spirit of Article 4 for most SME contexts. Keep attendance records.


Your Compliance Checklist for Synthetic Data

Inventory first. List every AI tool used to generate text, image, audio, or video outputs in your organisation. This is the foundation of your Article 26 compliance.

Classify the output. Does it interact with or depict natural persons? Does it feed a downstream model that makes decisions about people? Each yes adds obligations.

Apply Article 50 disclosure. Any AI-generated content that reaches end users, customers, or the public needs a disclosure label or statement. Default to disclosing.

Check the downstream model's risk tier. If synthetic data trains a model used in HR, credit, or customer triage, run an Annex III check. High-risk classification means conformity assessments, human oversight systems, and registration in the EU database.

Document everything. A simple spreadsheet covering tool, use case, volume, oversight assigned, and disclosure method is your Article 26 record. Start it today.

Run a DPIA if in doubt. If synthetic data derives from real personal data or is used to build models that affect individuals, coordinate with your DPO and the AVG process before you begin.


If you are unsure which of your synthetic data workflows already triggers obligations, the fastest first step is a structured compliance check. Run the free 2-minute AI Act compliance check at comply.khairos.ai to get a prioritised view of where your organisation stands and which articles apply to your specific use cases.

Need help getting compliant?

The free 2-minute compliance check shows you exactly where your gaps are. No email gate to see your score.

Start the free check →